Performance optimization and numerical validation of browser-based CREST: Revision 2

Optimization of compilation, electronic-state management and s/p integrals, with performance comparisons against the native implementation and Revision 1.

2026-10-09 · Mol2Mat · revision 2

Overview

Revision 1 enabled browser execution of the original GFN2-xTB and CREST implementations and parallelized independent trajectories in quick conformer searches using Web Workers. Although parallel scaling was established, execution overhead within each computation instance continued to limit overall performance. Revision 2 addresses hot-object compilation, electronic-state restoration, integral evaluation and task scheduling while preserving the scientific parameters and sampling workflow. [1]

With ten child Workers on an Apple M1 Pro, median full-task times were 4.99, 8.76 and 24.53 s for ethanol, n-butane and n-hexane, respectively, corresponding to reductions of 27.3%, 40.2%, 45.5% relative to Revision 1. Each system was measured three times. Timing included the complete quick search, result validation and property recomputation for up to ten conformers. The browser revisions produced identical conformer XYZ files, random-seed sequences, energy-and-gradient evaluation counts, and retained conformer coordinates, energies and electronic properties.

The native implementation and Revision 1 were measured on 8 October 2026, and Revision 2 on 9 October. The comparison uses the same hardware, browser version and chemical inputs and constitutes a historical same-machine comparison. The contributions of individual optimizations were evaluated in separate paired experiments.

Performance bottleneck analysis

In Revision 1, increasing concurrency from one to ten Workers yielded an approximately 5.44-fold speedup for n-hexane under the same independent-trajectory protocol. The ten-Worker full task took 44.98 s, compared with 68.97 s for the native single-thread reference. These results indicate substantial benefit from concurrency, with further improvement requiring lower computation and data-transfer costs per instance. The diagnosis therefore examined linked objects, numerical hotspots and trajectory-state management.

Tracing the members of the linked archives showed that some tblite and CREST hotspots still used earlier O0-compiled objects; the O2 link setting had not provided the corresponding object-level optimizations. Integral evaluation for s/p shells also retained general array descriptors, temporary allocation and transformation wrappers, adding address-computation costs to fixed-size calculations. Missing alias constraints on output pointers further limited compiler reuse of angular-momentum and array-dimension information.

Additional trajectory overhead arose at four stages. Worker-pool capacity was set by the first batch and could not grow for larger subsequent optimization batches. Every trajectory copied and returned electronic state. Restoring a saved wavefunction first performed a self-consistent-charge (SCC) calculation whose result was subsequently overwritten. Search steps requesting only energy and gradients still evaluated electronic properties. These costs affected concurrency utilization, state transfer and redundant computation.

Figure 1 · Computational workflow and optimization targets
The upper row shows trajectory scheduling and state transfer; the lower row shows numerical evaluation and electronic-state restoration. Arrows indicate execution or data dependencies; diagram dimensions do not represent time fractions.

Compilation and workflow optimization

The intermediate representation (IR) of hot objects was recompiled at O2 while retaining existing application binary interface (ABI) compatibility adjustments. Archive updates replaced only reviewed members and preserved member order and duplicate object names. Compilation retained strict floating-point semantics, with fast-math, reassociation and fused multiply-add (FMA) disabled. Serial and parallel CREST cores used the same scientific objects, with identical trajectory lengths, SCF iteration limits, convergence thresholds and numerical precision.

Fixed-size integral implementations were generated for s/s, s/p, p/s and p/p shell combinations. They wrote directly into caller-provided output arrays and reused one-dimensional integrals within each primitive Gaussian pair. The noalias attribute was restricted to output arrays verified as non-overlapping at their call sites; read-only contracted Gaussian-type orbital (CGTO) inputs could still share a basis object. Higher angular momenta retained the general implementation. The p-orbital order, gradient raising relations and integral reference origins were preserved. [4]

Electronic-state restoration allocated the wavefunction container before importing saved state, eliminating the initialization SCC calculation overwritten by the restored data. Energy-and-gradient (E/G) search entries omitted unrequested property postprocessing, while final conformer charges, dipole moments and orbital properties were fully evaluated. The Worker pool expanded at batch boundaries within its resource budget and reused existing instances. Electronic state was returned only by trajectories requiring state continuity into subsequent stages. Stage synchronization, random sequences and result aggregation order were preserved.

Table 3 · Optimizations and preserved computational conditions
ComponentOptimizationPreserved conditions
Object compilationO2 recompilation of hot IR; selected archive-member replacementABI, floating-point semantics, convergence thresholds
IntegralsDirect output, one-dimensional integral reuse, noalias, fixed-size implementationsGeneral higher-angular-momentum implementation, orbital order, reference origins
Electronic stateSaved-state import, demand-driven property postprocessingSCC solution, E/G results, final properties
SchedulingWorker-pool growth, selective electronic-state returnResource budgets, stage synchronization, random sequences, aggregation order

Performance comparison across three versions

Figure 2 and Table 1 compare the native single-thread implementation, Revision 1 and Revision 2. Both browser revisions used ten child Workers. Each group contains three measurements; bars show medians and whiskers span the minimum to maximum. Native and Revision 1 values are the 8 October references, and Revision 2 values were measured on 9 October.

The ratios of Revision 1 to Revision 2 median times were 1.38, 1.67 and 1.83 for ethanol, n-butane and n-hexane, respectively. The ratios of native reference times to Revision 2 times were 1.58, 2.40 and 2.81. The latter compare wall times for different execution paths and do not directly measure per-core or parallel-scaling efficiency.

Native timing extended from process launch to complete search-file output. Browser timing additionally included result validation and property recomputation for up to ten conformers, while excluding downloads and parent-Worker initialization. Native searches retained the SCC state-cache chain, whereas browser searches initialized independent trajectories from round-entry electronic state; the two protocols may produce different finite conformer ensembles. [1–3] The browser-revision comparison thus uses consistent task boundaries and sampling protocols, while the native reference characterizes the observed runtime of a different execution path.

Figure 2 · Complete quick-task times across three versions
Each group has n=3n=3; bars show medians and whiskers span the minimum to maximum. Native and Revision 1 values are the 8 October references, and Revision 2 values were measured on 9 October. Both browser revisions used ten child Workers.
Table 1 · Full-task times across three versions
MoleculeAtomsNative / sRevision 1 / sRevision 2 / sReduction from R1R1/R2
Ethanol97.87786.85814.985927.3%1.38×
n-Butane1421.000914.63158.756840.2%1.67×
n-Hexane2068.970844.980724.525445.5%1.83×

Experimental setup and computational parameters

The test systems were neutral, closed-shell ethanol, n-butane and n-hexane, evaluated using GFN2-xTB and the complete CREST quick conformer search. [5,6] The hardware was an Apple M1 Pro with eight performance cores, two efficiency cores and 32 GiB of memory. The software environment comprised macOS 27.0 (26A428) and Chrome 154.0.8037.98 with JSPI enabled. Browser execution used ten child computation Workers and one coordinating Worker.

Table 2 lists the computational parameters and execution conditions of all three versions. The browser revisions used identical starting geometries, random-seed sequences and search workflows. The optimizations preserved sampling duration, convergence conditions and requested final properties. Removed work was limited to overwritten restoration precomputations and postprocessing not requested by search steps. Reported E/G evaluation counts exclude restore-container precomputations.

Table 2 · Computational parameters and execution conditions
ParameterNativeRevision 1Revision 2
Method / workflowGFN2-xTB / CREST quickGFN2-xTB / CREST quickGFN2-xTB / CREST quick
Software versionsCREST 3.0.2 / tblite 0.7.0CREST 3.0.2 / tblite 0.7.0CREST 3.0.2 / tblite 0.7.0
Electronic temperature300 K300 K300 K
Accuracy parameter / search SCF limit1.0 / 5001.0 / 5001.0 / 500
Charge / UHF / phase0 / 0 / gas0 / 0 / gas0 / 0 / gas
Execution concurrency1 native thread10 child Workers10 child Workers
Starting geometry / random controlIdentical RDKit geometry; seed 20261003, reset sequence +37×iIdentical geometry and seed sequenceIdentical geometry and seed sequence
Hot-object compilationGNU Fortran 16.2.0 / OpenBLASWASM objects compiled at mixed optimization levelsO2 hot-object recompilation; fixed-size s/p integrals
Final property recomputationOutside native search timingUp to 10 conformers; SCF limit 150Up to 10 conformers; SCF limit 150
Electronic-state strategyNative SCC cache chainIndependent trajectories initialized from round-entry stateSame protocol; optimized restoration and return
Precisionf64f64f64
Repetitions / date3 / 2026-10-083 / 2026-10-083 / 2026-10-09
Timing boundaryProcess launch to search outputSearch, validation, propertiesSearch, validation, properties

Performance contributions of individual optimizations

Paired experiments assessed compilation coverage, CREST core recompilation and integral optimization separately. Additional compilation coverage and demand-driven postprocessing reduced search time by approximately 1–2% in the corresponding experiment; CREST core recompilation reduced it by approximately 12–13%. The latest output-alias constraints and fixed-size integral optimization reduced complete-search time by approximately 7–8% in pairs using identical inputs and random sequences. The experiments used baselines from different stages, and interactions between changes preclude adding these percentages.

Figure 3 compares three kernel versions using captured n-hexane molecular-dynamics (MD) and conformer-optimization tasks. The experiment used a single Node computation instance, with warm-up followed by interleaved execution of the baseline, noalias-only, and combined noalias and fixed-size s/p versions. Each version was measured twice for MD and eight times for optimization. Noalias alone reduced median times by approximately 9.9% and 5.3%, respectively; the combined optimization increased these reductions to approximately 11.7% and 9.0%. Numerical results and electronic states were bitwise identical across versions.

Preceding profiles attributed approximately one quarter of the corresponding task time to integral-related code. The relationship between local acceleration and overall runtime can be expressed using the Amdahl model:

T2T1≈(1−f)+fk\frac{T_2}{T_1} \approx (1-f) + \frac{f}{k}

Here, T1T_1 and T2T_2 denote the overall runtimes before and after optimization, ff is the original time fraction of the optimized component, and kk is its local speedup. Remaining SCC work, optimization and stage synchronization limit the overall gain. Task replay estimates the local contribution of integral optimization, while complete-search pairs evaluate its effect within the search workflow.

Figure 3 · Task-level performance of integral optimizations
Captured n-hexane tasks were interleaved after warm-up in a single Node instance. Each version has n=2n=2 for MD and n=8n=8 for conformer optimization. Times are normalized to the corresponding baseline-kernel median; whiskers span the minimum to maximum.

Numerical agreement and execution validation

All nine full browser tasks terminated normally. Starting geometries, complete conformer XYZ files, random-seed sequences, E/G evaluation counts, and full-precision retained coordinates and electronic properties matched the corresponding Revision 1 results. The E/G counts remained 10,593, 13,071 and 22,385 for ethanol, n-butane and n-hexane, respectively. Same-coordinate native validation retained the Revision 1 candidate-conformer evidence and confirmed that the new results matched the corresponding coordinates, energies and properties.

Local integral validation covered 25 shell-pair combinations from s through g for both values and gradients. Maximum absolute differences were 2.78×10−162.78\times 10^{-16} and 7.77×10−167.77\times 10^{-16}, respectively, below the 10−1210^{-12} component tolerance. Butane, hexane and caffeine were also used to validate cold starts, state reuse, geometry updates, property requests and electronic-state restoration, covering seven phases per system. Caffeine served as a kernel numerical control; the CREST parallel measurements remained limited to the three benchmark systems.

The updated kernel was integrated into the local browser workflow and validated through actual Worker execution. Checks covered rejection under insufficient resource budgets, cancellation and retry, durable result storage and recovery after reload, and serial fallback without JSPI. Execution records were bound to the specific computational assets and numerical controls. These checks complement the numerical-agreement evidence and establish completion of the tested local workflows.

Table 4 · Numerical-agreement validation
ValidationScopeResult
Full tasks3 molecules × 3 runsInputs, XYZ, seeds, calls, retained coordinates and properties identical
Value / gradient integrals25 shell pairs eachMaximum differences 2.78×10−162.78\times 10^{-16} / 7.77×10−167.77\times 10^{-16}
State transitions3 molecules × 7 phasesBitwise E/G, wavefunctions and requested properties
Native same-coordinate validationRevision 1 candidate conformers and new execution tasksPassed at original tolerances; matching coordinates and properties

Discussion and further optimization

The performance improvement in Revision 2 reflects more efficient execution within each instance and greater concurrency utilization in later task batches. Ten Workers represent ten independent computation instances scheduled by the browser and operating system, rather than ten identical, pinned physical cores. Heterogeneous M1 Pro cores, background load and thermal state may affect wall time. The present results cover three neutral, closed-shell systems with three repetitions per group. Charged systems, other search presets and other devices require separate measurements.

Potential targets for further optimization include basis and workspace construction for new geometries, remaining SCC linear algebra, CREST optimizers and general integral paths. Candidate changes will be assessed using their profiled time fractions, numerical agreement and paired complete-search results. The concurrency configurations for GFN2 single-point calculations and ASE geometry optimization were not remeasured in this revision; their results remain those reported in Revision 1.

Data and reproducibility

The accompanying data include individual timings, version comparisons, random-seed sequences, conformer outputs, environment records, kernel and Worker checksums, and task-replay and compilation records. Reproduction scripts use independent output directories to preserve the original measurements. Public files retain all numerical results; absolute workspace paths have been anonymized.

CREST 浏览器实现的性能优化与数值验证:Revision 2

编译、电子态管理与 s/p 积分优化,以及与原生实现和 Revision 1 的性能比较。

2026-10-09 · Mol2Mat · revision 2

概述

Revision 1 实现了原生 GFN2-xTB 与 CREST 的浏览器运行,并通过 Web Workers 并行执行 quick 构象搜索中的独立轨迹。其并行扩展已得到验证,但单个计算实例的执行开销仍限制整体性能。Revision 2 在保持科学参数与采样流程不变的条件下,对热点代码编译、电子态恢复、积分计算和任务调度进行优化。[1]

在 Apple M1 Pro 上,采用十个子 Worker 时,乙醇、正丁烷和正己烷的完整任务中位耗时分别为 4.99、8.76 和 24.53 s,较 Revision 1 分别降低 27.3%、40.2%、45.5%。每个体系测量三次,计时包括完整 quick 搜索、结果校验及最多十个构象的性质复算。两版浏览器实现的构象 XYZ 文件、随机种子序列、能量与梯度调用次数,以及保留构象的坐标、能量和电子性质均一致。

原生实现与 Revision 1 的测量日期为 2026 年 10 月 8 日,Revision 2 的测量日期为 10 月 9 日。比较采用相同硬件、浏览器版本和化学输入,属于同机历史对照;各项优化的性能贡献另以配对实验评估。

性能瓶颈分析

Revision 1 中,正己烷在相同独立轨迹协议下由一个 Worker 扩展至十个 Worker,获得约 5.44 倍加速。十 Worker 的完整任务耗时仍为 44.98 s,原生单线程参照为 68.97 s。这表明并发调度已提供显著收益,而进一步优化需要降低单个实例的计算与数据传递开销。诊断因此同时覆盖链接对象、数值热点和轨迹状态管理。

对链接归档成员的追溯表明,部分 tblite 与 CREST 热点代码仍采用早期的 O0 编译对象;链接阶段设置 O2 未使这些对象获得相应优化。s/p 壳层的积分计算还沿用通用数组描述符、临时内存分配和变换封装,其固定规模计算承担了额外的寻址开销。输出指针缺少别名约束,进一步限制了编译器对角动量与数组维度信息的复用。

轨迹执行中的额外开销主要来自四个环节:Worker 池容量由首批任务确定,后续较大的优化批次无法扩容;各轨迹均复制并返回电子态;已保存波函数的恢复过程先执行一次随后被覆盖的自洽电荷(SCC)计算;仅请求能量与梯度的搜索步骤仍进行电子性质后处理。上述开销分别涉及并行资源利用、状态传递和重复计算。

图 1 · 计算流程与优化环节
上层表示轨迹调度与状态传递,下层表示数值计算与电子态恢复。箭头表示执行或数据依赖,图形尺寸不对应耗时比例。

编译与计算流程优化

热点代码的中间表示(IR)采用 O2 重新编译,并保留现有应用二进制接口(ABI)兼容处理。归档更新仅替换已审查成员,保留成员顺序及同名对象。编译保持严格浮点语义,禁用 fast-math、浮点重排和融合乘加(FMA)。串行与并行 CREST 核心使用相同科学计算对象,轨迹长度、SCF 迭代上限、收敛阈值和数值精度保持一致。

积分计算为 s/s、s/p、p/s 和 p/p 壳层组合分别生成固定尺寸实现,直接写入调用方提供的输出数组,并复用同一原始高斯函数对的一维积分。noalias 属性仅用于经调用点确认互不重叠的输出数组;只读收缩高斯型轨道(CGTO)输入仍允许共享基组对象。高角动量组合保留通用实现,p 轨道顺序、梯度升阶关系及积分参考原点不变。[4]

电子态恢复先分配波函数容器,再导入保存状态,消除被恢复数据覆盖的初始化 SCC 计算。搜索中的能量与梯度(E/G)入口省略未请求的性质后处理,最终构象的电荷、偶极矩和轨道性质仍完整计算。Worker 池在批次边界按资源预算扩容并复用已有实例;仅需向后续阶段传递状态的轨迹返回电子态。阶段同步、随机序列和结果汇总顺序保持一致。

表 3 · 优化措施与保持不变的计算条件
优化环节优化措施保持不变的条件
对象编译热点 IR 以 O2 重编译;指定归档成员替换ABI、浮点运算语义、收敛阈值
积分直接输出、一维积分复用、noalias、固定尺寸实现高角动量通用实现、轨道顺序、参考原点
电子态保存态导入、按需性质后处理SCC 求解、E/G 结果、最终性质
调度Worker 池扩容、按需电子态返回资源预算、阶段同步、随机序列、汇总顺序

三版本性能比较

图 2 与表 1 比较原生单线程实现、Revision 1 和 Revision 2。两版浏览器实现均使用十个子 Worker。每组包含三次测量,柱高为中位数,误差棒表示最小值至最大值。原生与 Revision 1 采用 10 月 8 日的参照数据,Revision 2 采用 10 月 9 日的测量数据。

Revision 1 与 Revision 2 的中位耗时之比分别为 1.38、1.67 和 1.83,对应乙醇、正丁烷和正己烷。原生参照耗时与 Revision 2 耗时之比分别为 1.58、2.40 和 2.81。后者比较不同执行路径的墙钟时间,不能直接用于推算单核效率或线程扩展效率。

原生计时从进程启动延续至完整搜索文件输出;浏览器计时还包括结果校验与最多十个构象的性质复算,但不包括下载和父 Worker 初始化。原生搜索沿用 SCC 状态缓存链,浏览器采用阶段入口电子态初始化的独立轨迹协议,两者可能生成不同的有限构象集合。[1–3] 因此,浏览器两版之间的比较具有一致的任务边界与采样协议,原生参照则用于描述不同实现的实际运行时间。

图 2 · 三版本完整 quick 任务耗时
每组 n=3n=3,柱高为中位数,误差棒表示最小值至最大值。原生实现与 Revision 1 为 10 月 8 日参照数据,Revision 2 为 10 月 9 日测量数据;两版浏览器实现均使用十个子 Worker。
表 1 · 三版本完整任务耗时
分子原子数原生实现 / sRevision 1 / sRevision 2 / s较 R1 降幅R1/R2
乙醇97.87786.85814.985927.3%1.38×
正丁烷1421.000914.63158.756840.2%1.67×
正己烷2068.970844.980724.525445.5%1.83×

实验设置与计算参数

测试体系为中性闭壳层的乙醇、正丁烷和正己烷,均采用 GFN2-xTB 与完整 CREST quick 构象搜索。[5,6] 硬件为 Apple M1 Pro(8 个性能核、2 个能效核),内存为 32 GiB;软件环境为 macOS 27.0(26A428)和 Chrome 154.0.8037.98,启用 JSPI。浏览器配置包含十个子计算 Worker 和一个协调 Worker。

表 2 列出三版本的计算参数与执行条件。两版浏览器实现采用相同起始几何、随机种子序列和搜索流程。优化未改变采样长度、收敛条件或最终请求的性质,减少的工作限于恢复过程中被覆盖的预计算及搜索步骤未请求的后处理。所报告的 E/G 调用次数不包含恢复容器时的预计算。

表 2 · 计算参数与执行条件
参数原生实现Revision 1Revision 2
方法 / 流程GFN2-xTB / CREST quickGFN2-xTB / CREST quickGFN2-xTB / CREST quick
软件版本CREST 3.0.2 / tblite 0.7.0CREST 3.0.2 / tblite 0.7.0CREST 3.0.2 / tblite 0.7.0
电子温度300 K300 K300 K
计算精度参数 / 搜索 SCF 上限1.0 / 5001.0 / 5001.0 / 500
电荷 / UHF / 相态0 / 0 / 气相0 / 0 / 气相0 / 0 / 气相
执行并发1 原生线程10 子 Worker10 子 Worker
起始几何 / 随机控制相同 RDKit 几何;种子 20261003,重置序列 +37×i相同几何与种子序列相同几何与种子序列
热点代码编译GNU Fortran 16.2.0 / OpenBLAS不同优化等级的 WASM 对象热点 O2 重编译;固定尺寸 s/p 积分
最终性质复算未纳入原生搜索计时最多 10 构象;SCF 上限 150最多 10 构象;SCF 上限 150
电子态策略原生 SCC 缓存链独立轨迹;阶段入口态初始化相同协议;优化恢复与返回
数值精度f64f64f64
重复次数 / 日期3 / 2026-10-083 / 2026-10-083 / 2026-10-09
计时范围进程启动至搜索输出搜索、校验、性质复算搜索、校验、性质复算

局部优化的性能贡献

配对实验分别评估了编译覆盖、CREST 核心重编译和积分优化的贡献。补充编译覆盖与按需后处理在相应实验中使搜索耗时降低约 1–2%;CREST 核心重编译使搜索耗时降低约 12–13%。最近一轮输出别名约束与固定尺寸积分优化,在相同输入和随机序列的完整搜索配对中使耗时降低约 7–8%。各实验使用不同阶段的基线,且优化之间存在相互作用,因此这些比例不具有可加性。

图 3 比较正己烷已捕获分子动力学(MD)与构象优化子任务在三个内核版本下的执行时间。实验在 Node 单实例环境中进行,预热后交错执行基线、仅 noalias 和 noalias 与固定尺寸 s/p 联合优化版本;MD 每版测量两次,构象优化每版测量八次。仅 noalias 使两类任务的中位耗时分别降低约 9.9% 和 5.3%;联合优化使降幅分别达到约 11.7% 和 9.0%。各版本的数值结果和电子态逐位一致。

此前的任务剖析显示,积分相关代码约占相应任务执行时间的四分之一。局部加速与整体耗时之间的关系可用 Amdahl 模型表示:

T2T1≈(1−f)+fk\frac{T_2}{T_1} \approx (1-f) + \frac{f}{k}

其中,T1T_1 与 T2T_2 分别为优化前后的整体耗时,ff 为被优化部分在原任务中的时间占比,kk 为其局部加速倍数。整体收益受剩余 SCC、优化器和阶段同步开销限制。子任务回放用于评估积分优化的局部贡献;完整搜索配对则用于评估该改动在实际搜索流程中的效果。

图 3 · 积分优化的子任务性能比较
正己烷子任务在 Node 单实例环境中预热后交错执行。MD 每版 n=2n=2,构象优化每版 n=8n=8;耗时以对应基线内核的中位数归一化,误差棒表示最小值至最大值。

数值一致性与运行验证

九次完整浏览器任务均正常终止。与 Revision 1 对应结果比较,起始几何、完整构象 XYZ 文件、随机种子序列、E/G 调用次数,以及保留构象的全精度坐标和电子性质均一致。乙醇、正丁烷和正己烷的 E/G 调用次数分别保持为 10,593、13,071 和 22,385。原生同坐标验证沿用 Revision 1 的候选构象证据,并核对了新版结果与对应坐标、能量和性质的一致性。

局部积分验证分别覆盖值积分与梯度积分各 25 组 s 至 g 壳层组合,最大绝对差为 2.78×10−162.78\times 10^{-16} 和 7.77×10−167.77\times 10^{-16},均低于 10−1210^{-12} 的分量容差。正丁烷、正己烷和咖啡因还用于冷启动、状态复用、几何更新、性质请求及电子态恢复验证,共覆盖每个体系七个阶段。咖啡因仅作为内核数值控制体系,CREST 并行测量范围仍为前三个测试体系。

更新后的内核已接入本地浏览器计算流程,并通过真实 Worker 执行验证。运行检查覆盖资源不足时的拒绝执行、取消与重试、结果持久化与刷新恢复,以及无 JSPI 条件下的串行回退;验证记录绑定具体计算资产及数值控制结果。这些检查补充了数值一致性证据,表明本地实现能够完成所测运行流程。

表 4 · 数值一致性验证
验证项目范围结果
完整任务3 分子 × 3 次输入、XYZ、种子、调用、保留坐标与性质完全一致
值积分 / 梯度积分各 25 个壳层对最大差 2.78×10−162.78\times 10^{-16} / 7.77×10−167.77\times 10^{-16}
电子态切换3 分子 × 7 阶段E/G、波函数、所请求性质逐位一致
原生同坐标验证Revision 1 候选构象及新版实际任务原有容差下通过;坐标与性质对应一致

讨论与优化方向

Revision 2 的性能提升来自单实例执行效率的改善与后续任务批次并发利用率的提高。十个 Worker 对应十个独立计算实例,其调度由浏览器与操作系统管理,不能等同于十个固定绑定、同速的物理核。M1 Pro 的异构核心、后台负载及热状态均可能影响墙钟时间。当前结果覆盖三个中性闭壳层体系,每组仅三次重复;对带电体系、其他搜索预设和其他设备的推广仍需独立测量。

后续优化的候选环节包括新几何下的基组与工作区构建、剩余 SCC 线性代数、CREST 优化器及通用积分路径。候选改动将结合热点占比、数值一致性和完整搜索配对结果进行评估。本轮未重新测量 GFN2 单点计算与 ASE 几何优化的并发配置,其结果仍以 Revision 1 为依据。

数据与复现方法

随文数据包括逐次计时、版本比较、随机种子序列、构象输出、实验环境、内核与 Worker 校验值,以及子任务回放和编译记录。复现脚本在独立输出目录运行,以保留原始测量数据。公开文件保留全部数值结果,绝对工作区路径已作匿名化处理。

Data and sources