Deploying and validating ASKCOS / Pistachio in the browser

Porting the original FP32 reaction model with step-by-step probability and search checks. Scheduling reduces two warmed task times by 28.5% and 33.8%; read-only state reuse reduces copying.

2026-10-06 · Mol2Mat · revision 1

Reaction candidate generation and migration validation

Reaction-prediction models read reactant SMILES and generate candidate products token by token. Running a model in the browser requires porting molecular preprocessing, the vocabulary, neural network and search procedure, then checking whether each step produces the same result. Prediction accuracy requires comparing those candidates with independent experimental ground truth. Reproducing the program and evaluating its chemistry therefore call for different evidence.

The port uses the ASKCOS Augmented Transformer checkpoint trained on Pistachio 23Q3 data, with the original FP32 CPU model as the numerical reference.[1] Checkpoint weights and the service revision identify the version being run. The platform also has a CTranslate2 int8 configuration, identified separately so that results remain traceable.

The comparison data contain 23 distinct inputs and 41 browser runs, with complete-vocabulary logits and search decisions saved at each decoding step.[4] Reusing read-only encoder state reduces copying, while changes to task scheduling shorten waits. Median times for two warmed Worker tests fall by 28.5% and 33.8%, with complete decoding decisions matching the earlier version.[5]

The five groups contain 8, 8, 8, 5 and 12 records, totalling 41 for 23 distinct inputs.
The v2 assessment contains 41 records over 23 distinct inputs. The original eight repeat on CPU, WebGPU and Safari simulator; additional CPU and holdout CPU sets contain five and twelve cases. Counts are migration-comparison records.

Chemical representation, sequence probability and search

RDKit 2026.03.6 parses and canonicalizes isomeric SMILES while preserving stereochemistry, isotopes and charge. Atom maps are removed from molecular objects. Reactants form a dot-separated mixture and use the pinned vocabulary; unknown tokens retain the original UNK mapping and produce warnings.

The model generates product tokens autoregressively. Beam search retains five paths, ranked by cumulative token log probability:

log⁡P(y∣x)=∑tlog⁡P(yt∣y<t,x).\log P(\mathbf y\mid\mathbf x)=\sum_t\log P(y_t\mid y_{<t},\mathbf x).

Native and browser execution share parent-beam, token, tie-ordering, EOS and maximum-length rules.

Candidate scores rank outputs within this model and search profile. EOS termination, maximum-length termination, chemical rejection and final acceptance are recorded separately. Experimental success and yield require evaluation with conditions and independent ground truth.

The figure shows two saved outputs. Acetyl chloride and ethylamine give an amide as the first candidate; isotope-labelled methanol on its own gives H₂. The browser reproduces both results. The second example also shows why evaluating chemical predictions requires reaction conditions.

An actual amide prediction and the anomalous H₂ candidate from an isotope-labelled input.
Saved first-ranked model candidates: acetyl chloride and ethylamine map to an amide above; isotope-labelled methanol alone maps to H₂ below. Isomeric SMILES determine structures; arrows denote model mappings.

Graph export of the encoder and incremental decoder

The original checkpoint has SHA-256 0fdec0430d2890b54c12e9c3a6ad593db97a735a202b109e0d1454958adcadf2; the service revision is 897f5c39a7262d857f3239ee898b815338cf474e. The OpenNMT-py 1.2.0 model has four encoder and four decoder layers, eight attention heads, hidden dimension 512 and feed-forward dimension 2048. The native reference environment is macOS Torch 2.5.1 CPU; the supplier environment is identified as Torch 1.12.1 Linux/MKL.

Export separates the encoder, first decoding step and cached decoding into three ONNX graphs, preserving FP32 weights and attention states. The pinned vocabulary and chemistry preprocessing jointly define inference with the graphs. Self-attention and cross-attention K/V buffers are checked separately during decoding.

The default browser path is single-threaded WASM CPU in ONNX Runtime Web 1.30.0, with separate RDKit WASM chemistry assets.[3] WebGPU coverage comprises the original eight cases; additional and holdout inputs use CPU. Evidence and configurations are identified by backend.

Hashes bind the model, vocabulary, graphs, runtime and Worker to an execution profile. Task management supplies cancellation, receipts and persistence. Cached candidates retain model identity, making outputs from different profiles traceable.

Comparing native and browser outputs at each layer

Comparing original FP32 Torch, native ONNX Runtime and browser outputs helps locate where discrepancies arise. Checks include encoder representations, complete-vocabulary logits at every saved step, self/cross-attention caches, beam parents, token paths, EOS, candidate order and cumulative scores. One original eight-case comparison contains 81903288 scalars.[4]

The 23 distinct inputs cover single species, mixtures, aromatic halides, salts, atom maps, chirality, isotopes, acylation, heteroaromatics, E/Z and controls lacking conditions. Forty-one records comprise eight original CPU cases, eight WebGPU cases, eight Safari simulator cases, five additional CPU cases and twelve CPU holdouts.

The twelve holdouts were computed after the versioned validation rules were defined and preserve independent full Torch and native ONNX decoding. Four constructed controls exercise ties, near ties, EOS and maximum length. This corpus evaluates migration and search boundaries; an experimental accuracy benchmark requires its own ground truth, split and matching rules.

Recorded coverage and distinct-input accounting
Execution groupRecordsInput scope
original-cpu8Original or additional CPU probes
original-gpu8Original eight on another backend
original-safari-simulator8Original eight on another backend
additional-cpu5Original or additional CPU probes
new-holdout-cpu12New holdout inputs
Total4123 distinct inputs

Checking search decisions alongside logit differences

The original elementwise rule is ∣a−b∣≤10−5+10−4∣a∣|a-b|\leq10^{-5}+10^{-4}|a|. The isotope input has three out-of-tolerance logits at zero-based steps 9, 12 and 13 in the old CPU report. WebGPU has five; native ONNX shows corresponding discrepancies, and Safari simulator CPU buffers match desktop CPU. These raw records are supplied alongside subsequent assessment.

Softmax maps logits to probabilities:

pi=ezi−max⁡jzj∑jezj−max⁡kzk,TV⁡(p,q)=12∑i∣pi−qi∣.p_i=\frac{e^{z_i-\max_j z_j}}{\sum_j e^{z_j-\max_k z_k}},\qquad \operatorname{TV}(p,q)=\frac12\sum_i|p_i-q_i|.

A constant row shift cancels under Softmax, motivating separate checks of raw-logit differences and decision-probability differences.

To check whether these differences change the search, v2 retains the encoder and cache tolerances and requires both maximum full-vocabulary probability error and row total variation ≤10⁻⁵. EOS masking at BOS, visited beam paths, final tokens, EOS, candidate structures and ordering must also match. Cumulative scores retain their original tolerance. This preserves the elementwise discrepancy records while directly checking probabilities and search decisions.

Across 41 records, maximum probability error is 3.3769×10⁻⁶ and maximum row total variation is 3.4359×10⁻⁶. All records pass v2 with matching search decisions. The comparison includes full-vocabulary probabilities at saved steps and the beam paths actually visited. New inputs and near-tied candidates still need step-by-step checks, since small probability changes can affect ordering.

Maximum probability differences and row total variations are below 10⁻⁵ for all 23 inputs.
Twenty-three distinct desktop CPU inputs; additional CPU records are used for repeated older inputs. Each point summarizes every saved decoding step and visited beam row, with EOS consistently masked at BOS. The eight-case WebGPU and simulator coverage is tabulated separately.
Preserved original criteria and v2 results
FieldAcceptance ruleRecorded result
Encoder and all cachesatol 1×10⁻⁵; rtol 1×10⁻⁴All pass
Full-vocabulary probability error≤ 1×10⁻⁵3.3768861e-06
Row total variation≤ 1×10⁻⁵3.4359221e-06
Paths, tokens, EOS, structures and ranksExact41 records agree
Old isotope logits outside toleranceOriginal criterion preservedCPU 3; WebGPU 5

Structural validity and prediction accuracy

RDKit checks 66 accepted candidates across 23 inputs, together with 148 chemistry-boundary fixtures for parsing and preprocessing. Checks cover sanitization, canonical isomeric SMILES, formulas and preservation of elements and isotopes. Experimental occurrence, suitable conditions and major-product assignment require corresponding experimental evidence.[4]

Saved candidates also show limitations of the original model: isotope-labelled methanol alone produces H₂; chiral lactate may produce a candidate with unspecified stereochemistry; a second acylation candidate can contain changes with insufficient condition support. [Cl-] may trigger the pinned vocabulary's UNK warning. The structure figures retain these outputs for assessment in their chemical context.

This model takes only reactant sequences as input, while solvent, catalysts, temperature and process conditions all affect real products. Evaluating top-k accuracy requires independent experimental ground truth and fixed rules for splitting, deduplication and product matching. Those choices determine how agreement between candidates and observed products should be interpreted.

Progress notifications and event-loop scheduling

A Worker task has to send progress, handle messages and return results as well as perform the calculation. Total time can be divided into:

Ttask=Tkernel+Tprotocol+∑k=1NyieldTyield,k.T_{\mathrm{task}}=T_{\mathrm{kernel}}+T_{\mathrm{protocol}}+\sum_{k=1}^{N_{\mathrm{yield}}}T_{\mathrm{yield},k}.

When the calculation itself is short, waiting repeatedly to yield the event loop can take most of the time.

The old task host awaited setTimeout(0) after each progress notification to give the event loop a chance to process cancellation. Repeated nested timers encounter a minimum delay, so these short waits accumulate with every notification. An isolation test kept the kernel and progress count fixed while changing only the yield mechanism, confirming the scheduling overhead.[5]

The new host schedules yields through a task-local MessageChannel, giving message handling a chance to run while avoiding repeated timer waits. It checks cancellation before and after each yield and preserves progress sequence, stages and result-commit order. On completion, it closes the ports and settles progress notifications. The host still controls termination of synchronous external kernels.

Read-only encoder state and beam reordering

At each step, beam search selects parent paths and reorders their states. Some of this copying can be avoided: encoder memory and cross-attention K/V depend only on the input reactants, so all five rows are identical after initial expansion. Self-attention K/V records the tokens already generated along each path and must follow the selected parents.

Let E denote encoder state, C cross-attention state and H_b the self-attention history of beam b. Within one input and encoder context:

Eb=E,Cb=C,Hb′=Hπ(b).E_b=E,\qquad C_b=C,\qquad H'_b=H_{\pi(b)}.

After initial expansion to beam width, E and C remain read-only and only H is gathered. Multi-input batches require state management by each row’s context identity.

Summing bytes copied by gather calls gives reductions from 14643200 to 5519360 and from 47185920 to 15851520, or 62.3% and 66.4%. Complete beam traces, EOS, candidates and scores agree, with FP32 weights, ordering and termination retained. This experiment measures copying; isolated model timings show no stable improvement. Worker pairs in the following section measure total task time.[5]

State-copy volume falls for both inputs with identical scientific decoding results.
Copy volume accumulated over gather calls in an actual-model Worker falls by 62.3% and 66.4%. MiB = 1048576 bytes. Separate Worker pairs measure task time.
Isolated accounting of state-reordering copies
InputOld copies / bytesNew copies / bytesReduction
acyl-primary-amine14643200551936062.3%
chiral-acylation471859201585152066.4%

Paired timings for two reaction inputs

Old and new distributions alternate in the same macOS Chromium session, with one warmup and three measured trials per version. The acyl-primary-amine median falls from 115.0 to 82.2 ms, a 28.5% reduction; chiral-acylation falls from 223.7 to 148.0 ms, a 33.8% reduction. The figure displays both medians and individual observations.[5]

Timing covers task protocol, progress, model calculation and result transfer. The cases retain 11 and 19 progress events respectively, with identical complete beam traces, EOS, candidates and scores. The main time saving comes from reducing repeated scheduling waits; read-only state reuse reduces separately accounted copying work.

The optimized distribution passes 17 broker / Worker numerical and lifecycle checks. Cancellation takes approximately 14.7 ms, retry succeeds and numerical-asset hashes match. Paired timings show faster tasks; the 23-input v2 comparison checks probabilities and decisions, with the original out-of-tolerance logits supplied alongside the data.

Measurements cover warmed CPU WASM tasks at these two decoding lengths, with n = 3 per version on a shared machine. Decoding length, growing attention state and beam termination jointly determine cost. Further comparisons should retain the same initialization and timing boundaries.

Optimized Worker tasks take less time with identical scientific outputs.
2026-10-10, macOS Chromium. Alternating old/new distributions, one warmup and n = 3 each. Timing covers task submission through result transfer, including protocol, progress and calculation. Bars: medians; points: observations.
Paired medians under a common timing boundary, n = 3
InputBefore / msAfter / msTime reduction
acyl-primary-amine115.082.228.5%
chiral-acylation223.7148.033.8%

Initial loading and repeated inference

FP32 model and runtime assets total approximately 155.9 MiB and are cached by hash after first loading. Later queries reuse these assets, so initial preparation and repeated inference are timed separately.

The original eight CPU cases record single wall times of approximately 0.0965–0.8026 s, with runtime preparation around 0.4346 s. These n = 1 probes include numerical sampling and checks. CPU and WebGPU preparation and sampling overheads are recorded separately; controlled performance comparison uses the common Worker boundary above.

The probe records 290717696 bytes of ONNX WASM linear memory and 16777216 bytes for RDKit. Tasks use a 120 s wall limit and a 1.5 GiB admission estimate; process RSS is a separate measurement. Inputs permit eight reactants, 100 heavy atoms, 512 characters and 256 source tokens.

Safari simulator supplies execution and numerical evidence for the original eight cases. Physical-phone CPU, memory and thermal behaviour require device measurements. Offline local execution depends on a complete cache of the necessary model, runtime and vocabulary assets.

Single-case wall times in the original CPU probe, n = 1
InputInstrumented probe / s
single-reactant0.1633
amide-reactants0.1457
aromatic-halogen0.1235
salt-charge0.0965
mapped-stereochemistry0.1620
isotope0.1924
medium0.2660
long-connected0.8026

Replayable probability, path and structural evidence

The bundle supplies 23 inputs, complete-vocabulary logits at every step, native and browser candidates and decisions, four constructed controls, 66 candidate chemistry checks, v2 rules and original out-of-tolerance reports. Compressed JSON retains full numerical values; CSVs summarize per-case probability error, total variation and candidates.[4]

Install NumPy after extraction and run python replay.py --article prediction to replay probability and decision checks. Repository tools/compute-build/prediction/export_original.py supplies graph export; assess_scientific_acceptance.py supplies the complete v2 assessment. Full encoder and attention buffers reside in repository archives; the public bundle supplies their hashes and field-level assessments.

Model rebuilding requires access to the pinned checkpoint and preparation of dependencies under their licences. optimization.json, optimization.csv and the independent optimization-replay.py supply paired observations, build identities, medians and scientific-output checks.[1,5]

反应预测模型的浏览器部署与数值验证:ASKCOS / Pistachio

移植原始 FP32 反应预测模型,逐步检查概率与搜索决策。调度优化使两例预热任务耗时降低 28.5% 和 33.8%,只读状态复用减少复制。

2026-10-06 · Mol2Mat · revision 1

反应候选生成与移植验证

反应预测模型读入反应物 SMILES,逐步生成候选产物。把模型放到浏览器中,需要一起移植分子预处理、词表、神经网络和搜索过程,并检查每一步是否得到相同结果。预测准确率则需要把这些候选与独立实验真值比较:程序复现和化学预测各有对应的评价依据。

这里移植的是 ASKCOS Augmented Transformer 服务中基于 Pistachio 23Q3 数据训练的检查点,数值对照使用原始 FP32 CPU 模型。[1] 模型权重与服务修订共同确定运行版本。平台另有 CTranslate2 int8 配置,两者分别标识,便于追溯结果。

对照数据包含 23 个不同输入和 41 次浏览器运行,保存了各解码步的全词表 logits 与搜索决策。[4] 在此基础上,复用只读编码器状态可减少复制,调整任务调度可缩短等待。两个预热 Worker 测试的中位耗时分别降低 28.5% 和 33.8%,完整解码决策与优化前相同。[5]

五组记录数为 8、8、8、5、12,合计 41,对应 23 个不同输入。
v2 评估含 41 次记录、23 个不同输入。原八例在 CPU、WebGPU 和 Safari 模拟器重复;另有追加 CPU 五例与留出 CPU 十二例。计数单位为移植对照记录。

化学表示、序列概率与搜索

RDKit 2026.03.6 解析输入并规范化异构 SMILES,保留立体化学、同位素和电荷。原子映射从分子对象中移除,反应物合并为点分隔混合物,再按固定词表分词;未知词沿用原 UNK 映射并生成警告。

模型自回归生成产物 token,beam search 保存五条路径,累计 token 对数概率决定排序:

log⁡P(y∣x)=∑tlog⁡P(yt∣y<t,x).\log P(\mathbf y\mid\mathbf x)=\sum_t\log P(y_t\mid y_{<t},\mathbf x).

原生与浏览器共同固定父 beam、token、并列排序、EOS 和最大长度规则。

候选分数在这一模型和搜索配置内用于排序。EOS 终止、最大长度终止、化学解析拒绝和最终接受分别记录。实验成功率与收率需由具有条件和真值的数据另行评价。

图中展示两个保存的输出。乙酰氯与乙胺输入得到的第一位候选是酰胺;单独输入 [13CH3]OH 时,第一位候选却是 H₂。浏览器复现了这两个结果,后一个例子也说明评价化学预测时需要结合反应条件。

实际预测的酰胺候选与同位素输入产生 H₂ 的异常候选。
保存的第一位模型候选:上为乙酰氯与乙胺生成的酰胺,下为单独 [13CH3]OH 输入生成的 H₂。异构 SMILES 决定结构,箭头表示模型映射。

编码器与增量解码器的图导出

原始检查点 SHA-256 为 0fdec0430d2890b54c12e9c3a6ad593db97a735a202b109e0d1454958adcadf2,服务修订固定为 897f5c39a7262d857f3239ee898b815338cf474e。OpenNMT-py 1.2.0 模型包含 4 层编码器、4 层解码器、8 个注意力头、512 维隐藏表示及 2048 维前馈层。原生参考环境为 macOS Torch 2.5.1 CPU,供应方环境标识为 Torch 1.12.1 Linux/MKL。

导出将编码器、首次解码与缓存解码拆为三个 ONNX 图,保存 FP32 权重及 attention 状态。固定词表和化学预处理与图共同组成推理定义。解码过程中分别核查 self-attention 和 cross-attention K/V 缓冲。

默认浏览器路径为 ONNX Runtime Web 1.30.0 的单线程 WASM CPU,化学处理使用独立 RDKit WASM。[3] WebGPU 的覆盖为原八例;追加与留出输入采用 CPU 路径。各后端的证据与配置分别记录。

模型、词表、图、运行时和 Worker 通过哈希绑定到执行配置。任务管理提供取消、回执与持久化,缓存候选保留模型身份,使不同执行配置的结果可追溯。

原生模型与浏览器的逐层对照

从原始 FP32 Torch 到原生 ONNX Runtime,再到浏览器,三套输出让误差发生的位置可以逐层检查。比较内容包括编码器表示、每个保存步的全词表 logits、self / cross-attention 缓冲,以及父 beam、token 路径、EOS、候选顺序和累计分数。原八例的一份比较含 81903288 个标量。[4]

23 个不同输入覆盖单一物种、多组分、芳香卤化物、盐、原子映射、手性、同位素、酰化、杂芳环、E/Z 和缺少条件的控制。41 次记录由原 CPU 八例、WebGPU 八例、Safari 模拟器八例、追加 CPU 五例和留出 CPU 十二例组成。

十二个留出输入在版本化验证规则确定后计算,并保存独立 Torch 与原生 ONNX 完整解码。四个构造控制检查并列、近并列、EOS 和最大长度。这一语料专用于移植与搜索边界验证;实验准确率基准需另行确定实验真值、分割与匹配规则。

记录覆盖与不同输入的计数
执行集合记录数输入范围
original-cpu8原始或追加 CPU 探针
original-gpu8原八例重复后端
original-safari-simulator8原八例重复后端
additional-cpu5原始或追加 CPU 探针
new-holdout-cpu12新增留出输入
合计4123 个不同输入

从 logits 差异检查搜索决策

原逐元素规则为 ∣a−b∣≤10−5+10−4∣a∣|a-b|\leq10^{-5}+10^{-4}|a|。同位素输入在旧 CPU 报告的解码步 9、12、13(从 0 计数)出现三个 logits 超限;WebGPU 有五个超限,原生 ONNX 存在对应差异,Safari 模拟器 CPU 缓冲与桌面 CPU 相同。这些原始记录与后续评估并列提供。

Softmax 将 logits 转为概率:

pi=ezi−max⁡jzj∑jezj−max⁡kzk,TV⁡(p,q)=12∑i∣pi−qi∣.p_i=\frac{e^{z_i-\max_j z_j}}{\sum_j e^{z_j-\max_k z_k}},\qquad \operatorname{TV}(p,q)=\frac12\sum_i|p_i-q_i|.

行常数偏移在 Softmax 下抵消,因此原始 logits 差与决策概率差分别检查。

为检查这些差异是否改变搜索,v2 在原编码器及缓存容差之外,要求全词表概率最大差与行总变差均 ≤10⁻⁵。BOS 时的 EOS 屏蔽、已访问 beam 路径、最终 token、EOS、候选结构及排序也须相同,累计分数沿用原容差。这样既保留逐元素差异的记录,也直接检查概率和搜索决策。

41 次记录的最大概率差为 3.3769×10⁻⁶,最大行总变差为 3.4359×10⁻⁶,全部通过 v2 检查,搜索决策也相同。比较包含保存步骤中的全词表概率和实际访问的 beam 路径;新输入及近并列候选仍需逐步检查,因为较小的概率变化也可能影响排序。

23 个输入的最大概率差和行总变差均小于 10⁻⁵。
桌面 CPU 的 23 个不同输入;旧输入的重复记录采用追加 CPU 记录。每点汇总所有保存的解码步与已访问 beam 行,并在 BOS 步一致屏蔽 EOS。WebGPU 与模拟器的八例覆盖另见表。
保留的原判据与 v2 结果
字段验收规则保存结果
编码器与全部 cacheatol 1×10⁻⁵; rtol 1×10⁻⁴全部满足
完整词表概率差≤ 1×10⁻⁵3.3768861e-06
行总变差≤ 1×10⁻⁵3.4359221e-06
路径、token、EOS、结构与排序精确一致41 次记录一致
旧同位素 logits 超限旧判据原样保留CPU 3;WebGPU 5

结构有效性与预测准确率的区别

RDKit 核查 23 个输入的 66 个接受候选,并配合 148 个化学边界夹具检查解析与预处理。项目包括结构消毒、规范异构 SMILES、分子式及元素和同位素保持。实验反应成立、条件适用性与主产物判断需要相应实验依据。[4]

保存的候选中也能看到原模型的局限:单独同位素甲醇生成 H₂,手性乳酸可能生成未指定手性的候选,酰化第二位候选可含条件依据不足的结构变化。[Cl-] 还可能触发固定词表的 UNK 警告。结构图保留这些输出,便于结合具体化学情境判断。

此模型只读入反应物序列,而溶剂、催化剂、温度和工艺条件都会影响真实产物。评价 top-k 准确率时,需要独立实验真值,并固定数据分割、去重和产物匹配规则,才能解释候选与实测产物的符合程度。

进度通知与事件循环调度

一次 Worker 任务除了计算,还要发送进度、处理消息和交回结果。总耗时可以分为:

Ttask=Tkernel+Tprotocol+∑k=1NyieldTyield,k.T_{\mathrm{task}}=T_{\mathrm{kernel}}+T_{\mathrm{protocol}}+\sum_{k=1}^{N_{\mathrm{yield}}}T_{\mathrm{yield},k}.

当计算本身很短时,反复让出事件循环所产生的等待,就可能占去大部分时间。

旧任务主机在每次进度通知后 await setTimeout(0),让事件循环有机会处理取消消息。连续嵌套计时器受最小延迟约束,这些短等待会随通知次数累加。保持内核和进度数量相同、只更换让步方式的隔离测试,确认了这一调度开销。[5]

新主机改用任务内的 MessageChannel 安排让步,让消息处理获得执行机会,同时省去反复使用计时器的等待。每次让步前后检查取消,进度序号、阶段和结果提交顺序保持一致。任务完成后关闭端口并结清进度通知;同步外部内核的终止仍由主机控制。

只读编码器状态与 beam 重排

beam search 每一步都会选择父路径,并重排对应的状态。这里有一部分复制可以省去:编码器 memory 与 cross-attention K/V 只取决于本次输入的反应物,首次扩展后的五行相同。self-attention K/V 则记录各路径已经生成的 token,必须随父 beam 重排。

设 E 为编码器状态,C 为 cross-attention 状态,H_b 为 beam b 的 self-attention 历史。在单输入、同一编码器上下文中:

Eb=E,Cb=C,Hb′=Hπ(b).E_b=E,\qquad C_b=C,\qquad H'_b=H_{\pi(b)}.

首次扩展到 beam 宽度后,E 与 C 保持只读,仅对 H 作 gather。多输入 batch 则需要按上下文身份管理各行状态。

按 gather 调用累加复制字节,两例分别从 14643200 降至 5519360 字节、从 47185920 降至 15851520 字节,减少 62.3% 与 66.4%。完整 beam trace、EOS、候选和分数相同,FP32 权重、排序与终止规则也沿用原定义。这个实验测得的是复制量;隔离的模型计时未显示稳定改善,总任务耗时则由下一节的 Worker 配对测量给出。[5]

两个输入状态复制量下降,科学解码结果保持一致。
实际模型 Worker 中按 gather 调用累计的复制量,两例减少 62.3% 与 66.4%。MiB = 1048576 bytes;任务耗时由独立 Worker 配对测量。
状态重排复制量的隔离计量
输入原复制 / bytes新复制 / bytes减少
acyl-primary-amine14643200551936062.3%
chiral-acylation471859201585152066.4%

两组反应输入的配对耗时

同一 macOS Chromium 会话交替执行旧、新分发包,各预热一次后测三次。acyl-primary-amine 中位耗时从 115.0 降至 82.2 ms,降低 28.5%;chiral-acylation 从 223.7 降至 148.0 ms,降低 33.8%。图中同时展示中位数和每次实测。[5]

计时覆盖任务协议、进度、模型计算与结果回传。两例分别保留 11 和 19 次进度事件,完整 beam trace、EOS、候选及分数逐项相同。主要时间收益来自减少反复调度等待,只读状态复用则减少单独记账的复制工作。

优化后的分发包通过 17 项 broker / Worker 数值和生命周期检查,取消约 14.7 ms 后重试成功,数值资产哈希也相同。配对计时显示任务更快;23 个输入的 v2 对照检查概率和决策,而原始 logits 超限记录仍随数据提供。

测量范围为这两个解码长度下的预热 CPU WASM 任务,每版本 n = 3,电脑同时承担其他工作。解码长度、attention 状态增长和 beam 终止共同影响成本,进一步的性能比较应保持相同初始化与计时边界。

优化后 Worker 任务耗时降低,科学输出保持一致。
2026-10-10,macOS Chromium;旧、新分发包交替执行,各预热一次后 n = 3。计时从任务提交至结果回传,覆盖协议、进度及计算。柱为中位数,点为实测值。
同计时边界的配对中位数,n = 3
输入优化前 / ms优化后 / ms耗时降低
acyl-primary-amine115.082.228.5%
chiral-acylation223.7148.033.8%

首次装载与重复推理

FP32 模型与运行时资源约 155.9 MiB,首次装载后按哈希缓存。后续查询复用这些资源,因此首次准备与重复推理的耗时分别记录。

原 CPU 八例的单次墙钟约为 0.0965–0.8026 s,运行时准备约 0.4346 s。这组 n = 1 探针包含数值取样和检查。WebGPU 与 CPU 各自的准备与取样开销分别记录,受控性能比较采用前述同边界 Worker 配对。

探针记录 ONNX WASM 线性内存 290717696 字节,RDKit 16777216 字节。任务采用 120 s 墙钟限制与 1.5 GiB 准入估算;进程 RSS 是另一个测量指标。输入限于 8 个反应物、100 个重原子、512 字符和 256 个源 token。

Safari 模拟器提供原八例的运行与数值证据。物理手机的 CPU、内存与热状态需要设备测量。离线本地执行依赖所需模型、运行时及词表已完整缓存。

原 CPU 探针的单次案例墙钟,n = 1
输入带检查探针 / s
single-reactant0.1633
amide-reactants0.1457
aromatic-halogen0.1235
salt-charge0.0965
mapped-stereochemistry0.1620
isotope0.1924
medium0.2660
long-connected0.8026

可重放的概率、路径与结构证据

数据包提供 23 个输入、每步全词表 logits、原生及浏览器候选与决策、四个构造控制、66 个候选化学核查,以及 v2 规则和原始超限报告。gzip JSON 保留完整数值,CSV 汇总逐例概率差、总变差和候选。[4]

解压后安装 NumPy,运行 python replay.py --article prediction 可重放概率与决策检查。仓库 tools/compute-build/prediction/export_original.py 提供图导出,assess_scientific_acceptance.py 提供完整 v2 核查。完整编码器与 attention 缓冲在仓库归档中,公开数据包提供其哈希及字段评估。

模型重建需要获得固定检查点,并按各依赖许可准备环境。optimization.json、optimization.csv 与独立 optimization-replay.py 提供配对观测、构建身份、中位数及科学输出检查。[1,5]

Data and sources