Running AiZynthFinder in the browser with sharded stock
Running the original template policy, chemistry and tree search in the browser. Stock shards provide on-demand lookup of 17.4 million full keys, with native comparisons for six target molecules.
2026-10-10 · Mol2Mat · revision 1
From target molecules to stocked precursors
Retrosynthesis works backwards from a target molecule to its precursors, then continues through intermediates until all terminal molecules are found in the specified stock. AiZynthFinder uses reaction templates to propose disconnections, a neural policy to rank templates, and tree search to allocate a limited computation budget.[1,2]
Mol2Mat ports the USPTO configuration of AiZynthFinder 4.4.1 into one Dedicated Worker. The original FP32 ONNX policy ranks templates; 42554 templates and RDKit / RDChiral generate precursors; single-objective Monte Carlo tree search organizes routes. All these steps run locally in the browser.
The port also gives the browser access to 17422831 complete InChIKeys in historical ZINC stock. A Bloom index screens queries before exact lookup in shards fetched on demand, reducing the initial data load. Native comparisons for six molecules check fingerprints, 300 policy probabilities, 107 reaction operations and 52 stock claims.[3]
Molecular fingerprints and the original template policy
RDKit parses and sanitizes the target, generates canonical isomeric SMILES and removes atom maps from molecular objects. The policy uses a radius-2, 2048-bit nonchiral Morgan fingerprint:
Probabilities rank templates; subsequent full-graph operations handle molecular stereochemistry.
ONNX Runtime Web 1.30.0 uses a single-thread WASM execution provider and original FP32 weights. Selection retains at most 50 templates, with the original cumulative cutoff semantics at 0.995 and FP32 accumulation. Each of the six inputs retains 50 templates; all 300 probabilities and indices can be checked individually.[2,3]
RDKit 2026.03.6 and a pinned RDChiral C++ commit are compiled to WASM; the bridge manages inputs, outputs and resources. Selected templates act on full molecular graphs, with sanitization, stereochemistry and deduplication producing zero or multiple precursor combinations. Policy ranking and applicable outcomes are consequently recorded separately.[4]
Single-objective tree search and budget constraints
JavaScript search retains node selection, lazy expansion, backpropagation and cycle pruning. Initialized children use
V_i is accumulated value, n_i visit count and N total sibling visits; policy priors enter initialization. The exploration term allocates work between higher-average-value and less-visited branches.
State scoring retains the upstream definition:
Here f_stock is stock coverage and the depth term mildly favours shallower routes. This score evaluates search state; experimental feasibility, selectivity and separation cost require chemical evidence.
Measurements use 80 iterations and at most four steps. Tasks additionally have a 45 s search budget; the interface permits up to five steps. The browser uses xorshift32 seed 20261010, while the Python reference uses NumPy global randomness. Numerical and chemical layers receive item-level comparisons; visited trajectories and route rankings are interpreted under their respective random processes.
Sharded stock and exact membership checks
The historical ZINC snapshot contains 17422831 complete InChIKeys in an approximately 598 MiB SQLite source. The browser uses a 16 MiB Bloom index and 64 sorted compressed shards. Five hashes screen definitely absent keys; Bloom positives trigger binary search for the full 27-byte key in the corresponding shard. Exact lookup determines membership.[3,5]
Compressed shards total 129073867 bytes, approximately 123.09 MiB, and are retrieved on demand, verified by SHA-256 and persistently cached. Decoded shards use an LRU capped at 32 MiB. Full approved assets, including model and runtime, total approximately 250.87 MiB; eager core loading totals approximately 127.77 MiB. Per-query transfer depends on cache state and required shards. The .bin format preserves received bytes and hashes.
Offline queries require complete caching of their necessary shards; a missing shard returns a resource error, preserving membership semantics. The stock represents a fixed historical collection. Current suppliers, prices, lead times and purity require updated procurement sources.
Reusing state across searches
Parsing, policy inference, template application and tree search run serially in one Dedicated Worker. The page receives progress and complete results through the task protocol. Cancellation ends the current task; completed records can be exported and recovered after reload, retaining task, model hash and input identity.
On repeated searches, the Worker reuses its ONNX session, parsed templates, molecular identities and stock hits, reducing initialization work. Template and identity caches limit their entry counts; neural-prediction and reaction-result Maps are rebuilt for each search. An LRU manages decoded stock, while compressed assets remain in persistent storage.
The probe records 367394816 bytes of ONNX WASM linear memory and 33554432 bytes of chemistry WASM memory. The task’s 1 GiB value is an admission estimate. JavaScript objects, persistent storage, graphics and process RSS require separate accounting. Six frozen targets define the molecular scope of the performance evidence.
Native comparisons of policy, chemistry and stock
The six targets are aspirin, acetanilide, ethyl benzoate, chiral ibuprofen, an E/Z alkene and caffeine. Native and browser fingerprint bits and top-50 template indices agree exactly. Across 300 saved probabilities, maximum absolute difference is 3.8743×10⁻⁷, below the 10⁻⁶ tolerance.[3]
Native RDChiral rechecks template expansions and steps in the returned routes, giving 107 reaction checks. RDKit and read-only SQLite check terminal-molecule stock membership 52 times. All checks pass. Counts refer to individual operations and therefore include repeated steps and precursors.
The E/Z fixture’s identity difference originates in parsing semantics. Upstream sanitize=False construction followed by sanitization omits E/Z assignment for this InChI; full browser parsing retains stereochemistry and produces a different complete InChIKey. Both keys and their parsing paths remain in the records. The browser key supports stock and cycle checks.
These checks examine policy ranking, template-generated precursors and terminal-stock membership separately. Each of the six inputs returns five candidates; random processes also affect the order of search visits and route rankings. Evaluating route quality requires an independent target set and experimental or expert judgement as a reference.
| Layer | Coverage | Result and boundary |
|---|---|---|
| Fingerprints and indices | 6 inputs, top 50 each | Exact |
| Policy probabilities | 300 | Maximum absolute difference 3.8743×10⁻⁷; tolerance 10⁻⁶ |
| Templates and returned steps | 107 | Native RDChiral passes; includes repeated checks |
| Terminal stock claims | 52 | Native SQLite passes; includes repeated precursors |
| Stereo identity | One E/Z alkene | Fully parsed stereo difference retained |
| Search ranking | Different random processes | Evaluated under each random process |
Aspirin candidate route and chemical assessment
The first-ranked aspirin route proposes acetic anhydride and salicylic acid through template 40152, with policy probability approximately 0.726163. Both precursors match historical stock and the state score is 0.997629. The figure is drawn directly from returned isomeric SMILES; the double-line arrow denotes retrosynthetic disconnection.
Both precursors are in stock, so the search has found a candidate route that satisfies its stopping condition. Taking it to the laboratory requires literature evidence and chemical judgement to establish stoichiometry, solvent, temperature, workup and expected yield, as well as functional-group compatibility and scale-up conditions.
As with forward ASKCOS, native comparisons check whether model and chemistry operations are reproduced. Comparing route quality further requires a fixed target set, stock, time budget and route-equivalence rules, with experimental or expert judgement to assess stereochemistry and reaction conditions.
Search times and initial loading
Six asset-ready searches in macOS Chrome 154 take approximately 0.66–1.56 s, each a single 80-iteration, four-step, single-thread FP32 record. Model calls range from 8 to 65 and template calls from 66 to 609, showing workload variation under the same iteration budget.
Aspirin page records take 5.217 s with an initially empty cache, 1.822 s on repeat and 2.852 s offline after cancellation, each with n = 1. The first starts at download approval and the others at submission; all end after automated detection and export, including polling and export. Internal searches take 2.156, 0.683 and 1.630 s. Figures and tables identify the two boundaries separately.
The offline test uses cached assets, with zero asset requests and zero page calculation API POSTs. Retry after cancellation, export and reload recovery also complete successfully. These timings come from single runs on desktop localhost. Public-network download speeds and phone memory and thermal conditions will affect first use and sustained search.
| State | Page through export / s | Internal search / s |
|---|---|---|
| cold | 5.217 | 2.1560 |
| warm | 1.822 | 0.6832 |
| offline-cached-after-cancel | 2.852 | 1.6303 |
Data downloads and native checks
The bundle contains six native references, browser chemistry and search outputs, page exports, a 300-row probability CSV, search and page-timing CSVs, build identities and a SHA-256 manifest. python replay.py uses the standard library to check hashes, fingerprints, indices, probability tolerances and source counts.[3]
Native reaction and stock checks use tools/compute-build/aizynth/verify_science.py, requiring RDKit, RDChiral, the original templates and pinned ZINC SQLite; reference.py provides the model-reference entry point. Native replay results accompany the data.
Rebuilding uses the pinned ONNX model, RDChiral C++ commit 014c11d243f5f5bb60ea1e52f3350befc92b20a6 and recorded Emscripten dependencies. Downloads include build, stock-packing and verification sources and third-party notices. Asset manifests identify the model, template and stock sources.
AiZynthFinder 的浏览器移植、库存分片与原生验证
在浏览器运行原始模板策略、化学操作与树搜索。库存分片按需查询 1742 万项完整键,六个目标分子的结果与原生实现逐层对照。
2026-10-10 · Mol2Mat · revision 1
从目标分子到库存前体
逆合成从目标分子向前体拆解,再继续拆解其中的中间体,直到路线末端都能在指定库存中找到。AiZynthFinder 用反应模板提出拆解方式,用神经策略排列模板优先级,再通过树搜索分配有限的计算预算。[1,2]
Mol2Mat 将 AiZynthFinder 4.4.1 的 USPTO 配置移植到一个 Dedicated Worker:原始 FP32 ONNX 策略负责排序,42554 个模板与 RDKit / RDChiral 负责生成前体,单目标 Monte Carlo tree search 负责组织路线。这些步骤都在浏览器本地运行。
移植的另一项工作是让浏览器访问历史 ZINC 库存中的 17422831 个完整 InChIKey。Bloom 索引先筛选,再按需读取分片并精确查找,减少首次加载的数据量。六个分子的原生对照检查了指纹、300 项策略概率、107 次反应操作和 52 次库存声明。[3]
分子指纹与原始模板策略
RDKit 解析并消毒目标,生成规范异构 SMILES,从分子对象移除原子映射。策略采用半径 2、2048 位的非手性 Morgan 指纹:
概率用于模板排序,分子的立体化学则由后续完整分子图处理。
ONNX Runtime Web 1.30.0 使用单线程 WASM 执行提供器和原始 FP32 权重。按概率最多保留 50 项,沿用累计阈值 0.995 的截断语义与 FP32 累加。六个输入各保留 50 个模板,300 项概率及索引可逐项检查。[2,3]
RDKit 2026.03.6 与固定 RDChiral C++ 提交编译为 WASM,桥接层管理输入、输出和资源生命周期。选中模板作用于完整分子图,经过结构消毒、立体化学处理与去重,产生零个或多个前体组合。策略排序和实际可展开结果因此分别记录。[4]
单目标树搜索与预算约束
JavaScript 搜索保留节点选择、延迟展开、回传和循环剪枝。初始化子节点的选择量为:
V_i 为累计值,n_i 为访问数,N 为兄弟节点总访问数;策略先验参与初始化。探索项在高平均值和较少访问的分支之间分配计算预算。
状态评分沿用上游定义:
f_stock 表示库存覆盖比例,深度项轻度偏好较浅路线。该分数评价搜索状态,路线的实验可行性、选择性和分离成本需要化学证据。
测量配置为 80 迭代、最多四步,任务另有 45 s 搜索预算,界面允许最多五步。浏览器使用 xorshift32 种子 20261010,Python 参考使用 NumPy 全局随机过程。因此数值与化学层逐项对照,随机访问轨迹与路线排名按各自随机规则解释。
分片库存与精确成员判断
历史 ZINC 快照包含 17422831 个完整 InChIKey,源 SQLite 约 598 MiB。浏览器采用 16 MiB Bloom 索引和 64 个排序压缩分片。五个哈希先筛去确定缺席的键,Bloom 阳性再对相应分片的完整 27 字节键作二分查找,精确查找决定成员身份。[3,5]
压缩分片共 129073867 字节,约 123.09 MiB,按需获取并校验 SHA-256 后持久缓存。已解压分片使用上限 32 MiB 的 LRU。含模型和运行时的完整批准资源约 250.87 MiB,提前加载的核心约 127.77 MiB;单次查询传输由缓存及所需分片决定。.bin 保持下载字节与哈希一致。
离线查询要求所需分片完整缓存;缺少分片返回资源错误,库存语义保持完整。此库存表示固定历史集合,当前供应商、价格、交期及纯度需要采购来源更新。
重复搜索中的状态复用
解析、策略推理、模板执行和树搜索在一个 Dedicated Worker 中串行运行,页面通过任务协议接收进度及完整结果。取消结束当前任务,已完成记录可导出并在刷新后恢复,任务、模型哈希和输入身份随记录保存。
重复搜索时,Worker 复用 ONNX 会话、已解析模板、分子身份与库存命中,减少初始化工作。模板和身份缓存限制条目数;神经预测与反应结果 Map 在每次搜索开始时重建。解压库存由 LRU 管理,压缩资产则保存在持久缓存中。
探针记录 ONNX WASM 线性内存 367394816 字节,化学 WASM 33554432 字节。任务的 1 GiB 为准入估算;JavaScript 对象、持久存储、图形与进程 RSS 需另行记账。性能证据的分子范围由六个固定目标定义。
策略、化学与库存的原生对照
六个目标为阿司匹林、乙酰苯胺、苯甲酸乙酯、手性布洛芬、E/Z 烯烃和咖啡因。原生与浏览器的指纹位、前 50 模板索引逐项相同;300 项保存概率的最大绝对差为 3.8743×10⁻⁷,低于 10⁻⁶ 容差。[3]
原生 RDChiral 复核模板展开与返回路线中的步骤,共完成 107 次反应检查;RDKit 与只读 SQLite 检查末端分子的库存归属,共 52 次。两项检查全部通过。次数按每次操作计,因而包含重复出现的步骤和前体。
E/Z 夹具的身份差异来自解析语义。上游采用 sanitize=False 后再消毒,生成 InChI 时遗漏该例 E/Z 赋值;浏览器完整解析保留立体化学,得到不同完整 InChIKey。两种键及来源路径均在记录中,浏览器键用于库存与循环判断。
这些检查分别检验策略排序、模板生成的前体和末端库存归属。六个输入各返回五条候选;随机搜索的访问顺序和排名还会受到随机过程影响。评价路线质量时,需要独立目标集及实验或专家判断作为参照。
| 层次 | 核查范围 | 结果与边界 |
|---|---|---|
| 指纹与模板索引 | 6 个输入,各前 50 项 | 精确一致 |
| 策略概率 | 300 | 最大绝对差 3.8743×10⁻⁷;容差 10⁻⁶ |
| 模板及返回步骤 | 107 | 原生 RDChiral 通过;含重复检查 |
| 末端库存声明 | 52 | 原生 SQLite 通过;含重复前体 |
| 立体身份 | E/Z 烯烃一例 | 保留完整解析差异 |
| 搜索排名 | 两种随机过程 | 按各自随机规则评价 |
阿司匹林候选路线与化学评价
阿司匹林第一位路线提出乙酸酐和水杨酸,模板索引为 40152,策略概率约 0.726163。两种前体命中历史库存,状态分数为 0.997629。图直接由返回路线的异构 SMILES 绘制,双线箭头表示逆合成拆解。
两个前体都在库存中,搜索因此得到一条满足终止条件的候选路线。将它用于实验时,还需从文献和化学判断中确定试剂用量、溶剂、温度、后处理及预期收率,并考虑官能团兼容性和放大条件。
与正向 ASKCOS 一样,原生对照检查模型和化学操作是否复现。若进一步比较路线质量,需要固定目标集、库存、时间预算和路线等价规则,并依据实验或专家判断评价立体化学与反应条件。
搜索耗时与首次加载
macOS Chrome 154 的六个资源就绪搜索约为 0.66–1.56 s,各为 80 迭代、最多四步、单线程 FP32 的单次记录。模型调用为 8–65 次,模板调用为 66–609 次,显示相同迭代预算下的工作量差异。
实际阿司匹林页面的首次空缓存、重复和取消后离线记录分别为 5.217、1.822、2.852 s,各 n = 1。首次从下载同意开始,其余从提交开始,均到自动化检测与导出结束,包含轮询及导出。对应内部搜索为 2.156、0.683、1.630 s,图表分别标明两类边界。
离线测试使用已缓存资产,资产请求和页面计算 API POST 均为零;取消后重试、导出和刷新恢复也正常完成。以上耗时来自桌面 localhost 的单次运行。公共网络的下载速度,以及手机的内存和热状态,会影响首次使用和持续搜索的体验。
| 状态 | 页面至导出 / s | 内部搜索 / s |
|---|---|---|
| cold | 5.217 | 2.1560 |
| warm | 1.822 | 0.6832 |
| offline-cached-after-cancel | 2.852 | 1.6303 |
数据下载与原生复核
数据包包含六套原生参考、浏览器化学与搜索输出、页面导出、300 项概率 CSV、搜索及页面计时 CSV,以及构建身份与 SHA-256 清单。python replay.py 使用标准库检查哈希、指纹、索引、概率容差和来源计数。[3]
原生反应与库存核查入口为 tools/compute-build/aizynth/verify_science.py,依赖 RDKit、RDChiral、原始模板库和固定 ZINC SQLite;reference.py 提供模型参考入口。原生回放结果随数据保存。
重建使用固定 ONNX 模型、RDChiral C++ 提交 014c11d243f5f5bb60ea1e52f3350befc92b20a6 及记录的 Emscripten 依赖。下载附构建、库存分片、验证源码和第三方声明,资产清单提供模型、模板和库存的获取身份。
Data and sources
- Complete saved results · JSON gzip
- Six-case numerical and search summary · CSV
- 300 native probability comparisons · CSV
- Six search observations · CSV
- Page timing observations · CSV
- Provenance and build identities · JSON
- Data, replay script and checksum manifest · tar.gz
- Asset volumes · CSV
- Genheden et al. AiZynthFinder: retrosynthetic planning (2020)
- AiZynthFinder 4.4.1: implementation and configuration documentation
- Mol2Mat: original results, source identities and native replay
- RDChiral: stereochemistry-aware reaction-template application
- Pinned historical ZINC stock snapshot: source file