Pith. sign in

REVIEW 3 major objections 6 minor 67 references

Most stock-terminated synthesis routes fail chemical-plausibility checks; specialized planners still beat LLMs on hard targets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 15:06 UTC pith:TZI3PLJD

load-bearing objection Solid, usable Solv-2 benchmark that shows stock termination overstates chemical validity; specialized CASP still beats LLMs on hard targets and is far cheaper. the 3 major comments →

arxiv 2607.04688 v1 pith:TZI3PLJD submitted 2026-07-06 cs.LG cs.AIcs.CEcs.CL

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

classification cs.LG cs.AIcs.CEcs.CL
keywords retrosynthesissynthesis planningchemical plausibilitybenchmarkSolv-NChemCensorCASPlarge language models
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Computer-aided synthesis planning has long been scored mainly by whether a planner can reach commercially available starting materials. That formal success often hides chemically invalid steps. This paper introduces URSA, a route-level evaluation framework that standardizes heterogeneous planner outputs and then scores every reaction for chemical plausibility against synthetic precedents, aiming to capture the Solv-2 layer of validity that chemists actually care about. On a hard set of novel molecules without published routes, the best specialized systems solve only about one-third of targets under the strict plausibility criterion; the best large language model trails them. On a more familiar set of drugs and clinical candidates the same specialized tools remain competitive or superior, and at far lower cost. The central claim is that the field is moving from mere navigability to chemical validity, and that stock-termination metrics alone substantially overstate practical synthesis capability.

Core claim

Stock-terminated synthetic routes frequently fail stricter Solv-1 and Solv-2 chemical-plausibility checks. On the hard URSA-expert-2026 set the strongest conventional planner reaches only 32 percent Solv-2 while the best LLM reaches 21 percent; on drugs and clinical candidates specialized CASP systems remain competitive or superior to frontier LLMs while being substantially cheaper. Multistep planning is not reducible to isolated reaction-plausibility judgment.

What carries the argument

URSA: a pipeline that standardizes routes (via RetroCast or LLM parsing), enforces stock termination and structural integrity, generates alternative subtrees to match reaction granularity, scores every step with ChemCensor precedent matching, and selects a best route under a Solv-2-first priority.

Load-bearing premise

The central ranking rests on treating ChemCensor's precedent match of reaction centers and functional-group signatures as a faithful automated stand-in for expert chemical-plausibility judgment.

What would settle it

A large independent panel of expert chemists re-labels the same best routes that URSA scored; if many ChemCensor-passing routes are rejected as implausible (or many failing routes accepted), the Solv-2 ranking of planners collapses.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces URSA, a route-level evaluation framework that scores retrosynthetic plans under the Solv-N hierarchy (stock termination plus syntactic, topological, and chemical-plausibility validity). Routes from heterogeneous CASP systems and LLMs are standardized (via RetroCast or LLM parsers), stock-checked against a shared commercial building-block set, optionally augmented by subtree generation, and scored with ChemCensor as an automated Solv-2 judge. The authors release expert-labeled reaction data (1000 reactions), two new target sets (URSA-expert-2026 and URSA-drugs&clinicals-2026), code, and annotated best routes. Empirically, ChemCensor outperforms LLM judges on the reaction set (accuracy 0.96 vs ≤0.83); on route planning, stock-terminated Solv-0 rates substantially exceed Solv-2 rates, conventional planners (notably RetroChimera-MCTS, AZF-MCTS) remain competitive or superior to frontier LLMs on both target sets, and specialized systems are far cheaper.

Significance. If the results hold, URSA supplies a practical, chemistry-aware alternative to stock-termination and exact route-reproduction metrics at a time when both specialized CASP tools and LLMs are being compared without shared validity criteria. The released open target sets, building-block stock, ChemCensor integration, annotated CDXML routes, and cost accounting are concrete community assets. The head-to-head finding that stock termination overstates chemical validity, and that conventional planners still match or beat LLMs under Solv-2 while remaining cheaper, is actionable for both method developers and medicinal-chemistry users. Strengths include the large multi-system evaluation, bootstrap CIs on Solv-2, dual open/closed precedent corpora (U2 vs U2P2), and explicit limitations on Solv-3 and best-route-only reporting.

major comments (3)
  1. §5.2 and Table 1: ChemCensor is adopted as the automated Solv-2 ground truth for all route rankings on the basis of a 1000-reaction expert-labeled set (accuracy 0.96, MCC 0.92). The manuscript does not report inter-annotator agreement, number of annotators per reaction, adjudication protocol, or how borderline selectivity cases were resolved. Because binary Solv-2 pass/fail is load-bearing for the central ranking claim (stock termination ≫ chemical validity; CASP vs LLM ordering), please add agreement statistics (e.g., Cohen’s κ or majority-vote rates) and a short description of the labeling SOP; if only single-annotator labels exist, state that limitation explicitly and discuss its effect on the reported accuracy gap versus LLMs.
  2. §4.2 best-route selection and Eq. (1): Primary metrics (Solv-0/1/2, CC*) are computed only on the single best route per target after allowing up to 10 candidates and subtree augmentation. Systems that emit more diverse or more finely granular routes can be systematically favored by the priority order (Solv-2 first, then average CC, then length). Please report, at least in the appendix, complementary statistics that are less selection-dependent—e.g., fraction of targets with any Solv-2 route among all candidates, mean Solv-2 rate over all valid stock-terminated routes, or success@k—so that the ranking is not driven solely by best-of-N selection.
  3. §3–§4 and Related Work: URSA’s Solv-2 component, route standardization (RetroCast), and Solv-N formalism are drawn from closely related prior work by overlapping author groups. The new expert reaction set and open target sets mitigate circularity for the ranking claim, but the manuscript should more clearly separate (i) what is newly validated here (ChemCensor vs LLM judges on the 1000-reaction set; planner rankings under a fixed protocol) from (ii) what is inherited infrastructure. A short paragraph stating independence assumptions and any held-out checks against ChemCensor’s own training/reference construction would strengthen confidence that the Solv-2 judge is not inadvertently tuned to the same systems being ranked.
minor comments (6)
  1. Figure 3 and Table 2: Solv-0/1/2 bars are clear, but adding absolute counts (n targets solved) next to percentages would help readers interpret small differences on URSA-expert-2026 (e.g., 32% vs 29%).
  2. §5.1 model list: Several model version names (GPT 5.x, Claude Opus 4.x, Grok 4.x) will age quickly; pin exact API model IDs and access dates in Appendix F for reproducibility.
  3. Appendix J (USPTO-190 critique) is useful but long; a one-paragraph summary in the main text with a pointer to the appendix would improve flow for readers who do not need the full cluster analysis.
  4. Eq. (1): Define the range and missing-value handling of CCi more explicitly when a reaction receives score 0 (implausible) versus when ChemCensor cannot extract a center.
  5. Typos / polish: “Utilitarian RetroSynthesis Assessment” capitalization is inconsistent in places; “URSA-Drugs&Clinicals-2026” vs “URSA-drugs&clinicals-2026” should be normalized; a few long sentences in §6.2 could be split for readability.
  6. Appendix M (Synthegy critique) is informative but somewhat polemical in tone relative to the rest of the paper; consider shortening to 1–2 concrete failure modes with figures retained.

Circularity Check

1 steps flagged

Mild self-citation of evaluation infrastructure (ChemCensor, Solv-N, RetroCast) from overlapping authors; the empirical planner ranking is not forced by construction.

specific steps
  1. self citation load bearing [§3 Preliminaries (Solv-N, ChemCensor); §4.2 Evaluation protocol; Table 1]
    "We introduce URSA ... at the practically urgent Solv-1/Solv-2 boundary. ... In URSA, ChemCensor is used as a candidate automated validator for applying Solv-2 plausibility checks across every step of a proposed route. ... ChemCensor achieves the highest accuracy on the expert-labeled reaction plausibility benchmark (0.96) ... support using deterministic ChemCensor as the automated Solv-2 validator in URSA"

    The Solv-N hierarchy [14], RetroCast standardization [12], and ChemCensor scorer [21] are all prior work with overlapping authors (Morgunov; Zagribelnyy et al.). The paper’s operational definition of route-level Solv-2 therefore rests on a self-cited stack. This is mild infrastructure self-reference, not a forced result: Table 1 supplies independent expert labels against which ChemCensor is chosen over LLM judges, and the planner rankings remain empirical measurements on new targets rather than identities of the cited tools.

full rationale

URSA’s measurement stack reuses Solv-N formalism, RetroCast route standardization, and ChemCensor Solv-2 scoring from papers with substantial author overlap (Morgunov on Solv-N/RetroCast/DMS; Zagribelnyy–Ilin–Kuznetsov–Bondarev–Shayakhmetov–Aladinskiy–Aliper–Zhavoronkov on ChemCensor). That is real self-reference of the evaluation apparatus. It is not, however, a reduction of the central claim to its inputs by construction: ChemCensor is independently validated on a newly collected 1000-reaction expert-labeled set (0.96 accuracy / 0.92 MCC vs LLM judges 0.73–0.83), the target sets are released, and the ranking of planners (stock termination ≫ Solv-2; specialized CASP competitive/superior and cheaper than frontier LLMs) is an empirical outcome that could have gone either way. There is no fitted parameter reappearing as a prediction, no uniqueness theorem forbidding alternatives, and no self-definitional identity between inputs and reported Solv-2 rates. Score 2 reflects non-load-bearing self-citation of infrastructure, not circular derivation of the main result.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 2 invented entities

The central ranking claims rest on (1) the Solv-N validity hierarchy taken as given, (2) ChemCensor as an automated Solv-2 oracle whose binary threshold and precedent database are fixed by the authors, (3) a hand-curated commercial building-block stock, and (4) expert binary labels on 1000 reactions that themselves depend on chemist judgment. No free continuous parameters are fitted to the planner rankings themselves; the free choices are discrete protocol decisions.

free parameters (3)
  • ChemCensor binarization threshold
    Scores on [0,1-5] are mapped 0→implausible, 1-5→plausible; this cut-off is chosen by the authors and directly defines Solv-2 success.
  • Shared CABB stock size/composition
    255 365 compounds assembled from Enamine + ASKCOS buyables + eMolecules Tier 1/2; stock definition controls the [task] termination constraint and therefore all Solv-N rates.
  • Max routes retained per target
    Up to 10 routes per model; truncation rule (score or random) affects which candidate is available for best-route selection.
axioms (3)
  • domain assumption Solv-N hierarchy (Solv-0 syntactic, Solv-1 topological/reaction-center, Solv-2 selectivity/plausibility, Solv-3 executability) correctly orders chemical validity for synthesis planning.
    Adopted from prior work by overlapping authors and used as the evaluation backbone throughout Sections 3–6.
  • domain assumption Matching reaction-center + functional-group signatures against USPTO/Pistachio precedents is a valid operationalization of chemical plausibility (Solv-2).
    ChemCensor is treated as the automated Solv-2 judge after a 1000-reaction expert comparison (Table 1).
  • domain assumption Expert binary labels on the 1000-reaction set are ground truth for plausibility discrimination.
    Used to select ChemCensor over LLM judges; labels are internal and not independently re-annotated.
invented entities (2)
  • URSA evaluation protocol (Solv-0/1/2 + CC* best-route selection + subtree augmentation) independent evidence
    purpose: Unified, chemistry-aware scoring of heterogeneous retrosynthesis outputs.
    New composite framework; independent evidence is the released code and the expert reaction set, but the protocol itself is defined by the paper.
  • URSA-expert-2026 and URSA-drugs&clinicals-2026 target sets independent evidence
    purpose: Provide prospective-style and medicinally relevant molecules without disclosed routes.
    New public benchmarks intended to replace USPTO-190; structures and labels released.

pith-pipeline@v1.1.0-grok45 · 31592 in / 2815 out tokens · 29318 ms · 2026-07-11T15:06:39.036337+00:00 · methodology

0 comments
read the original abstract

Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Recent progress has produced both specialized deep-learning retrosynthesis systems and general-purpose large language models, but objective comparison remains difficult due to the lack of flexible, chemically interpretable benchmarking protocols. In the current study, we are introducing the URSA (Utilitarian RetroSynthesis Assessment) evaluation framework that provides the opportunity to benchmark the synthetic routes not only from a formal perspective, such as convergence to commercially available starting materials, but also from a chemical plausibility perspective, mimicking the way expert chemists evaluate the reactions and routes. The study covers a comprehensive evaluation of both conventional end-to-end retrosynthesis solutions and LLMs for the synthesis planning task on a set of novel, diverse target molecules with undisclosed synthetic routes, which represent realistic tasks in the daily drug design routine. We find that while LLMs can support high-level strategic planning, they currently underperform specialized retrosynthesis models in reliably solving synthesis planning tasks.

Figures

Figures reproduced from arXiv: 2607.04688 by Alex Aliper, Alex Zhavoronkov, Anton Morgunov, Arkadii Lin, Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Rim Shayakhmetov, Vladimir Aladinskiy.

Figure 1
Figure 1. Figure 1: URSA benchmark system. A. The benchmarking pipeline for both conventional synthesis planning tools and LLMs. B. The URSA benchmark components. 4.2 Evaluation protocol formalization URSA evaluates the necessary precondition for executable synthesis planning under Solv-N[task] formalism: whether a proposed route is stock-terminated [task], structurally (Solv-0) and topologically (Solv-1) valid, and chemicall… view at source ↗
Figure 2
Figure 2. Figure 2: Examples of multi-step reaction entities stored in the USPTO reaction database. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Assessment of routes predicted by conventional retrosynthetic tools ( [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Clusters of similar molecules from USPTO-190. Maximum common substructure (MCS) is highlighted independently for each cluster. 27 [PITH_FULL_IMAGE:figures/full_fig_p027_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The biggest cluster of highly similar structures from [PITH_FULL_IMAGE:figures/full_fig_p028_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Example of the subtree generation process. [PITH_FULL_IMAGE:figures/full_fig_p029_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: ChemCensor functionalities. A. The input reaction. B. Process of reaction center extraction with different RCN highlighted in blue. C. The respective FG signatures annotated for the extracted centers. ChemCensor estimates how closely a single reaction resembles a reference corpus of synthetic precedents and uses that resemblance as a proxy for chemical plausibility. Besides precedent estimation, it also pe… view at source ↗
Figure 8
Figure 8. Figure 8: Reaction erroneously understood as an SNAr reaction by the Synthegy framework. [PITH_FULL_IMAGE:figures/full_fig_p032_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Mechanism supporting the correct interpretation of the highlighted reaction. [PITH_FULL_IMAGE:figures/full_fig_p032_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Reaction unfairly penalized by Synthegy. [PITH_FULL_IMAGE:figures/full_fig_p033_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: An example from literature supporting the reaction penalized by Synthegy. [PITH_FULL_IMAGE:figures/full_fig_p033_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 6 canonical work pages

  1. [1]

    Grand challenges for predictive modeling in small molecule drug discovery.ChemRxiv, 2026(0304), 2026

    Connor W Coley, Pankaj Daga, Marco De Vivo, Willem Jespers, Ashutosh S Jogalekar, S Roy Kimura, Lucien Koenekoop, Anne-Grete Märtson, Timothy R Newhouse, Soumya Ray, Riccardo Sabatini, David C Thompson, and Woody Sherman. Grand challenges for predictive modeling in small molecule drug discovery.ChemRxiv, 2026(0304), 2026. doi: 10.26434/chemrxiv.15000615/v...

  2. [2]

    E. J. Corey and W. Todd Wipke. Computer-Assisted Design of Complex Organic Syntheses: Pathways for molecular synthesis can be devised with a computer and equipment for graphical communication.Science, 166(3902):178–192, 1969. ISSN 0036-8075, 1095-9203. doi: 10.1126/science.166.3902.178. URL https://www.science.org/doi/10.1126/science. 166.3902.178

  3. [3]

    Planning chemical syntheses with deep neural networks and symbolic AI.Nature, 555(7698):604–610, March 2018

    Marwin H S Segler, Mike Preuss, and Mark P Waller. Planning chemical syntheses with deep neural networks and symbolic AI.Nature, 555(7698):604–610, March 2018. doi: 10.1038/ nature25978. URLhttps://www.nature.com/articles/nature25978

  4. [4]

    Choure, Mun Hong Fong, Jihye Roh, Itai Levin, Kevin Yu, Joonyoung F

    Zhengkai Tu, Sourabh J. Choure, Mun Hong Fong, Jihye Roh, Itai Levin, Kevin Yu, Joonyoung F. Joung, Nathan Morgan, Shih-Cheng Li, Xiaoqi Sun, Huiqian Lin, Mark Murnin, Jordan P. Liles, Thomas J. Struble, Michael E. Fortunato, Mengjie Liu, William H. Green, Klavs F. Jensen, and Connor W. Coley. Askcos: Open-source, data-driven synthesis planning.Accounts o...

  5. [5]

    Retro*: Learning retrosynthetic planning with neural guided a* search, 2020

    Binghong Chen, Chengtao Li, Hanjun Dai, and Le Song. Retro*: Learning retrosynthetic planning with neural guided a* search, 2020. URLhttps://arxiv.org/abs/2006.15820

  6. [6]

    Self-improved retrosynthetic planning, 2021

    Junsu Kim, Sungsoo Ahn, Hankook Lee, and Jinwoo Shin. Self-improved retrosynthetic planning, 2021. URLhttps://arxiv.org/abs/2106.04880

  7. [7]

    A data-driven group retrosynthesis planning model inspired by neurosymbolic programming.Nature Communi- cations, 16(1):192, Jan 2025

    Xuefeng Zhang, Haowei Lin, Muhan Zhang, Yuan Zhou, and Jianzhu Ma. A data-driven group retrosynthesis planning model inspired by neurosymbolic programming.Nature Communi- cations, 16(1):192, Jan 2025. ISSN 2041-1723. doi: 10.1038/s41467-024-55374-9. URL https://doi.org/10.1038/s41467-024-55374-9

  8. [8]

    Retrograph: Retrosynthetic planning with graph search

    Shufang Xie, Rui Yan, Peng Han, Yingce Xia, Lijun Wu, Chenjuan Guo, Bin Yang, and Tao Qin. Retrograph: Retrosynthetic planning with graph search. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’22, page 2120–2129, New York, NY , USA, 2022. Association for Computing Machinery. ISBN 9781450393850. doi: 10.1145/35...

  9. [9]

    Grasp: Navigating retrosynthetic planning with goal-driven policy

    Yemin Yu, Ying Wei, Kun Kuang, Zhengxing Huang, Huaxiu Yao, and Fei Wu. Grasp: Navigating retrosynthetic planning with goal-driven policy. In S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors,Advances in Neu- ral Information Processing Systems, volume 35, pages 10257–10268. Curran Associates, Inc., 2022. URL https://proceedings....

  10. [10]

    Yu Shee, Anton Morgunov, Haote Li, and Victor S. Batista. Directmultistep: Direct route generation for multistep retrosynthesis.Journal of Chemical Information and Modeling, 65(8): 3903–3914, 2025. doi: 10.1021/acs.jcim.4c01982. URL https://doi.org/10.1021/acs. jcim.4c01982

  11. [11]

    Models matter: the impact of single-step retrosynthesis on synthesis planning.Digit

    Paula Torren-Peraire, Alan Kai Hassen, Samuel Genheden, Jonas Verhoeven, Djork-Arné Clevert, Mike Preuss, and Igor V Tetko. Models matter: the impact of single-step retrosynthesis on synthesis planning.Digit. Discov., 3(3):558–572, 2024. doi: 10.1039/D3DD00252G. URL https://pubs.rsc.org/en/content/articlelanding/2024/dd/d3dd00252g

  12. [12]

    Anton Morgunov and Victor S. Batista. Procrustean bed for ai-driven retrosynthesis: A unified framework for reproducible evaluation, 2025. URL https://arxiv.org/abs/2512.07079. 10

  13. [13]

    PaRoutes: towards a framework for benchmarking retrosynthesis route predictions.Digital Discovery, 1(4):527–539, 2022

    Samuel Genheden and Esben Bjerrum. PaRoutes: towards a framework for benchmarking retrosynthesis route predictions.Digital Discovery, 1(4):527–539, 2022. ISSN 2635-098X. doi: 10.1039/D2DD00015F. URLhttp://xlink.rsc.org/?DOI=D2DD00015F

  14. [14]

    The syntax of matter: Synthesis planning as the foundation of generative chemistry.ChemRxiv, 2026(0326), 2026

    Anton Morgunov, Yu Shee, Alexander V Soudackov, and Victor S Batista. The syntax of matter: Synthesis planning as the foundation of generative chemistry.ChemRxiv, 2026(0326), 2026. doi: 10.26434/chemrxiv.15001278/v1. URL https://chemrxiv.org/doi/abs/10.26434/ chemrxiv.15001278/v1

  15. [15]

    Krzysztof Maziarz, Austin Tripp, Guoqing Liu, Megan Stanley, Shufang Xie, Piotr Gai ´nski, Philipp Seidl, and Marwin H. S. Segler. Re-evaluating retrosynthesis algorithms with syntheseus. Faraday Discuss., 256:568–586, 2025. doi: 10.1039/D4FD00093E. URL http://dx.doi. org/10.1039/D4FD00093E

  16. [16]

    SynthArena: A unified evaluation framework for AI-driven retrosynthesis, 2026

    ischemist. SynthArena: A unified evaluation framework for AI-driven retrosynthesis, 2026. syntharena.ischemist.com

  17. [17]

    Trustworthy retrosynthesis: Eliminating halluci- nations with a diverse ensemble of reaction scorers, 2025

    Michal Sadowski, Tadija Radusinovi ´c, Maria Wyrzykowska, Lukasz Sztukiewicz, Jan Rzymkowski, Paweł Włodarczyk-Pruszy´nski, Mikołaj Sacha, Piotr Kozakowski, Ruard van Workum, and Stanislaw Kamil Jastrzebski. Trustworthy retrosynthesis: Eliminating halluci- nations with a diverse ensemble of reaction scorers, 2025. URL https://arxiv.org/abs/ 2510.10645

  18. [18]

    Synthelite: Chemist-aligned and feasibility-aware synthesis planning with llms, 2025

    Nguyen Xuan-Vu, Daniel Armstrong, Milena Wehrbach, Andres M Bran, Zlatko Jonˇcev, and Philippe Schwaller. Synthelite: Chemist-aligned and feasibility-aware synthesis planning with llms, 2025. URLhttps://arxiv.org/abs/2512.16424

  19. [19]

    Chemical reasoning in LLMs unlocks strategy-aware synthesis planning and reaction mechanism elucidation.Matter, 0(102812):102812, April 2026

    Andres M Bran, Théo A Neukomm, Daniel Armstrong, Zlatko Jonˇcev, and Philippe Schwaller. Chemical reasoning in LLMs unlocks strategy-aware synthesis planning and reaction mechanism elucidation.Matter, 0(102812):102812, April 2026. doi: 10.1016/j.matt.2026.102812. URL https://www.cell.com/matter/fulltext/S2590-2385(26)00175-X

  20. [20]

    Synthstrategy: Extracting and formalizing latent strategic insights from llms in organic chemistry, 2025

    Daniel Armstrong, Zlatko Jon ˇcev, Andres M Bran, and Philippe Schwaller. Synthstrategy: Extracting and formalizing latent strategic insights from llms in organic chemistry, 2025. URL https://arxiv.org/abs/2512.01507

  21. [21]

    When single answer is not enough: Rethinking single-step retrosynthesis benchmarks for llms, 2026

    Bogdan Zagribelnyy, Ivan Ilin, Maksim Kuznetsov, Nikita Bondarev, Roman Schutski, Thomas MacDougall, Rim Shayakhmetov, Zulfat Miftakhutdinov, Mikolaj Mizera, Vladimir Aladinskiy, Alex Aliper, and Alex Zhavoronkov. When single answer is not enough: Rethinking single-step retrosynthesis benchmarks for llms, 2026. URLhttps://arxiv.org/abs/2602.03554

  22. [22]

    Chemical reactions from US patents (1976-Sep2016)

    Daniel Lowe. Chemical reactions from US patents (1976-Sep2016). 6 2017. doi: 10.6084/ m9.figshare.5104873.v1. URL https://figshare.com/articles/dataset/Chemical_ reactions_from_US_patents_1976-Sep2016_/5104873

  23. [23]

    SciFinder: Retrosynthesis software, 2026

    Chemical Abstracts Service (CAS). SciFinder: Retrosynthesis software, 2026. cas.org/. . . /synthesis-planning

  24. [24]

    Reaxys, 2026

    Elsevier. Reaxys, 2026. elsevier.com/solutions/reaxys

  25. [25]

    Scalfani, Daniel Probst, Kazuya Ujihara, guillaume godin, Rachel Walker, Juuso Lehtivarjo, Axel Pahl, Francois Berenger, jasondbiggs, and strets123

    Greg Landrum, Paolo Tosco, Brian Kelley, Ric, David Cosgrove, sriniker, gedeck, Riccardo Vianello, NadineSchneider, Eisuke Kawashima, Gareth Jones, Dan N, Andrew Dalke, Brian Cole, Matt Swain, Samo Turk, AlexanderSavelyev, Alain Vaucher, Maciej Wójcikowski, Ichiru Take, Vincent F. Scalfani, Daniel Probst, Kazuya Ujihara, guillaume godin, Rachel Walker, Ju...

  26. [26]

    Grok 4.1 model card, November 2025

    xAI. Grok 4.1 model card, November 2025. URL https://data.x.ai/ 2025-11-17-grok-4-1-model-card.pdf

  27. [27]

    Grok 4.3 model card, November 2026

    xAI. Grok 4.3 model card, November 2026. URL https://docs.x.ai/developers/ models/grok-4.3. 11

  28. [28]

    Gemini 3.1 pro, February 2026

    Gemini Team. Gemini 3.1 pro, February 2026. URL https://deepmind.google/models/ model-cards/gemini-3-1-pro/

  29. [29]

    Gpt-5.1 instant and gpt-5.1 thinking system card addendum, November 2025

    OpenAI. Gpt-5.1 instant and gpt-5.1 thinking system card addendum, November 2025. URL https://cdn.openai.com/pdf/4173ec8d-1229-47db-96de-06d87147e07e/5_ 1_system_card.pdf

  30. [30]

    Update to gpt-5 system card: Gpt-5.2, December 2025

    OpenAI. Update to gpt-5 system card: Gpt-5.2, December 2025. URL https://cdn.openai. com/pdf/3a4153c8-c748-4b71-8e31-aecbde944f8d/oai_5_2_system-card.pdf

  31. [31]

    Introducing gpt -5.4, March 2026

    OpenAI. Introducing gpt -5.4, March 2026. URL https://openai.com/index/ introducing-gpt-5-4/

  32. [32]

    Introducing gpt -5.5, April 2026

    OpenAI. Introducing gpt -5.5, April 2026. URL https://openai.com/index/ introducing-gpt-5-5/

  33. [33]

    System card: Claude sonnet 4.5, September 2025

    Anthropic. System card: Claude sonnet 4.5, September 2025. URL https://www.anthropic. com/claude-sonnet-4-5-system-card

  34. [34]

    System card: Claude sonnet 4.6, February 2026

    Anthropic. System card: Claude sonnet 4.6, February 2026. URL https://www.anthropic. com/claude-sonnet-4-6-system-card

  35. [35]

    System card: Claude opus 4.5, November 2025

    Anthropic. System card: Claude opus 4.5, November 2025. URL https://www.anthropic. com/claude-opus-4-5-system-card

  36. [36]

    System card: Claude opus 4.6, February 2026

    Anthropic. System card: Claude opus 4.6, February 2026. URL https://www.anthropic. com/claude-opus-4-6-system-card

  37. [37]

    System card: Claude opus 4.7, April 2026

    Anthropic. System card: Claude opus 4.7, April 2026. URL https://www.anthropic.com/ claude-opus-4-7-system-card

  38. [38]

    System card: Claude opus 4.8, May 2026

    Anthropic. System card: Claude opus 4.8, May 2026. URL https://www.anthropic.com/ claude-opus-4-8-system-card

  39. [39]

    Deepseek-v3.2: Pushing the frontier of open large language models, 2025

    DeepSeek-AI. Deepseek-v3.2: Pushing the frontier of open large language models, 2025. URL https://arxiv.org/abs/2512.02556

  40. [40]

    Model card: Qwen3.5, March 2026

    QwenTeam. Model card: Qwen3.5, March 2026. URL https://qwen.ai/blog?id=qwen3. 5

  41. [41]

    Kimi k2.5: Visual agentic intelligence, 2026

    Kimi Team. Kimi k2.5: Visual agentic intelligence, 2026. URL https://arxiv.org/abs/ 2602.02276

  42. [42]

    Glm-5: from vibe coding to agentic engineering, 2026

    GLM-5-Team. Glm-5: from vibe coding to agentic engineering, 2026. URL https://arxiv. org/abs/2602.15763

  43. [43]

    Multistep retrosynthesis combining a disconnection aware triple transformer loop with a route penalty score guided tree search.Chem

    David Kreutter and Jean-Louis Reymond. Multistep retrosynthesis combining a disconnection aware triple transformer loop with a route penalty score guided tree search.Chem. Sci., 14(36): 9959–9969, September 2023. doi: 10.1039/D3SC01604H. URL https://pubs.rsc.org/ en/content/articlelanding/2023/sc/d3sc01604h

  44. [44]

    Deep retrosynthetic reaction prediction using local reactivity and global attention.JACS Au, 1(10):1612–1620, October 2021

    Shuan Chen and Yousung Jung. Deep retrosynthetic reaction prediction using local reactivity and global attention.JACS Au, 1(10):1612–1620, October 2021. doi: 10.1021/jacsau.1c00246. URLhttps://pubs.acs.org/doi/10.1021/jacsau.1c00246

  45. [45]

    Aizynthfinder: a fast, robust and flexible open-source software for ret- rosynthetic planning.Journal of Cheminformatics, 12(1):70, Nov 2020

    Samuel Genheden, Amol Thakkar, Veronika Chadimová, Jean-Louis Reymond, Ola Engkvist, and Esben Bjerrum. Aizynthfinder: a fast, robust and flexible open-source software for ret- rosynthetic planning.Journal of Cheminformatics, 12(1):70, Nov 2020. ISSN 1758-2946. doi: 10.1186/s13321-020-00472-1. URLhttps://doi.org/10.1186/s13321-020-00472-1

  46. [46]

    AiZynthFinder 4.0: developments based on learnings from 3 years of industrial application.J

    Lakshidaa Saigiridharan, Alan Kai Hassen, Helen Lai, Paula Torren-Peraire, Ola Engkvist, and Samuel Genheden. AiZynthFinder 4.0: developments based on learnings from 3 years of industrial application.J. Cheminform., 16(57), May 2024. doi: 10.1186/s13321-024-00860-x. URLhttps://link.springer.com/article/10.1186/s13321-024-00860-x. 12

  47. [47]

    Chemist- aligned retrosynthesis by ensembling diverse inductive bias models, 2025

    Krzysztof Maziarz, Guoqing Liu, Hubert Misztela, Austin Tripp, Junren Li, Aleksei Kornev, Piotr Gai´nski, Holger Hoefling, Mike Fortunato, Rishi Gupta, and Marwin Segler. Chemist- aligned retrosynthesis by ensembling diverse inductive bias models, 2025. URL https: //arxiv.org/abs/2412.05269

  48. [48]

    RetroChimera.https://github.com/microsoft/retrochimera, 2025

    Microsoft. RetroChimera.https://github.com/microsoft/retrochimera, 2025

  49. [49]

    Synplanner: An end-to-end tool for synthesis planning

    Tagir Akhmetshin, Dmitry Zankov, Philippe Gantzer, Dmitry Babadeev, Anna Pinigina, Timur Madzhidov, and Alexandre Varnek. Synplanner: An end-to-end tool for synthesis planning. Journal of Chemical Information and Modeling, 65(1):15–21, 2025. doi: 10.1021/acs.jcim. 4c02004. URLhttps://doi.org/10.1021/acs.jcim.4c02004. PMID: 39739735

  50. [50]

    Enamine: Chemical supplier, 2026

    Enamine. Enamine: Chemical supplier, 2026. enamine.net/building-blocks

  51. [51]

    eMolecules, 2026

    eMolecules. eMolecules, 2026. emolecules.com/products/building-blocks

  52. [52]

    Pistachio

    NextMove Software. Pistachio. https://www.nextmovesoftware.com/pistachio.html. Commercial reaction database

  53. [53]

    Daylight Chemical Information Systems

    Inc. Daylight Chemical Information Systems. SMARTS — a language for describing molec- ular patterns, 2007. URL https://www.daylight.com/dayhtml/doc/theory/theory. smarts.html

  54. [54]

    Extraction of organic chemistry grammar from unsupervised learning of chemical reactions.Sci

    Philippe Schwaller, Benjamin Hoover, Jean-Louis Reymond, Hendrik Strobelt, and Teodoro Laino. Extraction of organic chemistry grammar from unsupervised learning of chemical reactions.Sci. Adv., 7(15):eabe4166, April 2021. URL https://www.science.org/doi/10. 1126/sciadv.abe4166

  55. [55]

    Huiyu Ren, Nicole A Bakas, Mitchell Vamos, Apirat Chaikuad, Allison S Limpert, Carina D Wimer, Sonja N Brun, Lester J Lambert, Lutz Tautz, Maria Celeridad, Douglas J Sheffler, Stefan Knapp, Reuben J Shaw, and Nicholas D P Cosford. Design, synthesis, and characterization of an orally active dual-specific ULK1/2 autophagy inhibitor that synergizes with the ...

  56. [56]

    Identify the reactants and products in the reaction

  57. [57]

    Analyze the reaction mechanism, including: • Bond formation and breaking • Electron movement • Intermediates (if any) Identify key structural changes that occur during the reaction. Evaluate the reaction’s characteristics: • Efficiency (yield, number of steps) • Selectivity • Reagents and conditions required • Potential side products Assess the reaction’s...

  58. [58]

    List the identified reactants and products separately

  59. [59]

    Reaction SMILES vs. heavy-atom accounting: confirm that every reactant or reagent whose heavy atoms are incorporated into the product is shown explicitly in the reaction SMILES (with species separated as usual, e.g. by . on the reactant side). Do not treat as a defect the omission of optional agents (e.g. 16 catalysts, additives, solvents) whose heavy ato...

  60. [60]

    Identify and list functional groups present in reactants and products

  61. [61]

    Describe the key steps in the reaction mechanism, including: • Initial bond breaking events • Formation of any intermediates • Final bond formation events

  62. [62]

    Highlight the main structural changes that occur

  63. [63]

    Discuss the electronic and steric factors that influence the reaction’s plausibility

  64. [64]

    Explain the mechanistic rationale for the proposed reaction

  65. [65]

    This structured approach will help ensure a thorough interpretation of the reaction and its citability

    Consider and describe any potential alternative pathways or competing reactions. This structured approach will help ensure a thorough interpretation of the reaction and its citability. After your analysis, provide a detailed assessment of the reaction in the following format: <assessment> Reaction Components: [List identified reactants and products] Funct...

  66. [66]

    dataset) benchmark results. Model URSA-expert-2026 URSA-drugs&clinicals-2026 Solv-0 Solv-1 Solv-2 CC* Solv-0 Solv-1 Solv-2 CC* Proprietary Foundation Models Grok-4.1 67 31 13 1.53 74 48 35 2.65 Grok-4.3 5 1 0 1.30 10 1 1 2.24 Gemini 3.1 Pro 54 37 24 1.83 70 56 49 3.36 GPT 5.1 15 1 1 0.44 11 1 1 0.81 GPT 5.2 11 1 0 0.82 34 12 8 1.91 GPT 5.4 20 2 1 0.89 29 ...

  67. [67]

    Two already merged nodes are exempt from collapsing3

    Collapse two single-component steps into one2. Two already merged nodes are exempt from collapsing3. Return all possible subtrees for the input route A B C Collapse A Collapse B Collapse C Collapse A&B Collapse A&C Collapse B&C Collapse A&B&C Figure 6: Example of the subtree generation process. 29 L ChemCensor Scoring Dynamicatoms RC1 RC2 RC3 RC4 Aryl chl...