REVIEW 2 major objections 2 minor 1 cited by
SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?
T0 review · 2 major / 2 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read Even the best large language models solve only 40.2 percent of tasks in a benchmark for small molecule drug design.
desk verdict New multi-turn benchmark for LLM drug design agents with 502 tasks, but the 40% solve rate needs human expert baselines to interpret what it actually shows. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
SMDD-Bench, the multi-turn long-horizon agentic benchmark that evaluates LLM agents on tasks requiring chemical and biological reasoning, 3D intuition, specialized tool use, and planning with limited oracle calls.
What would settle it
Finding that human experts solve substantially fewer than 40.2 percent of the tasks or that a model achieves near-complete success would challenge the benchmark's ability to measure the intended capabilities.
Extended reading notes
Core claim
SMDD-Bench is a challenging multi-turn benchmark of 502 guaranteed-solvable tasks in five categories that span wide chemical space and 102 protein targets, on which even the strongest tested LLM solves only 40.2 percent.
Load-bearing premise
The 502 task instances represent real-world small molecule drug design challenges and that their guaranteed-solvable status holds without direct comparison to human expert performance.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SMDD-Bench, a multi-turn agentic benchmark with 502 guaranteed-solvable task instances spanning five task types (2D Pharmacophore Identification, Interaction Point Discovery, Scaffold Hopping, Lead Optimization, Fragment Assembly) across 102 protein targets. It evaluates seven frontier LLMs on these tasks, which require chemical reasoning, 3D intuition, specialized tool use, and planning under limited oracle calls, and reports that the strongest model (GPT5.4) solves only 40.2% of instances. A public leaderboard at smddbench.com is provided to standardize evaluation of LLM agents for real-world small-molecule drug design.
Significance. If the solvability guarantee holds, the benchmark would provide a much-needed standardized, large-scale testbed for long-horizon LLM agents in computational drug design, moving beyond ad-hoc or single-turn evaluations. The public leaderboard and coverage of diverse chemistries and targets are concrete strengths that could drive reproducible progress in the field.
major comments (2)
- [Benchmark construction] Benchmark construction section: The central claim that GPT5.4 solves only 40.2% of tasks demonstrates LLM limitations on real-world SMDD rests on the assertion that all 502 instances are 'guaranteed-solvable' by domain experts under identical multi-turn and oracle-call constraints. No human expert performance baselines, success rates, or inter-annotator agreement on solvability are reported; without these, the 40.2% figure cannot be unambiguously interpreted as a capability gap rather than task over-constraint.
- [Task validation and scoring methodology] Task validation and scoring methodology: The paper states tasks are 'guaranteed-solvable' and span five types with 102 targets, yet provides no details on how solvability was verified (e.g., via expert annotation or oracle-based construction) or on the precise scoring rules and oracle-call limits per task type. This information is load-bearing for assessing whether the benchmark fairly measures LLM performance.
minor comments (2)
- [Abstract and introduction] The abstract and introduction would benefit from a brief explicit statement of the oracle-call budget and termination criteria used in the multi-turn setting.
- [Figures] Figure captions for the task-type illustrations should include the exact number of instances per type to allow quick verification of the 502 total.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback on our manuscript. We address each major comment below and will revise the manuscript to provide greater transparency on benchmark construction and validation.
read point-by-point responses
-
Referee: [Benchmark construction] Benchmark construction section: The central claim that GPT5.4 solves only 40.2% of tasks demonstrates LLM limitations on real-world SMDD rests on the assertion that all 502 instances are 'guaranteed-solvable' by domain experts under identical multi-turn and oracle-call constraints. No human expert performance baselines, success rates, or inter-annotator agreement on solvability are reported; without these, the 40.2% figure cannot be unambiguously interpreted as a capability gap rather than task over-constraint.
Authors: The solvability guarantee derives from the task construction process, in which each instance was created by starting from a known successful molecule and working backwards using the oracle to confirm a valid action sequence exists within the call limits. This provides a construction-based assurance rather than post-hoc empirical validation. We agree that human expert baselines would aid interpretation and will add a detailed description of the construction method plus a limitations discussion noting the absence of a full human study. We do not believe the tasks are over-constrained, as the oracle budgets were chosen to reflect realistic expert workflows. revision: partial
-
Referee: [Task validation and scoring methodology] Task validation and scoring methodology: The paper states tasks are 'guaranteed-solvable' and span five types with 102 targets, yet provides no details on how solvability was verified (e.g., via expert annotation or oracle-based construction) or on the precise scoring rules and oracle-call limits per task type. This information is load-bearing for assessing whether the benchmark fairly measures LLM performance.
Authors: We agree that explicit details on verification and scoring are needed for reproducibility. We will revise the Task validation and scoring methodology section to describe the oracle-based construction process used to verify solvability, along with the exact success criteria and per-task oracle-call limits for each of the five task types. revision: yes
- Conducting a full human expert performance study with inter-annotator agreement across all 502 tasks, which would require substantial additional expert time and resources beyond the scope of the current work.
Circularity Check
No circularity: empirical benchmark with no derivation chain
full rationale
The paper presents SMDD-Bench as a new empirical evaluation suite of 502 task instances across five task types. No equations, fitted parameters, or self-referential derivations appear in the abstract or described structure. The central performance claim (GPT5.4 solves 40.2%) is a direct measurement on the constructed tasks rather than a prediction obtained by fitting or re-deriving from the benchmark definition itself. The 'guaranteed-solvable' assertion is an empirical construction claim, not a self-definitional loop. This is a standard benchmark paper whose results stand or fall on external validation rather than internal reduction.
Assumptions & free parameters
Cite this review
Pith. "Pith review of SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?." pith.science (2026). https://pith.science/paper/V3FYH6XJ
@misc{pith2026260521740,
author = {Pith},
title = {Pith review of: SMDD-Bench: Can LLMs Solve Real-World Small Molecule Drug Design Tasks?},
year = {2026},
howpublished = {\url{https://pith.science/paper/V3FYH6XJ}},
note = {Machine review of arXiv:2605.21740}
}
read the original abstract
LLM agents have incredible potential for scientific discovery applications. However, the performance of LLM agents on real-world, small molecule drug design (SMDD) tasks across diverse chemistries and targets is unclear. Current evaluation methods are either ad hoc, too simple for real-world discovery, limited in scale, or restricted to single-turn question answering. In effort to standardize the evaluation of LLM agents on small molecule design, we introduce SMDD-Bench, a challenging, multi-turn, long-horizon agentic benchmark consisting of 502 guaranteed-solvable task instances spanning 5 task types: 2D Pharmacophore Identification, Interaction Point Discovery, Scaffold Hopping, Lead Optimization, and Fragment Assembly. SMDD-Bench tasks span a wide region of chemical space and involve 102 unique protein targets. Completely solving the benchmark would require having strong chemical and biological reasoning and 3D intuition, understanding specialized tool use, and displaying planning expertise over a limited number of oracle calls. We benchmark 7 frontier open and closed source LLMs and find even the most performant LLM, GPT5.4, solves only 40.2\% of tasks. We hope SMDD-Bench provides a standardized testbed to invigorate the field towards training and evaluating LLM agents for fully autonomous computational drug design. We host a public leaderboard at smddbench.com .
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows
CAi Copilot, a three-layer LLM agent, converts broad molecular-design requests into executed, evidence-traceable workflows and outperforms five baseline agents on 45 curated tasks plus external benchmarks.
Reference graph
Works this paper leans on
-
[1]
In Figure 5, we plot each of the reference molecule baseline property values for optimization properties and hold-constant properties in the Lead Optimization task. Blood brain barrier penetration can either be optimized to enhance penetration or to decrease penetration depending on the real world application. Both directions of optimization are included ...
-
[2]
**C3 Ketone Interaction:** The ligand's C3 ketone acts as a hydrogen bond acceptor, interacting with the hydrogen bond donors on the protein: Arg752 (NH1/NH2) and Gln711 (NE2).,→
-
[3]
**C17 Hydroxyl Interaction:** The ligand's C17 hydroxyl group (acting as both a donor and an acceptor) interacts with Asn705 (ND2/OD1) and Thr877 (OG1).,→ 40
-
[4]
**Hydrophobic Core:** The steroid backbone sits in a large hydrophobic pocket, likely interacting with residues like Met742 and Met745.,→ So, to find those three key points, I'll need to define the optimal positions for the *ligand's* pharmacophore features based on the *protein's* residues.,→ * **Point 1:** A Ligand Acceptor positioned to interact with A...
-
[5]
Then, I'll project a vector from this midpoint towards the center of the pocket
**Pole 1 (Acceptor):** I'll find the midpoint between the hydrogen bond donors on Arg752 (NH2) and Gln711 (NE2). Then, I'll project a vector from this midpoint towards the center of the pocket. I'll place my Acceptor point roughly 2.5 Angstroms along this vector, representing an ideal hydrogen bond distance. ,→ ,→ ,→
-
[6]
I'll find the midpoint between Asn705 (ND2) and Thr877 (OG1)
**Pole 2 (Donor):** I'll apply the same logic here. I'll find the midpoint between Asn705 (ND2) and Thr877 (OG1). Then, I'll project a vector from this midpoint towards the pocket center, placing the Donor point 2.5 Angstroms along it. ,→ ,→
-
[7]
eq": true} Tanimoto Similaritypassed0.10714285714285714 {
**Pole 3 (Hydrophobic):** For the hydrophobic core, I'll simply use the coordinates provided for the center of the binding pocket.,→ Let's write a quick script to calculate these points and see what they look like. Tool Call: python(turn 5, 0.21s) Reasoning: Refine the coordinates by projecting along the axis between the two poles, which better represents...
-
[8]
**Decrease clearance**: Typically, reducing lipophilicity, adding polar groups, or reducing metabolic liability can help,→
Show all 14 references
-
[9]
**Decrease binding affinity**: This is tricky - we need to make the molecule bind less well (more negative affinity means weaker binding),→
-
[10]
CCOC(=O)[C@H](CCc1ccccc1)O[C@@H](C)C(=O)N1[C@H](C(=O)O)C[C@H]2CCCC[C@@H]21
**Decrease BBB**: Reduce lipophilicity, add polar groups, reduce LogP Let me think about modifications: - Replace ethyl ester with something more polar (like methyl ester or remove it) - Add polar groups to reduce LogP and BBB - Modify the phenyl ring to reduce lipophilicity L...
-
[11]
Binding Probability: ~0.95 (Requirement: > 0.7). Check
-
[12]
(Requirement: < 0.65)
Scaffold Similarity: [...] 9/30 = 0.3. (Requirement: < 0.65). Check
-
[13]
Limitations
Morgan Similarity: 0.29. (Requirement: < 0.5). Check. [...] I'm going to proceed with the SMILES string I've generated. [ERROR — No Exploration]:The agent treats the first passing molecule as the final solution. The entire “design” process consumed a single Boltz call on a mol...
-
[14]
Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects
Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals...
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.