{"id":"73297683-a6d5-4dd8-a3f9-e46338704b20","arxiv_id":"2605.21740","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"SMDD-Bench is a new multi-turn benchmark for LLM agents on small molecule drug design where the best tested model solves 40.2% of 502 tasks.","lead":"SMDD-Bench creates a standardized multi-turn benchmark of 502 tasks across five drug design categories and 102 protein targets to test LLM agents on realistic small-molecule tasks. A smart generalist might read it to see how far current frontier models still are from autonomous computational drug design.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No human expert baselines to confirm 'guaranteed-solvable' status of the 502 tasks","rationale":"The reader's weakest assumption matches the load-bearing gap exactly. Without expert baselines the benchmark's difficulty cannot be anchored, so the central performance claim remains uninterpretable as a real-world limitation. This is not a minor omission; it directly affects whether the 40.2% number supports the paper's conclusion about LLM agents for drug design.","tokens_in":1730,"tokens_out":319,"duration_ms":17953,"concrete_test":"Sample 50 tasks uniformly across the 5 types; have 3-5 medicinal chemists attempt each under identical tool/oracle limits and time budget; report solve rate and failure modes. If expert solve rate <70% on the subset, re-evaluate the 'guaranteed-solvable' claim and rescale the LLM percentages.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (GPT5.4 solves 40.2%) is only evidence of LLM limitation on real-world SMDD if the tasks are verifiably solvable by domain experts under the same oracle-call and multi-turn constraints. The paper states the instances are 'guaranteed-solvable' and span 5 task types with 102 targets, yet provides no expert human performance data or inter-annotator agreement on solvability. If the guarantee rests solely on synthetic construction or oracle access without human validation, the 40.2% figure could indicate task over-constraint rather than a genuine capability gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SMDD-Bench, a multi-turn agentic benchmark with 502 guaranteed-solvable task instances spanning five task types (2D Pharmacophore Identification, Interaction Point Discovery, Scaffold Hopping, Lead Optimization, Fragment Assembly) across 102 protein targets. It evaluates seven frontier LLMs on these tasks, which require chemical reasoning, 3D intuition, specialized tool use, and planning under limited oracle calls, and reports that the strongest model (GPT5.4) solves only 40.2% of instances. A public leaderboard at smddbench.com is provided to standardize evaluation of LLM agents for real-world small-molecule drug design.","tokens_in":1856,"tokens_out":543,"duration_ms":21629,"significance":"If the solvability guarantee holds, the benchmark would provide a much-needed standardized, large-scale testbed for long-horizon LLM agents in computational drug design, moving beyond ad-hoc or single-turn evaluations. The public leaderboard and coverage of diverse chemistries and targets are concrete strengths that could drive reproducible progress in the field.","major_comments":[{"comment":"Benchmark construction section: The central claim that GPT5.4 solves only 40.2% of tasks demonstrates LLM limitations on real-world SMDD rests on the assertion that all 502 instances are 'guaranteed-solvable' by domain experts under identical multi-turn and oracle-call constraints. No human expert performance baselines, success rates, or inter-annotator agreement on solvability are reported; without these, the 40.2% figure cannot be unambiguously interpreted as a capability gap rather than task over-constraint.","section":"Benchmark construction"},{"comment":"Task validation and scoring methodology: The paper states tasks are 'guaranteed-solvable' and span five types with 102 targets, yet provides no details on how solvability was verified (e.g., via expert annotation or oracle-based construction) or on the precise scoring rules and oracle-call limits per task type. This information is load-bearing for assessing whether the benchmark fairly measures LLM performance.","section":"Task validation and scoring methodology"}],"minor_comments":[{"comment":"The abstract and introduction would benefit from a brief explicit statement of the oracle-call budget and termination criteria used in the multi-turn setting.","section":"Abstract and introduction"},{"comment":"Figure captions for the task-type illustrations should include the exact number of instances per type to allow quick verification of the 502 total.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive feedback on our manuscript. We address each major comment below and will revise the manuscript to provide greater transparency on benchmark construction and validation.","responses":[{"response":"The solvability guarantee derives from the task construction process, in which each instance was created by starting from a known successful molecule and working backwards using the oracle to confirm a valid action sequence exists within the call limits. This provides a construction-based assurance rather than post-hoc empirical validation. We agree that human expert baselines would aid interpretation and will add a detailed description of the construction method plus a limitations discussion noting the absence of a full human study. We do not believe the tasks are over-constrained, as the oracle budgets were chosen to reflect realistic expert workflows.","revision_made":"partial","referee_comment":"[Benchmark construction] Benchmark construction section: The central claim that GPT5.4 solves only 40.2% of tasks demonstrates LLM limitations on real-world SMDD rests on the assertion that all 502 instances are 'guaranteed-solvable' by domain experts under identical multi-turn and oracle-call constraints. No human expert performance baselines, success rates, or inter-annotator agreement on solvability are reported; without these, the 40.2% figure cannot be unambiguously interpreted as a capability gap rather than task over-constraint."},{"response":"We agree that explicit details on verification and scoring are needed for reproducibility. We will revise the Task validation and scoring methodology section to describe the oracle-based construction process used to verify solvability, along with the exact success criteria and per-task oracle-call limits for each of the five task types.","revision_made":"yes","referee_comment":"[Task validation and scoring methodology] Task validation and scoring methodology: The paper states tasks are 'guaranteed-solvable' and span five types with 102 targets, yet provides no details on how solvability was verified (e.g., via expert annotation or oracle-based construction) or on the precise scoring rules and oracle-call limits per task type. This information is load-bearing for assessing whether the benchmark fairly measures LLM performance."}],"tokens_in":1472,"tokens_out":501,"duration_ms":43776,"standing_objections":["Conducting a full human expert performance study with inter-annotator agreement across all 502 tasks, which would require substantial additional expert time and resources beyond the scope of the current work."]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is SMDD-Bench itself: 502 task instances across five categories (pharmacophore ID, interaction points, scaffold hopping, lead optimization, fragment assembly) on 102 targets, run in a multi-turn agent setting with oracle limits. That setup is more realistic than the single-turn or ad-hoc tests it cites, and the public leaderboard is a practical step.\n\nThe authors did the work of spanning chemical space and defining concrete task types that require planning and 3D reasoning. Benchmarking seven models and reporting the 40.2% top score gives a concrete starting point for the field.\n\nThe soft spot is the \"guaranteed-solvable\" claim. The stress-test note is right that no human expert performance data is mentioned under the same multi-turn and oracle constraints. Without that calibration, the 40% figure could reflect task construction choices rather than a clear LLM limitation. If solvability was verified only through synthetic generation or oracle access, the benchmark risks measuring something narrower than real-world drug design difficulty.\n\nTask construction and scoring details will matter a lot for anyone trying to use the numbers. The paper is aimed at groups building agentic systems for computational chemistry. Readers working on evaluation protocols or training objectives for scientific LLMs will find the task breakdown useful even if they end up adjusting the benchmark.\n\nIt deserves peer review. The benchmark framing is worth referee scrutiny on the validation side, and the field needs standardized tests like this once the details are tightened.","headline":"New multi-turn benchmark for LLM drug design agents with 502 tasks, but the 40% solve rate needs human expert baselines to interpret what it actually shows.","tokens_in":2359,"tokens_out":381,"would_cite":false,"duration_ms":22813,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Even the best large language models solve only 40.2 percent of tasks in a benchmark for small molecule drug design.","keywords":["LLM agents","small molecule drug design","benchmark","computational drug discovery","agentic evaluation","pharmacophore identification","scaffold hopping"],"falsifier":"Finding that human experts solve substantially fewer than 40.2 percent of the tasks or that a model achieves near-complete success would challenge the benchmark's ability to measure the intended capabilities.","tokens_in":2647,"feed_emoji":"💊","tokens_out":402,"duration_ms":41380,"temperature":0.7,"pith_summary":"The paper creates SMDD-Bench as a standardized test with 502 multi-turn tasks across five categories of small molecule design work. These tasks involve 102 different protein targets and demand chemical reasoning, three-dimensional thinking, tool handling, and step-by-step planning. Testing seven leading models shows that the top performer completes just 40.2 percent of the tasks. This result indicates that current systems lack the skills for fully independent computational drug design. The benchmark is meant to drive progress by providing a common way to measure and improve LLM agents in this area.","feed_headline":"LLMs solve only 40% of tasks on small molecule drug design benchmark","feed_subtitle":"New testbed with 502 instances shows top models finish under half, pointing to gaps in reasoning and planning for discovery work.","key_machinery":"SMDD-Bench, the multi-turn long-horizon agentic benchmark that evaluates LLM agents on tasks requiring chemical and biological reasoning, 3D intuition, specialized tool use, and planning with limited oracle calls.","core_discovery":"SMDD-Bench is a challenging multi-turn benchmark of 502 guaranteed-solvable tasks in five categories that span wide chemical space and 102 protein targets, on which even the strongest tested LLM solves only 40.2 percent.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLMs solve 40% on SMDD-Bench drug design tasks","SMDD-Bench shows LLMs solve 40% of tasks","Top models hit 40% on multi-turn SMDD-Bench","LLMs manage 40% on 502 drug design benchmark tasks"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 502 task instances represent real-world small molecule drug design challenges and that their guaranteed-solvable status holds without direct comparison to human expert performance.","fun_headline_variants_meta":{"raw":{"variants":["LLMs solve 40% on SMDD-Bench drug design tasks","SMDD-Bench shows LLMs solve 40% of tasks","Top models hit 40% on multi-turn SMDD-Bench","LLMs manage 40% on 502 drug design benchmark tasks"]},"model":"grok-4.3","cost_usd":0.009545,"raw_usage":{"total_tokens":4261,"prompt_tokens":671,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":95449500,"prompt_tokens_details":{"text_tokens":671,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3516,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":671,"tokens_out":74,"duration_ms":38126,"temperature":1.0,"reasoning_tokens":3516,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T16:54:45.787552+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Finding that human experts solve substantially fewer than 40.2 percent of the tasks or that a model achieves near-complete success would challenge the benchmark's ability to measure the intended capabilities.","supporting_citations":[],"review_version":2}