Pith. sign in

REVIEW 3 major objections 6 minor 53 references

General-purpose LLMs can follow explicit 3D spatial instructions to design ligand molecules, but their binding poses remain physically less plausible than those from specialized diffusion models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 15:50 UTC pith:6FT5RJHD

load-bearing objection Useful benchmark with a robust negative result; the positive claim that LLMs can satisfy spatial constraints is inflated by a placement-only success metric. the 3 major comments →

arxiv 2607.18144 v1 pith:6FT5RJHD submitted 2026-07-20 cs.LG cs.AIcs.CL

Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints

classification cs.LG cs.AIcs.CL
keywords structure-based drug design3D molecule generationlarge language modelsspatial constraintsbenchmarkpharmacophoreanchor fragmentsSimplified SDF
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper asks whether general-purpose large language models can reason about 3D geometry well enough to generate small molecules that satisfy realistic spatial constraints in protein binding. To answer it, the authors build 3D-Fit, a benchmark that conditions generation on a protein pocket plus up to three additional spatial requirements: anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. The central finding is that frontier LLMs show an emerging ability to follow such multi-condition spatial instructions, and adding more conditions tends to improve their pocket placement. However, their raw poses contain severe steric clashes and internal geometry problems, and even after local optimization their docking scores and physical validity trail state-of-the-art diffusion models. A fair reader would take away that off-the-shelf LLMs are not yet competitive structure-based drug designers, but they form a flexible substrate for instruction-following 3D generation.

Core claim

On the paper's own terms, the discovery is a comparative capability profile: LLMs can parse compact textual descriptions of spatially explicit molecular generation conditions and produce syntactically valid Simplified SDF output that places atoms within 0.5 Å of requested anchor and pharmacophore coordinates at meaningful success rates, and they can juggle pocket, fragment, pharmacophore, and interaction conditions simultaneously. Yet their generated conformations are physically rough: raw UniDock scores are poor, PoseBusters intra- and inter-molecular pass rates are low, and only local pose optimization brings docking scores into the moderate range. Diffusion baselines that support the same

What carries the argument

The load-bearing machinery is 3D-Fit, a benchmark built on token-efficient textual condition representations——pockets written as sequential residue atom coordinates, anchor fragments as element-plus-sphere constraints, pharmacophore points as typed coordinates——paired with a Simplified SDF output format that removes redundant header fields and uses explicit atom indices and bonds. The mechanism that carries the argument is the benchmark's ability to convert diverse 3D conditions into plain-text instructions an LLM can be prompted with, together with an evaluation protocol that cleanly separates condition satisfaction (atomic placement) from physical validity (PoseBusters filters and UniDock

Load-bearing premise

The paper's success metric for anchor fragments and pharmacophore points counts a condition as satisfied if atoms of the right element land within 0.5 Å of the prescribed coordinates, without checking whether the requested chemical fragment's bond network is actually present.

What would settle it

A reader could falsify the central capability claim by testing a strong LLM on anchor-fragment conditions where the requested fragment is a rare or synthetically unusual substructure, then checking whether the generated molecule's 2D graph contains the exact fragment rather than just scattered atoms near the coordinates; if placement-success rates collapse under 2D connectivity verification, the reported spatial instruction-following would be largely placement mimicry.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the paper is right, general-purpose LLMs can serve as drop-in generators for multi-condition 3D molecular design prompts without retraining, as long as downstream local pose optimization is applied.
  • Adding more explicit seed-ligand conditions (anchor fragments, pharmacophore points, interactions) tends to improve LLM pocket placement, so richer condition sets may be a practical path to better LLM-generated poses.
  • A simplified, indexed SDF-like text output outperforms SMILES+XYZ for this task, indicating that output format design is a first-order factor in LLM 3D molecular generation.
  • Current LLM raw poses are not reliable enough for direct structure-based drug design; they require substantial repositioning, whereas diffusion models' poses are stable under optimization.
  • The benchmark enables scaling to heterogeneous condition combinations that no single diffusion model currently supports natively.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit: because anchor and pharmacophore success only checks 3D placement of atoms of the right element, the reported 'condition satisfied' rates likely overstate chemical reconstruction fidelity; a stricter check on 2D connectivity of the fragment would probably lower the numbers.
  • The success of textual condition-following suggests that stronger reasoning or domain-specific fine-tuning of LLMs could narrow the gap with diffusion models, especially by training on simplified SDF and explicit spatial reasoning tasks.
  • One could extend 3D-Fit to deliberately conflict conditions, such as anchor and pharmacophore constraints that cannot be satisfied simultaneously, to quantify how LLMs prioritize validity over condition satisfaction; the prompt already instructs prioritization but the paper does not measure it.
  • The same benchmarking strategy could be adapted to other 3D design domains, such as materials or protein design, wherever constraints are naturally expressed as point-placement requirements.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces 3D-Fit, a benchmark for evaluating general-purpose LLMs on 3D molecule generation conditioned on protein pockets together with additional spatial constraints: anchor fragments, pharmacophore points, and mandatory protein–ligand interactions. It proposes token-efficient textual representations of these conditions and a Simplified SDF output format, then evaluates a range of proprietary and open-weight LLMs against specialized diffusion baselines on CrossDocked2020 and PLINDER. The headline findings are that LLMs can parse the spatial instructions and frequently place atoms near the requested coordinates, but their raw docking scores are poor and, even after local UniDock optimization, remain behind the best diffusion models; LLM outputs also show weaker internal and external physical validity. The paper concludes that LLMs are increasingly capable of multi-condition generation but do not yet match diffusion models in physical plausibility and binding quality.

Significance. If the results hold, 3D-Fit provides a valuable, standardized benchmark for LLM-based 3D molecular design that goes beyond pocket-only generation to multi-constraint settings. The study's strengths include the breadth of models (11 LLMs and 8 diffusion baselines), use of two established test sets, an ablation of output formats, a code/data link, and explicit acknowledgment of limitations. The conclusion that LLMs lag diffusion models in docking quality and physical plausibility is robust: raw UniDock scores for all LLMs are consistently poor and improve only after local optimization, and optimized scores remain behind MolSnapper, PocketXMol, and PDMD. However, the positive claim that LLMs 'satisfy' anchor fragment and pharmacophore constraints is weakened by the placement-only success metric, which does not verify chemical identity or connectivity. With a stricter metric, the qualitative conclusions about LLM spatial reasoning would need to be reassessed.

major comments (3)
  1. [§5.3, §3] The condition-success metric for Anchor Fragments and Pharmacophore Points is purely 3D atom placement within 0.5 Å; 2D connectivity and feature identity are explicitly not checked. This conflicts with §3's definition of anchor fragments as 'chemically significant ligand substructures' and pharmacophores as 'key functional features.' Consequently the high SRs in Tables 2–4 can be obtained by copying the coordinate list from the prompt and emitting isolated atoms, without forming the requested substructure. §6's observation that LLMs do best when conditions can be 'directly copied, mirrored, or locally reconstructed' indicates this is not a hypothetical concern. The Section 8 claim that LLMs 'satisfy explicit local 3D constraints' is therefore inflated. Please re-evaluate with substructure matching for anchors and functional-group checks for pharmacophores, and report whether the qualitat
  2. [§7] The paper states as a limitation that no statistical significance testing was performed. For a comparative benchmark across many models and conditions, this matters: the claims 'adding more spatial conditions often improves pocket-related metrics' (§6) and 'still remain behind the best pocket-specific diffusion models' (§6) rest on differences in median scores and SRs without any error bars or hypothesis tests. Given the test sets have 948 and 505 complexes, paired bootstrap confidence intervals or a Wilcoxon signed-rank test on per-complex UniDock scores are feasible and should be reported for at least the main comparisons (Tables 1 and 4). Without this, small differences (e.g., among GPT-5.5 and Opus 4.8 in Table 2) cannot be distinguished from noise.
  3. [§5.1/§5.4] The experimental protocol does not specify the number of generated molecules per test complex or how the reported percentages are aggregated across models and targets. The success rate is defined 'on all outputs' but the paper does not state whether one sample per complex or multiple samples were generated, nor how many outputs were produced in total. This information is essential for interpreting the binary metrics, particularly the many 0% and low SRs. Please add the sampling count, total output count per model/dataset, and a description of how multiple samples per complex are aggregated (e.g., per-complex success then averaging, or pooled).
minor comments (6)
  1. [§7] Typo: 'enviornment' should be 'environment'.
  2. [Appendix A] In the pocket description example, 'Residue 109, V AL:' should be 'Residue 109, VAL:'.
  3. [Appendix A prompt] The prompt says 'close the tags and generate a new molecule from scratch in a new set of mol tags,' but the format uses <sdf> tags; make the terminology consistent.
  4. [Reference [5]] Reference [5] has 'URLhttps://' missing a space; should be 'URL https://'.
  5. [§5.1] DeepSeek v3.2 is listed as an open-weight LLM but does not appear in any results table; clarify whether it was excluded and why.
  6. [Tables 1–4] The capitalization of 'UniDock' is inconsistent (e.g., 'Unidock' in headers). Please standardize.

Circularity Check

0 steps flagged

No significant circularity: the benchmark verdict rests on external docking/PoseBusters metrics and held-out diffusion baselines, not on fitted inputs or self-citations.

full rationale

This is an empirical benchmark, not a derivation chain, and no load-bearing step reduces to its own inputs. The condition-success metric is explicitly defined as 3D placement only ('We do not evaluate the 2D-connectivity of Anchor Fragments or other features, only 3D placement', §5.3), and the same 0.5 Å tolerance appears in the prompt and the scorer; but that is a fixed evaluation criterion aligned with the task definition, not a parameter fitted to the results being predicted. The paper's central conclusions are anchored to external, independently defined metrics (UniDock scores, PoseBusters filters, RDKit parsing) and to comparisons with specialist diffusion baselines on held-out CrossDocked2020/PLINDER test splits, so the finding that LLMs lag on physical plausibility is not forced by any fitted quantity. Self-citations to BindGPT and nach0-pc are used only as Related Work context and as an alternative output-format baseline in Appendix C; the Simplified SDF vs Enumerated SMILES+XYZ comparison is an independent experiment reported in Table 5, and no conclusion depends on those works' internal validity. The paper itself flags the fragment-conditioning limitation ('We consider fragments as a 3D constraint independent of chemical validity and the protein environment', §7); this is a construct-validity caveat on the positive claim, not circularity. No uniqueness theorem, ansatz-by-citation, or renamed known result is load-bearing. Score 0.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The benchmark introduces hand-chosen evaluation thresholds rather than fitted physical parameters. The strongest assumptions are metric validity: UniDock and PoseBusters are treated as ground truth for binding quality and plausibility, and the anchor/pharmacophore success criterion is placement-only by explicit design. These choices are declared in the paper, but they are load-bearing for the claim that LLMs 'satisfy' spatial constraints.

free parameters (4)
  • 0.5 Å tolerance radius = 0.5 Å
    Used in textual condition descriptions and condition-success evaluation for anchor atoms and pharmacophore points; hand-chosen; directly sets LLM success rates.
  • 10 Å pocket cutoff = 10 Å
    Residues kept if any atom is within 10 Å of the reference ligand center (§4.3); controls pocket size input and affects task difficulty.
  • Minimum heavy atoms = >25 heavy atoms
    Ligand size filter applied to both datasets (§4.3); biases the test set toward larger, drug-like molecules.
  • CrossDocked binding-quality filters = RMSD<0.5 Å, Vina<-6
    Applied only to CrossDocked2020 (§4.3); prunes low-quality complexes and may raise the apparent quality of the benchmark set.
axioms (5)
  • domain assumption UniDock score is a valid measure of binding pose quality and of model ranking
    The central 'LLMs lag diffusion models' conclusion is based on median UniDock scores before and after local optimization (§5.3, Tables 1-4).
  • domain assumption PoseBusters checks capture physical plausibility of generated molecules
    PoseBusters intra- and inter-molecular filters are used as validity metrics; if these filters are not informative, per-model validity comparisons change.
  • ad hoc to paper Satisfying a condition by placing any atom of the right element within 0.5 Å counts as success; 2D connectivity is not required
    Explicitly stated in §5.3: 'We do not evaluate the 2D-connectivity of Anchor Fragments or other features, only 3D placement.' This assumption underpins the high anchor/pharmacophore success rates for LLMs.
  • domain assumption Random sampling of one interaction, one pharmacophore, and one BRICS fragment represents the range of real multi-constraint drug design
    The multi-condition benchmark (§4.3) uses one random constraint of each type; the authors acknowledge this in §7 as a limitation.
  • domain assumption LLM API outputs with temperature 1.0 and top_p 1.0 reflect model capability rather than lucky or unlucky sampling
    Sampling strategy is given in §5.1, but no multiple seeds or variance estimates are reported, so score stability is unknown.

pith-pipeline@v1.3.0-alltime-deepseek · 22938 in / 12494 out tokens · 120181 ms · 2026-08-01T15:50:15.289858+00:00 · methodology

0 comments
read the original abstract

Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.

Figures

Figures reproduced from arXiv: 2607.18144 by Alex Aliper, Alex Zhavoronkov, Maksim Kuznetsov, Maxim Malkov, Rim Shayakhmetov, Roman Schutski, Thomas MacDougall, Vladimir Aladinskiy.

Figure 2
Figure 2. Figure 2: Visualization of spatial conditions and their corresponding textual descriptions. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Zwitterionic form of Serine amino acid as SDF and Simplified SDF. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Generated structures for 5TBO pocket: raw poses in magenta, UniDock-optimized in cyan. B Generated Structures and Unidock Score Distributions [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Example of Enumerated SMILES+XYZ representation [PITH_FULL_IMAGE:figures/full_fig_p021_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Distributions of Unidock scores for all targets and condition sets, compared between models. [PITH_FULL_IMAGE:figures/full_fig_p023_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 16 canonical work pages

  1. [1]

    Pocket2Mol: Efficient molecular sampling based on 3D protein pockets

    Xingang Peng, Shitong Luo, Jiaqi Guan, Qi Xie, Jian Peng, and Jianzhu Ma. Pocket2Mol: Efficient molecular sampling based on 3D protein pockets. In Kamalika Chaudhuri, Ste- fanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors,Proceed- ings of the 39th International Conference on Machine Learning, volume 162 ofProceed- ings of Machi...

  2. [2]

    3D equivariant diffusion for target-aware molecule generation and affinity prediction

    Jiaqi Guan, Wesley Wei Qian, Xingang Peng, Yufeng Su, Jian Peng, and Jianzhu Ma. 3D equivariant diffusion for target-aware molecule generation and affinity prediction. In The Eleventh International Conference on Learning Representations, 2023. URL https: //openreview.net/forum?id=kJqXEPXMsE0

  3. [3]

    Blundell, Pietro Lio, Max Welling, Michael Bronstein, and Bruno Correia

    Arne Schneuing, Charles Harris, Yuanqi Du, Kieran Didi, Arian Jamasb, Ilia Igashov, Weitao Du, Carla Gomes, Tom L. Blundell, Pietro Lio, Max Welling, Michael Bronstein, and Bruno Correia. Structure-based drug design with equivariant diffusion models.Nature Computational Science, 4(12):899–909, Dec 2024. ISSN 2662-8457. doi: 10.1038/s43588-024-00737-x. URL...

  4. [4]

    A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets.Nature Communications, 15(1):2657, Mar 2024

    Lei Huang, Tingyang Xu, Yang Yu, Peilin Zhao, Xingjian Chen, Jing Han, Zhi Xie, Hailong Li, Wenge Zhong, Ka-Chun Wong, and Hengtong Zhang. A dual diffusion model enables 3d molecule generation and lead optimization based on target pockets.Nature Communications, 15(1):2657, Mar 2024. ISSN 2041-1723. doi: 10.1038/s41467-024-46569-1. URL https: //doi.org/10....

  5. [5]

    Aligning target-aware molecule diffusion models with exact energy optimization

    Siyi Gu, Minkai Xu, Alexander S Powers, Weili Nie, Tomas Geffner, Karsten Kreis, Jure Leskovec, Arash Vahdat, and Stefano Ermon. Aligning target-aware molecule diffusion models with exact energy optimization. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URLhttps://openreview.net/forum?id=EWcvxXtzNu

  6. [6]

    SGEDiff: a subgraph-enriched diffusion model for structure-based 3d molecular generation.Journal of Cheminformatics, 17(1):175, Dec 2025

    Changda Gong, Jiaojiao Fang, Yan Tang, Guixia Liu, Yun Tang, and Weihua Li. SGEDiff: a subgraph-enriched diffusion model for structure-based 3d molecular generation.Journal of Cheminformatics, 17(1):175, Dec 2025. ISSN 1758-2946. doi: 10.1186/s13321-025-01123-z. URLhttps://doi.org/10.1186/s13321-025-01123-z

  7. [7]

    Steering semi-flexible molecular diffusion model for structure-based drug design with reinforcement learning.Science Advances, 12(16):eady9955, 2026

    Xudong Zhang, Sanqing Qu, Fan Lu, Jianmin Wang, Zhixin Tian, Shangding Gu, Yan- ping Zhang, Alois Knoll, Shaorong Gao, Guang Chen, and Changjun Jiang. Steering semi-flexible molecular diffusion model for structure-based drug design with reinforcement learning.Science Advances, 12(16):eady9955, 2026. doi: 10.1126/sciadv.ady9955. URL https://www.science.org...

  8. [8]

    DiffDec: Structure-aware scaffold decoration with an end-to-end diffusion model.Journal of Chemical Information and Modeling, 64(7):2554–2564, 2024

    Junjie Xie, Sheng Chen, Jinping Lei, and Yuedong Yang. DiffDec: Structure-aware scaffold decoration with an end-to-end diffusion model.Journal of Chemical Information and Modeling, 64(7):2554–2564, 2024. doi: 10.1021/acs.jcim.3c01466. URL https://doi.org/10.1021/ acs.jcim.3c01466. PMID: 38267393

  9. [9]

    Protein-ligand interaction prior for binding- aware 3d molecule diffusion models

    Zhilin Huang, Ling Yang, Xiangxin Zhou, Zhilong Zhang, Wentao Zhang, Xiawu Zheng, Jie Chen, Yu Wang, Bin CUI, and Wenming Yang. Protein-ligand interaction prior for binding- aware 3d molecule diffusion models. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=qH9nrMNTIW

  10. [10]

    Yael Ziv, Fergus Imrie, Brian Marsden, and Charlotte M. Deane. MolSnapper: Conditioning diffusion for structure-based drug design.Journal of Chemical Information and Modeling, 65 (9):4263–4273, 2025. doi: 10.1021/acs.jcim.4c02008. URL https://doi.org/10.1021/ acs.jcim.4c02008. PMID: 40248896

  11. [11]

    Interaction-constrained 3d molec- ular generation using a diffusion model enables structure-based pharmacophore model- ing for drug design.npj Drug Discovery, 3(1):8, Mar 2026

    Masami Sako, Nobuaki Yasuo, and Masakazu Sekijima. Interaction-constrained 3d molec- ular generation using a diffusion model enables structure-based pharmacophore model- ing for drug design.npj Drug Discovery, 3(1):8, Mar 2026. ISSN 3005-1452. doi: 10.1038/s44386-026-00040-x. URLhttps://doi.org/10.1038/s44386-026-00040-x. 11

  12. [12]

    Cassady, Michael A

    Debjyoti Bhattacharya, Harrison J. Cassady, Michael A. Hickner, and Wesley F. Reinhart. Large language models as molecular design engines.Journal of Chemical Information and Modeling, 64(18):7086–7096, 2024. doi: 10.1021/acs.jcim.4c01396. URL https://doi.org/10.1021/ acs.jcim.4c01396. PMID: 39231030

  13. [13]

    Conversational drug editing using retrieval and domain feedback

    Shengchao Liu, Jiongxiao Wang, Yijin Yang, Chengpeng Wang, Ling Liu, Hongyu Guo, and Chaowei Xiao. Conversational drug editing using retrieval and domain feedback. InThe Twelfth International Conference on Learning Representations, 2024. URL https://openreview. net/forum?id=yRrPfKyJQ2

  14. [14]

    Baker, Ziqi Chen, Xia Ning, and Huan Sun

    Botao Yu, Frazier N. Baker, Ziqi Chen, Xia Ning, and Huan Sun. LlaSMol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. InFirst Conference on Language Modeling, 2024. URL https://openreview. net/forum?id=lY6XTF9tPv

  15. [15]

    Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files,

    Daniel Flam-Shepherd and Alán Aspuru-Guzik. Language models can generate molecules, materials, and protein binding sites directly in three dimensions as XYZ, CIF, and PDB files,

  16. [16]

    Artem Zholus, Maksim Kuznetsov, Roman Schutski, Rim Shayakhmetov, Daniil Polykovskiy, Sarath Chandar, and Alex Zhavoronkov. BindGPT: A scalable framework for 3d molecular design via language modeling and reinforcement learning.Proceedings of the AAAI Conference on Artificial Intelligence, 39(24):26083–26091, Apr. 2025. doi: 10.1609/aaai.v39i24.34804. URLh...

  17. [17]

    nach0-pc: Multi-task language model with molecular point cloud encoder.Proceedings of the AAAI Conference on Artificial Intelligence, 39(23): 24357–24365, Apr

    Maksim Kuznetsov, Airat Valiev, Alex Aliper, Daniil Polykovskiy, Elena Tutubalina, Rim Shayakhmetov, and Zulfat Miftahutdinov. nach0-pc: Multi-task language model with molecular point cloud encoder.Proceedings of the AAAI Conference on Artificial Intelligence, 39(23): 24357–24365, Apr. 2025. doi: 10.1609/aaai.v39i23.34613. URL https://ojs.aaai.org/ index....

  18. [18]

    Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B

    Paul G. Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B. Iovanisci, Ian Snyder, and David R. Koes. Three-dimensional convolutional neural networks and a cross- docked data set for structure-based drug design.Journal of Chemical Information and Modeling, 60(9):4200–4215, 2020. doi: 10.1021/acs.jcim.0c00411. URL https://doi.org/10.1021/ a...

  19. [19]

    PLINDER: The protein-ligand interactions dataset and resource

    Janani Durairaj, Yusuf Adeshina, Zhonglin Cao, Xuejin Zhang, Vladas Oleinikovas, Thomas Duignan, Zachary McClure, Xavier Robin, Gabriel Studer, Daniel Kovtun, Emanuele Rossi, Guoqing Zhou, Srimukh Veccham, Clemens Isert, Yuxing Peng, Prabindh Sundareson, Mehmet Akdel, Gabriele Corso, Hannes Stärk, Gerardo Tauriello, Zachary Carpenter, Michael Bron- stein,...

  20. [20]

    The PDBbind database: Methodologies and updates.Journal of Medicinal Chemistry, 48(12):4111–4119,

    Renxiao Wang, Xueliang Fang, Yipin Lu, Chao-Yie Yang, and Shaomeng Wang. The PDBbind database: Methodologies and updates.Journal of Medicinal Chemistry, 48(12):4111–4119,

  21. [21]

    Smith, Anthony J

    Swapnil Wagle, Richard D. Smith, Anthony J. Dominic, Debarati DasGupta, Sunil Kumar Tripathi, and Heather A. Carlson. Sunsetting Binding MOAD with its last data update and the addition of 3d-ligand polypharmacology tools.Scientific Reports, 13(1):3008, Feb 2023. ISSN 2045-2322. doi: 10.1038/s41598-023-29996-w. URL https://doi.org/10.1038/ s41598-023-29996-w

  22. [22]

    Haitao Lin, Guojiang Zhao, Odin Zhang, Yufei Huang, Lirong Wu, Cheng Tan, Zicheng Liu, Zhifeng Gao, and Stan Z. Li. CBGBench: Fill in the blank of protein-molecule complex binding graph. InThe Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=mOpNrrV2zH. 12

  23. [23]

    Haoyang Liu, Yifei Qin, Zhangming Niu, Mingyuan Xu, Jiaqiang Wu, Xianglu Xiao, Jinping Lei, Ting Ran, and Hongming Chen. How good are current pocket-based 3d generative models?: The benchmark set and evaluation of protein pocket-based 3d molecular generative models.Journal of Chemical Information and Modeling, 64(24):9260–9275, 2024. doi: 10.1021/acs.jcim...

  24. [24]

    Durian: A comprehensive benchmark for structure-based 3d molecular generation.Journal of Chemical Information and Modeling, 65(1):173–186, 2025

    Dou Nie, Huifeng Zhao, Odin Zhang, Gaoqi Weng, Hui Zhang, Jieyu Jin, Haitao Lin, Yufei Huang, Liwei Liu, Dan Li, Tingjun Hou, and Yu Kang. Durian: A comprehensive benchmark for structure-based 3d molecular generation.Journal of Chemical Information and Modeling, 65(1):173–186, 2025. doi: 10.1021/acs.jcim.4c02232. URL https://doi.org/10.1021/ acs.jcim.4c02...

  25. [25]

    MolDiff: Addressing the atom- bond inconsistency problem in 3D molecule diffusion generation

    Xingang Peng, Jiaqi Guan, Qiang Liu, and Jianzhu Ma. MolDiff: Addressing the atom- bond inconsistency problem in 3D molecule diffusion generation. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 ofPro- ceedin...

  26. [26]

    Haitao Lin, Yufei Huang, Odin Zhang, Siqi Ma, Meng Liu, Xuanjing Li, Lirong Wu, Jishui Wang, Tingjun Hou, and Stan Z. Li. DiffBP: generative diffusion of 3d molecules for target protein binding.Chem. Sci., 16:1417–1431, 2025. doi: 10.1039/D4SC05894A. URL http: //dx.doi.org/10.1039/D4SC05894A

  27. [27]

    Binding-adaptive diffusion models for structure-based drug design.Proceedings of the AAAI Conference on Artificial Intelligence, 38(11):12671–12679, Mar

    Zhilin Huang, Ling Yang, Zaixi Zhang, Xiangxin Zhou, Yu Bao, Xiawu Zheng, Yuwei Yang, Yu Wang, and Wenming Yang. Binding-adaptive diffusion models for structure-based drug design.Proceedings of the AAAI Conference on Artificial Intelligence, 38(11):12671–12679, Mar. 2024. doi: 10.1609/aaai.v38i11.29162. URL https://ojs.aaai.org/index.php/ AAAI/article/view/29162

  28. [28]

    An effective fragment-based dual conditional diffusion framework for molecular generation.Briefings in Bioinformatics, 27(1): bbaf727, 01 2026

    Haotian Chen, Yiting Shen, Jichun Li, and Weizhong Zhao. An effective fragment-based dual conditional diffusion framework for molecular generation.Briefings in Bioinformatics, 27(1): bbaf727, 01 2026. ISSN 1477-4054. doi: 10.1093/bib/bbaf727. URL https://doi.org/10. 1093/bib/bbaf727

  29. [29]

    Berman, John Westbrook, Zukang Feng, Gary Gilliland, T

    Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, T. N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne. The protein data bank.Nucleic Acids Research, 28 (1):235–242, 01 2000. ISSN 0305-1048. doi: 10.1093/nar/28.1.235. URL https://doi.org/ 10.1093/nar/28.1.235

  30. [30]

    Nourse, W

    Arthur Dalby, James G. Nourse, W. Douglas Hounshell, Ann K. I. Gushurst, David L. Grier, Burton A. Leland, and John Laufer. Description of several chemical structure file formats used by computer programs developed at molecular design limited.Journal of Chemical Information and Computer Sciences, 32(3):244–255, May 1992. ISSN 0095-2338. doi: 10.1021/ci000...

  31. [31]

    nach0: multimodal natural and chemical languages foundation model

    Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov, Daniil Polykovskiy, Annika Brundyn, Aastha Jhunjhunwala, Anthony Costa, Alex Aliper, Alán Aspuru-Guzik, and Alex Zhavoronkov. nach0: multimodal natural and chemical languages foundation model. Chem. Sci., 15:8380–8389, 2024. doi: 10.1039/D4SC00966E. URL http://dx.doi.org/10. 1039/D4SC00966E

  32. [32]

    Meireles, and Ivet Bahar

    Ahmet Bakan, Lidio M. Meireles, and Ivet Bahar. ProDy: Protein dynamics inferred from theory and experiments.Bioinformatics, 27(11):1575–1577, 06 2011. ISSN 1367-4803. doi: 10. 1093/bioinformatics/btr168. URLhttps://doi.org/10.1093/bioinformatics/btr168

  33. [33]

    Scalfani, Daniel Probst, Kazuya Ujihara, guillaume godin, Rachel Walker, Juuso Lehtivarjo, Axel Pahl, Francois Berenger, jasondbiggs, and strets123

    Greg Landrum, Paolo Tosco, Brian Kelley, Ric, David Cosgrove, sriniker, gedeck, Riccardo Vianello, NadineSchneider, Eisuke Kawashima, Gareth Jones, Dan N, Andrew Dalke, Brian Cole, Matt Swain, Samo Turk, AlexanderSavelyev, Alain Vaucher, Maciej Wójcikowski, Ichiru Take, Vincent F. Scalfani, Daniel Probst, Kazuya Ujihara, guillaume godin, Rachel Walker, Ju...

  34. [34]

    Oleg Trott and Arthur J. Olson. AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading.Journal of Computational Chemistry, 31(2):455–461, 2010. doi: https://doi.org/10.1002/jcc.21334. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/jcc.21334

  35. [35]

    ProLIF: a library to encode molecular interactions as fingerprints.Journal of Cheminformatics, 13(1):72, Sep 2021

    Cédric Bouysset and Sébastien Fiorucci. ProLIF: a library to encode molecular interactions as fingerprints.Journal of Cheminformatics, 13(1):72, Sep 2021. ISSN 1758-2946. doi: 10.1186/s13321-021-00548-6. URLhttps://doi.org/10.1186/s13321-021-00548-6

  36. [36]

    On the art of compiling and using ’drug-like’ chemical fragment spaces.ChemMedChem, 3(10):1503–1507,

    Jörg Degen, Christof Wegscheid-Gerlach, Andrea Zaliani, and Matthias Rarey. On the art of compiling and using ’drug-like’ chemical fragment spaces.ChemMedChem, 3(10):1503–1507,

  37. [37]

    Ligand- based pharmacophore modeling using novel 3d pharmacophore signatures.Molecules, 23 (12), 2018

    Alina Kutlushina, Aigul Khakimova, Timur Madzhidov, and Pavel Polishchuk. Ligand- based pharmacophore modeling using novel 3d pharmacophore signatures.Molecules, 23 (12), 2018. ISSN 1420-3049. doi: 10.3390/molecules23123094. URL https://www.mdpi. com/1420-3049/23/12/3094

  38. [38]

    GPT-5.5 system card, April 2026

    OpenAI. GPT-5.5 system card, April 2026. URL https://deploymentsafety.openai. com/gpt-5-5/gpt-5-5.pdf

  39. [39]

    GPT-5.4 thinking system card, March 2026

    OpenAI. GPT-5.4 thinking system card, March 2026. URL https://deploymentsafety. openai.com/gpt-5-4-thinking/gpt-5-4-thinking.pdf

  40. [40]

    System card: Claude opus 4.7, April 2026

    Anthropic. System card: Claude opus 4.7, April 2026. URL https://cdn.sanity.io/ files/4zrzovbb/website/037f06850df7fbe871e206dad004c3db5fd50340.pdf

  41. [41]

    System card: Claude opus 4.6, February 2026

    Anthropic. System card: Claude opus 4.6, February 2026. URL https://www-cdn. anthropic.com/6a5fa276ac68b9aeb0c8b6af5fa36326e0e166dd.pdf

  42. [42]

    System card: Claude sonnet 4.6, February 2026

    Anthropic. System card: Claude sonnet 4.6, February 2026. URL https://www-cdn. anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75.pdf

  43. [43]

    Gemini 3.1 pro model card, February 2026

    Gemini Team. Gemini 3.1 pro model card, February 2026. URL https://storage. googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card. pdf

  44. [44]

    Grok 4.1 model card, November 2025

    xAI. Grok 4.1 model card, November 2025. URL https://data.x.ai/ 2025-11-17-grok-4-1-model-card.pdf

  45. [45]

    Qwen3.5: Towards native multimodal agents, February 2026

    Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URL https: //qwen.ai/blog?id=qwen3.5

  46. [46]

    DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Ha...

  47. [47]

    GLM-5: from vibe coding to agentic engineering, 2026

    GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunx- iang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Z...

  48. [48]

    Unified modeling of 3d molecular generation via atomic interactions with pocketXMol.Cell, 189(7):1904–1922.e28, Apr 2026

    Xingang Peng, Ruihan Guo, Fenglin Guo, Ziyi Wang, Jiayu Sun, Jiaqi Guan, Yinjun Jia, Yan Xu, Yanwen Huang, Muhan Zhang, Jian Peng, Xinquan Wang, Chuanhui Han, Zihua Wang, and Jianzhu Ma. Unified modeling of 3d molecular generation via atomic interactions with pocketXMol.Cell, 189(7):1904–1922.e28, Apr 2026. ISSN 0092-8674. doi: 10.1016/j.cell. 2026.01.003...

  49. [49]

    Morris, and Charlotte M

    Martin Buttenschoen, Garrett M. Morris, and Charlotte M. Deane. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences.Chem. Sci., 15:3130–3139, 2024. doi: 10.1039/D3SC04185A. URL http://dx.doi.org/10.1039/ D3SC04185A. 15

  50. [50]

    Uni- Dock: GPU-accelerated docking enables ultralarge virtual screening.Journal of Chemical Theory and Computation, 19(11):3336–3345, 2023

    Yuejiang Yu, Chun Cai, Jiayue Wang, Zonghua Bo, Zhengdan Zhu, and Hang Zheng. Uni- Dock: GPU-accelerated docking enables ultralarge virtual screening.Journal of Chemical Theory and Computation, 19(11):3336–3345, 2023. doi: 10.1021/acs.jctc.2c01145. URL https://doi.org/10.1021/acs.jctc.2c01145. PMID: 37125970. 16 A Benchmarking templates General template f...

  51. [2005]

    URL https://doi.org/10.1021/jm048957q

    doi: 10.1021/jm048957q. URL https://doi.org/10.1021/jm048957q. PMID: 15943484

  52. [2008]

    URL https://chemistry-europe

    doi: https://doi.org/10.1002/cmdc.200800178. URL https://chemistry-europe. onlinelibrary.wiley.com/doi/abs/10.1002/cmdc.200800178

  53. [2023]

    URLhttps://arxiv.org/abs/2305.05708