Pith. sign in

REVIEW 4 major objections 6 minor 127 references

FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction

T0 review · 4 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read FLOWR.ROOT claims that a single SE(3)-equivariant flow-matching model can jointly generate pocket-aware 3D ligands and predict binding affinity at FEP-competitive accuracy, enabling affinity-guided sampling and fast project-specific adaptat

desk verdict Solid generation paper with a leaky affinity benchmark; the authors are honest about the FEP+ overlap, but the abstract still overstates it. read the letter →

arxiv 2510.02578 v6 pith:2N4TOUOA submitted 2025-10-02 q-bio.BM cs.LG

classification q-bio.BMcs.LG
keywords flowmatchingSE(3)-equivariantmodel3Dligandgenerationbindingaffinitypredictionpocket-conditionedfragmentelaborationimportancesamplingstructure-baseddrugdesign
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to establish that one SE(3)-equivariant flow-matching model can do both structure-aware 3D ligand generation and binding-affinity prediction jointly, and that the joint training is what makes each task better rather than a convenient packaging. The authors argue that a three-stage recipe—large-scale pretraining on mixed-fidelity data, refinement on curated complexes, then project-specific parameter-efficient fine-tuning—is the right way to build a drug-design foundation model. If true, structure-based drug design could collapse the usual generate-then-score pipeline into a single model that generates potency-guided molecules, ranks them, and estimates its own confidence, at speeds orders of magnitude beyond alchemical free-energy methods. The load-bearing caveat, which the paper itself concedes, is that the strongest affinity benchmark may overlap the training data, so the headline affinity numbers may not be evidence of generalization.

What carries the argument

The central object is an SE(3)-equivariant flow-matching transport map with a pocket encoder and a ligand decoder carrying three output heads: structure, affinity, and confidence. Continuous flow matching transports coordinates while discrete flow matching handles atom types, bonds, charges, and hybridization. Training follows a three-stage curriculum: pretraining on roughly 1.5 billion small molecules and 2.5 million mixed-fidelity protein-ligand complexes, fine-tuning on curated co-crystal data, and project-specific adaptation via parameter-efficient fine-tuning. Inference-time steering uses sequential Monte Carlo importance sampling to resample particles by a reward such as predicted affi

What would settle it

Compute the overlap between the FEP-type benchmark complexes and the pretraining/fine-tuning corpora (by ligand and target identity, and by protein-ligand interaction similarity). If the overlap is high, the benchmark comparison cannot support the generalization claim. Alternatively, prospectively rank a set of newly synthesized analogs for a target absent from all training data: if the Kendall tau between predicted and measured affinities falls below about 0.4, the generalization claim fails.

Watch

Extended reading notes

Core claim

FLOWR.ROOT is an SE(3)-equivariant flow-matching model that jointly predicts ligand structure, binding affinity, and pose confidence conditioned on a protein pocket. The central claim is that this joint training is not incremental: it lets affinity gradients shape the generative distribution, so that inference-time importance sampling can push generation toward higher predicted potency, and a small amount of project-specific fine-tuning can shift the model onto an unseen structure-activity landscape. Evidence offered: state-of-the-art validity and low strain on GEOM-DRUGS, CrossDocked2020, and SPINDR; affinity correlations of R2=0.84 (pIC50) and Pearson=0.78 (aggregated) on SPINDR; accuracy

Load-bearing premise

The affinity results are taken as evidence of generalization, but the free-energy perturbation benchmark likely overlaps the training data, and the paper provides no overlap analysis; if the overlap is substantial, the reported superiority is memorization, not prediction.

Editorial extensions

If this is right

  • A single model can replace separate generative and scoring stages: generated ligands carry affinity and confidence estimates at no extra cost.
  • Because affinity and structure are trained jointly, fine-tuning the affinity head on project data also adjusts the generated chemical distribution, enabling rapid campaign-specific adaptation.
  • Inference-time importance sampling shifts the distribution of generated ligands toward higher predicted affinity while preserving chemical validity and geometry, as demonstrated on SPINDR and in the CK2-alpha/CLK3 selectivity study.
  • The speed advantage over alchemical free-energy methods makes it practical to rank large generative libraries on the fly, not just post hoc.
  • Quantum-mechanical validation on TYK2, ER-alpha, and BACE1 suggests that predicted affinities track real binding-energy trends, supporting use as a prioritization filter.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The FEP-type benchmark results likely overstate generalization: the paper concedes the benchmark overlaps the training data, and no overlap analysis is provided; the honest reading is that the model is very fast and competitive on near-distribution data, while out-of-distribution affinity prediction still requires fine-tuning (the paper's own in-house experiment shows negligible zero-shot correlat
  • The claimed benefit of joint training could be tested directly by ablating the affinity head during generation and measuring whether steering effectiveness drops; the paper does not isolate this mechanism.
  • The Hsp90 appendix result suggests the model does not capture explicit solvation layers; a testable extension would be conditioning generation on water positions, which could improve affinity ranking in solvent-dominated pockets.
  • A prospective corollary: the most valuable deployment is not zero-shot prediction but continuous learning loops where each round of experimental data is fed back via fine-tuning; the paper gestures at this but does not demonstrate a multi-round loop.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents FLOWR.ROOT, an SE(3)-equivariant flow-matching model that jointly learns 3D ligand generation, binding-affinity prediction (pIC50, pKi, pKd, pEC50), and confidence estimation. The model is trained in three stages: large-scale pre-training on small molecules and mixed-fidelity protein–ligand complexes, fine-tuning on curated high-fidelity datasets (SPINDR, HiQBind), and project-specific adaptation via parameter-efficient fine-tuning. The authors report state-of-the-art unconditional and pocket-conditional ligand generation on GEOM-DRUGS, CrossDocked2020, and SPINDR, competitive affinity prediction on SPINDR, and superior benchmark performance on the Schrödinger FEP+ and OpenFE benchmarks compared with AEV-PLIG, OpenFE, FEP+, and Boltz-2. The paper also demonstrates inference-time importance-sampling steering of generated molecules toward predicted higher affinity, and validates generated compounds through QM calculations in case studies on CK2α/CLK3, TYK2, ERα, and BACE1.

Significance. If the claims hold, FLOWR.ROOT would be a notable contribution: a single backbone unifying structure-aware generation, affinity prediction, confidence scoring, and transferable fine-tuning, with publicly available code and data. The generation results are extensive and internally consistent, with large improvements in PoseBusters validity, strain energy, and Vina scores over strong baselines. The QM-based case studies (especially the CK2α/CLK3 selectivity analysis) provide external evidence beyond the model's own predictions and are a genuine strength. The affinity-prediction contribution, however, is weakened by the authors' own admission that the FEP+/OpenFE benchmark overlaps with training data, and the inference-time steering demonstration is partly circular because the reward is the model's own affinity head. These issues are fixable within scope, so the paper merits revision rather than rejection.

major comments (4)
  1. [§5.5, Figs. 5–6] The abstract and conclusion state that FLOWR.ROOT 'outperforms recent models on the Schrödinger FEP+/OpenFE benchmark,' but §5.5 itself says the FEP+ dataset 'is likely significantly overlapping with the training data in terms of ligand and target space' and that the results should not be read as generalization. This is a load-bearing inconsistency: the affinity-superiority claim rests on a benchmark that may substantially overlap with Stage 1/Stage 2 training data. No overlap analysis (e.g., ligand Tanimoto similarity, target sequence identity, or complex-level redundancy) is provided. As written, the headline comparison against Boltz-2 is not a matched out-of-distribution test. Please either add a quantitative overlap analysis and report results on a non-overlapping subset, or explicitly reframe all FEP+/OpenFE statements as benchmark-specific recall, removing the unqualified 'superior
  2. [§5.4, Fig. 4; §5.8, Fig. 8] The importance-sampling steering evaluation is circular for the claim that the model 'steers design toward higher-affinity compounds.' In Fig. 4 the reward is FLOWR.ROOT's own affinity head, so showing that the mean predicted pIC50 increases is a self-consistency check, not evidence of improving true binding affinity. The CK2α/CLK3 study (Fig. 8–9) partly mitigates this by using QM binding energies as an external metric, but the SPINDR steering results remain unexplained by any external validation. Please either add an external evaluation for the steered molecules (e.g., QM, docking, or experimental data) or clearly label Fig. 4 as illustrating reward optimization against the learned model, not as evidence of true affinity improvement. In addition, the importance weight in §4.3 is written as exp(ŷ_i)/Σ exp(ŷ_j) without a temperature parameter, while the text refers to 'exp(λr)'; clarify
  3. [§4.2, Eq. (2)] The overall loss function as written, L_total = λ_c MSE + λ_t CE + λ_ch CE + λ_h CE + λ_b CE, contains only structure losses. The affinity loss (Huber on pIC50/pKi/pKd/pEC50) and the pLDDT confidence cross-entropy loss, both described later in the same section, are absent from the equation. Since the paper's central novelty is 'jointly trains its confidence and affinity prediction modules with structure generation,' the total loss must include these terms, or the text should explicitly explain why they are omitted from the equation. This is a technical inconsistency that should be corrected.
  4. [§5.6, Fig. 7] The domain-adaptation evaluation uses a random split of ~1,000 in-house ligands for fine-tuning and testing. A random split can overstate adaptation performance because analogs of the same chemical series and similar SAR can appear in both training and test partitions. As the authors frame fine-tuning as the route to 'unseen structure-activity landscapes,' a temporal or scaffold-based split—or at minimum a similarity-based analysis of train/test overlap—should be provided to support that claim. The pre-fine-tuning result (Pearson 0.39, R² −2.18) already shows poor generalization, which makes the split choice more consequential.
minor comments (6)
  1. [§5.5] Typo: 'high-fidelty' should be 'high-fidelity'.
  2. [§5.8] The sentence 'we excluded all complexes for which CLK3 binding energies were above 40.0 kcal/mol, as these resulted from structural clashes ( ˚A5.6 % of all complexes)' contains a garbled symbol placement; the percentage and the Angstrom symbol are mixed. Please fix the formatting.
  3. [§4.2] In the pairwise interaction definition, the text says 'Both S and P are stacked to the existing pairwise message tensor,' but P is not defined. The cross-product tensor is denoted C. This is either a typo (P should be C) or a missing definition.
  4. [§4.2] The normalization constant C=100 in the affinity head z_lig computation is introduced without explanation. Please state how it is chosen and whether it is a fixed hyperparameter or learned.
  5. [Tables 1–3] The phrase 'non-pretrained FLOWR.ROOT base model' is ambiguous. Table 1 appears to evaluate a ligand-only model trained from scratch, while Tables 2–3 evaluate a 'base model' that may or may not have received the full Stage 1 pre-training. Define the training protocol for each table explicitly.
  6. [References] The reference for O'Boyle et al. (Open Babel) contains an inadvertent insertion of 'Del Moral, Arnaud Doucet, and Ajay Jasra' in the author list. Also, some URLs appear malformed (e.g., the DOI in the PoseBusters reference).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: generation and affinity claims rest on independent benchmarks and explicit caveats; self-citations are not load-bearing.

full rationale

The paper's central derivation chain is not circular. Unconditional and pocket-conditional generation are evaluated on external benchmarks (GEOM-DRUGS and CROSSDOCKED2020) against independently published baselines, with SPINDR as an additional comparison. Affinity prediction is based on a multi-stage training pipeline and evaluated on held-out SPINDR test splits and the external FEP+/OpenFE benchmarks. The authors explicitly flag the main threat to the FEP+/OpenFE comparison: 'the Schrodinger FEP+ dataset (and with that the OpenFE dataset) is likely significantly overlapping with the training data in terms of ligand and target space... we do not expect these results to indicate significant generalization capabilities' (§5.5). This is an acknowledged data-overlap validity limitation, not a construction-level circularity. The inference-time steering figure (§5.4) measures the model's own predicted pIC50 under importance-sampling weights derived from that same predicted affinity; this is a self-consistency demonstration rather than an independent empirical prediction, and the paper separately provides independent quantum-mechanical validation (TYK2, ERα, BACE1, CK2α/CLK3) for the steering claims. Self-citations to FLOWR, SPINDR, and PILOT refer to prior architecture and baseline work; no uniqueness theorem or prior result by the authors is invoked to force the conclusions. No equation is defined in terms of the quantity it is used to predict, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 7 free parameters · 6 assumptions · 0 invented entities

No new physical entities, forces, or conserved quantities are introduced. The model heads (affinity, confidence) are architectural components, not invented entities. The most significant assumptions are the comparability of affinity endpoints and the correctness of the split/leakage handling, both of which are only partially validated.

free parameters (7)
  • Pocket cutoff radius = 7 Å
    Used to extract protein pockets; chosen by hand, not fitted. Affects all pocket-conditional generation and affinity predictions.
  • Pocket size limits = 10-800 atoms
    Data filtering thresholds to keep complexes tractable.
  • pLDDT distance thresholds and bins = τ={0.5,1,2,4} Å, 50 bins
    Confidence head design choices; not fitted to data.
  • Importance sampling weight/steering duration = not specified (steering from 0.3 to 0.5 of trajectory)
    Hyperparameters for inference-time steering; reward exponent λ is not reported.
  • Multi-task loss weights = unspecified
    Weights balancing coordinate, atom type, charge, hybridization, bond, affinity losses; critical for training but not reported.
  • Affinity head normalization constant C = 100
    Arbitrary normalization in the ligand pooling (Eq. in 4.2).
  • LoRA rank / trainable params = ~10M trainable
    Parameter-efficient fine-tuning configuration; rank not specified.
assumptions (6)
  • standard math Flow matching transport map and ODE integration are valid for molecular generation
    The model relies on continuous/discrete flow matching theory (Lipman et al.; Campbell et al.) without re-deriving it.
  • domain assumption SE(3)-equivariant message passing captures relevant protein-ligand interactions
    The architecture assumes equivariant features are sufficient for pocket conditioning and affinity.
  • domain assumption Affinity labels from different datasets can be treated as comparable despite different assay types
    The model trains separate heads per endpoint, but the evaluation combines them (median) and converts to ΔG, assuming they measure the same underlying binding free energy.
  • domain assumption The Plinder split prevents information leakage across all aggregated datasets
    The paper states it follows Plinder splits, but does not demonstrate that other datasets (SAIR, BindingNet, HiQBind) do not share complexes with the test set. This is critical for the affinity evaluation.
  • domain assumption Semi-empirical GFN2-xTB binding energies are a proxy for experimental affinity in case studies
    Used to validate predicted affinities; GFN2-xTB is an approximation.
  • domain assumption Schrödinger FEP+ experimental values are reliable ground truth
    Used as benchmark; assumes the curated FEP+ data is accurate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction." pith.science (2026). https://pith.science/paper/2N4TOUOA

@misc{pith2026251002578,
  author       = {Pith},
  title        = {Pith review of: FLOWR.root: A flow matching based foundation model for joint multi-purpose structure-aware 3D ligand generation and affinity prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2N4TOUOA}},
  note         = {Machine review of arXiv:2510.02578}
}
abstract

We present FLOWR.root, an SE(3)-equivariant flow-matching model for pocket-aware 3D ligand generation with joint potency and binding affinity prediction and confidence estimation. The model supports de novo generation, interaction- and pharmacophore-conditional sampling, fragment elaboration and replacement, and multi-endpoint affinity prediction (pIC50, pKi, pKd, pEC50). Training combines large-scale ligand libraries with mixed-fidelity protein-ligand complexes, refined on curated co-crystal datasets and adapted to project-specific data through parameter-efficient finetuning. The base FLOWR.root model achieves state-of-the-art performance in unconditional 3D molecule and pocket-conditional ligand generation. On HiQBind, the pre-trained and finetuned model demonstrates highly accurate affinity predictions, and outperforms recent state-of-the-art methods such as Boltz-2 on the FEP+/OpenFE benchmark with substantial speed advantages. However, we show that addressing unseen structure-activity landscapes requires domain adaptation; parameter-efficient LoRA finetuning yields marked improvements on diverse proprietary datasets and PDE10A. Joint generation and affinity prediction enable inference-time scaling through importance sampling, steering design toward higher-affinity compounds. Case studies validate this: selective CK2$\alpha$ ligand generation against CLK3 shows significant correlation between predicted and quantum-mechanical binding energies. Scaffold elaboration on ER$\alpha$, TYK2, and BACE1 demonstrates strong agreement between predicted affinities and QM calculations while confirming geometric fidelity. By integrating structure-aware generation, affinity estimation, property-guided sampling, and efficient domain adaptation, FLOWR.root provides a comprehensive foundation for structure-based drug design from hit identification through lead optimization.

Figures

Figures reproduced from arXiv: 2510.02578 by the authors.

Figure 1
Figure 1. Overview of the dataset generation pipeline used in this work. Dataset generation workflow comprising data filtering, curation via Schrodinger’s LigPrep and PrepWizard, building of metadata-annotated internal representation, and calculation of molecule statistics. To comprehensively train and evaluate FLOWR.ROOT for structure-aware ligand design, we leverage a diverse collection of public datasets spanning both smal… view at source ↗
Figure 2
Figure 2. Graphical overview of the FLOWR.ROOT framework. FLOWR.ROOT is a flow matching-based framework for joint prediction of 3D ligand structure, binding affinity, and confidence estimation. The model follows a multi-stage training paradigm: large-scale pre-training on small molecules and mixed-fidelity protein-ligand complexes, followed by high-fidelity dataset training, with optional project-specific domain adaptation. D… view at source ↗
Figure 3
Figure 3. Top left: Correlation plot of FLOWR.ROOT-predicted pIC50 in kcal/mol vs. experimental pIC50 binding affinities across protein-ligand complexes on the SPINDR test set, with shaded regions indicating 0.5 and 1 kcal/mol error boundaries, and color denoting density of predictions (the darker the denser). Error bars are reported as standard deviations from five seed runs. Top right: Correlation with experimental pKi affi… view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: Top left: Inference-time steering via importance sampling on the SPINDR test set using the FLOWR.ROOT model comparing the distribution of pIC50 predictions of generated ligands across test set targets between un-guided, and mild to strongly guided steering. Top right: …
Figure 5
Figure 5. Figure 5: Top left: Correlation plot of FLOWR.ROOT-predicted pKd in kcal/mol vs. experimental binding affinities across protein-ligand complexes of the SCHRODINGER FEP+ BENCHMARK dataset, with shaded regions indicating 0.5 and 1 kcal/mol error boundaries, and color denoting dens…
Figure 6
Figure 6. Figure 6: Top left: Correlation plot of FLOWR.ROOT-predicted pKd in kcal/mol vs. experimental binding affinities across protein-ligand complexes of the OPENFE INDUSTRYBENCHMARK dataset, with shaded regions indicating 0.5 and 1 kcal/mol error boundaries, and color denoting densit…
Figure 7
Figure 7. Figure 7: Evaluation of FLOWR.ROOT’s pIC50 and pKi affinity prediction accuracy on a challenging in-house project￾specific dataset comprising around 1, 000 experimentally tested ligands after finetuning. We examine the performance of FLOWR.ROOT on a hold-out test set generated b…
Figure 8
Figure 8. Figure 8: Kernel density estimation plot comparing joint optimization (purple, maximizing CK2 [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Quantum mechanical analysis of the joint ligand design to achieve selectivity towards CK2 [PITH_FULL_IMAGE:figures/full_fig_p022_9.png]
Figure 10
Figure 10. Figure 10: FLOWR.ROOT binding affinity validation using quantum mechanical calculations. Benchmark cases of the TYK2 kinase (top block) and ERα (lower block). Top Left: correlation plot between FLOWR.ROOT predicted affinities and the quantum mechanical binding energies for TYK2.…
Figure 11
Figure 11. Figure 11: FLOWR.ROOT binding affinity validation using quantum mechanical calculations. Benchmark case of BACE1. Left: correlation plot between FLOWR.ROOT predicted affinities and the respective quantum mechanical binding energies. Middle: example of the best binder in the seri…
Figure 12
Figure 12. Figure 12: Distributions of ligand atom types, protein atom types, and protein residue types across multiple datasets. [PITH_FULL_IMAGE:figures/full_fig_p033_12.png]
Figure 13
Figure 13. Figure 13: Overview of affinity data distributions and dataset coverage. [PITH_FULL_IMAGE:figures/full_fig_p034_13.png]
Figure 14
Figure 14. Figure 14: Comparison of key continuous molecular properties of ligands—including molecular weight, logP, molar [PITH_FULL_IMAGE:figures/full_fig_p035_14.png]
Figure 15
Figure 15. Figure 15: Comparison of the distributions of discrete ligand features such as hydrogen bond acceptors and donors, [PITH_FULL_IMAGE:figures/full_fig_p036_15.png]
Figure 16
Figure 16. Figure 16: On the left, the structure of anaplastic lymphoma kinase (ALK, PDB ID: 2XB7, blue) with docked [PITH_FULL_IMAGE:figures/full_fig_p037_16.png]
Figure 17
Figure 17. Figure 17: FLOWR.ROOT validation on Hsp90. Top Block: test1 validation, with more ligands but without explicit waters. Left Block: correlation between FLOWR.ROOT affinity metrics and QM binding energies. Middle Block: schematic representation of the best binder in the series. Ri…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

127 extracted references · 38 canonical work pages

  1. [1]

    Harder, B

    Robert Abel, Lingle Wang, Edward D. Harder, B. J. Berne, and Richard A. Friesner. Advancing drug discovery through enhanced free energy calculations. Accounts of Chemical Research, 50 0 (7): 0 1625--1632, 2017. doi:10.1021/acs.accounts.7b00083

  2. [2]

    Ballard, Joshua Bambrick, Sebastian W

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J. Ballard, Joshua Bambrick, Sebastian W. Bodenstein, David A. Evans, Chia-Chun Hung, Michael O'Neill, David Reiman, Kathryn Tunyasuvunakool, Zachary Wu, Akvil \. e Z emgulyt \. e , Eirini Arvaniti, Charles Beattie, Ottavia Bert...

  3. [3]

    Albergo and Eric Vanden-Eijnden

    Michael S. Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. ICLR, 2023. arXiv:2209.15571

  4. [4]

    Stochastic interpolants: A unifying framework for flows and diffusions

    Michael S Albergo, Nicholas M Boffi, and Eric Vanden-Eijnden. Stochastic interpolants: A unifying framework for flows and diffusions. arXiv preprint, 2023. arXiv:2303.08797

  5. [5]

    Bodkin, Stefan Knapp, and Philip C

    Matteo Aldeghi, Alexander Heifetz, Michael J. Bodkin, Stefan Knapp, and Philip C. Biggin. Accurate calculation of the absolute free energy of binding for drug molecules. Chem. Sci., 7: 0 207--218, 2016. doi:10.1039/C5SC02678D. URL http://dx.doi.org/10.1039/C5SC02678D

  6. [6]

    Evaluating the use of absolute binding free energy in fragment optimization

    Ibrahim Alibay et al. Evaluating the use of absolute binding free energy in fragment optimization. Communications Chemistry, 5 0 (1): 0 93, 2022

  7. [7]

    Guided docking as a data generation approach facilitates structure-based machine learning on kinases

    Michael Backenk \"o hler, Joschka Gro , Verena Wolf, and Andrea Volkamer. Guided docking as a data generation approach facilitates structure-based machine learning on kinases. Journal of Chemical Information and Modeling, 64 0 (10): 0 4009--4020, May 2024. ISSN 1549-9596. doi:10.1021/acs.jcim.4c00055. URL https://doi.org/10.1021/acs.jcim.4c00055

  8. [8]

    Ballester and John B

    Pedro J. Ballester and John B. O. Mitchell. A machine learning approach to predicting protein--ligand binding affinity with applications to molecular docking. Bioinformatics, 26 0 (9): 0 1169--1175, 2010. doi:10.1093/bioinformatics/btq112

Show all 127 references
  1. [9]

    Gfn2-xtb---an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions

    Christoph Bannwarth, Sebastian Ehlert, and Stefan Grimme. Gfn2-xtb---an accurate and broadly parametrized self-consistent tight-binding quantum chemical method with multipole electrostatics and density-dependent dispersion contributions. Journal of Chemical Theory and Computat...

  2. [10]

    Unprecedented selectivity and structural determinants of a new class of protein kinase CK2 inhibitors in clinical trials for the treatment of cancer

    Roberto Battistutta, Giorgio Cozza, Fabrice Pierre, Elena Papinutto, Graziano Lolli, Stefania Sarno, Sean E O'Brien, Adam Siddiqui-Jain, Mustapha Haddach, Kenna Anderes, David M Ryckman, Flavio Meggio, and Lorenzo A Pinna. Unprecedented selectivity and structural determinants ...

  3. [11]

    Bolton, Jie Chen, Sunghwan Kim, Lianyi Han, Siqian He, Wenyao Shi, Vahan Simonyan, Yan Sun, Paul A

    Evan E. Bolton, Jie Chen, Sunghwan Kim, Lianyi Han, Siqian He, Wenyao Shi, Vahan Simonyan, Yan Sun, Paul A. Thiessen, Jiyao Wang, Bo Yu, Jian Zhang, and Stephen H. Bryant. Pubchem3d: a new resource for scientists. Journal of Cheminformatics, 3 0 (1): 0 32, Sep 2011. ISSN 1758-...

  4. [12]

    Absolute binding free energies: A quantitative approach for their calculation

    Stefan Boresch, Franz Tettinger, Martin Leitgeb, and Martin Karplus. Absolute binding free energies: A quantitative approach for their calculation. The Journal of Physical Chemistry B, 107 0 (35): 0 9535--9551, 2003. doi:10.1021/jp0217839

  5. [13]

    Prolif: a library to encode molecular interactions as fingerprints

    C \'e dric Bouysset and S \'e bastien Fiorucci. Prolif: a library to encode molecular interactions as fingerprints. Journal of Cheminformatics, 13 0 (1): 0 72, Sep 2021. ISSN 1758-2946. doi:10.1186/s13321-021-00548-6. URL https://doi.org/10.1186/s13321-021-00548-6

  6. [14]

    Deane, and Garrett M

    Fergus Boyles, Charlotte M. Deane, and Garrett M. Morris. Learning from the ligand: using ligand-based features to improve binding affinity prediction. Bioinformatics, 36 0 (3): 0 758--764, 2020. doi:10.1093/bioinformatics/btz665

  7. [15]

    Morris, and Charlotte M

    Martin Buttenschoen, Garrett M. Morris, and Charlotte M. Deane. Posebusters: Ai-based docking methods fail to generate physically valid poses or generalise to novel sequences, 2024. URL http://dx.doi.org/10.1039/D3SC04185A

  8. [16]

    Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design, 2024

    Andrew Campbell, Jason Yim, Regina Barzilay, Tom Rainforth, and Tommi Jaakkola. Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design, 2024. URL https://arxiv.org/abs/2402.04997

  9. [17]

    Relative binding free energy calculations in drug discovery: Recent advances and practical considerations

    Zoe Cournia, Bryce Allen, and Woody Sherman. Relative binding free energy calculations in drug discovery: Recent advances and practical considerations. Journal of Chemical Information and Modeling, 57 0 (12): 0 2911--2937, 2017. doi:10.1021/acs.jcim.7b00564

  10. [18]

    Julian Cremer, Tuan Le, Frank Noé, Djork-Arné Clevert, and Kristof T. Schütt. Pilot: equivariant diffusion for pocket-conditioned de novo ligand generation with multi-objective guidance via importance sampling. Chem. Sci., 15: 0 14954--14967, 2024. doi:10.1039/D4SC03523B. URL ...

  11. [19]

    Flowr: Flow matching for structure-aware de novo, interaction- and fragment-based ligand generation

    Julian Cremer, Ross Irwin, Alessandro Tibot, Jon Paul Janet, Simon Olsson, and Djork-Arné Clevert. Flowr: Flow matching for structure-aware de novo, interaction- and fragment-based ligand generation. arXiv preprint, 2025

  12. [20]

    Daniel Crusius, Flaviu Cipcigan, and Philip C. Biggin. Are we fitting data or noise? analysing the predictive power of commonly used datasets in drug-, materials-, and molecular-discovery. Faraday Discussions, 256: 0 304--321, 2025. doi:10.1039/D4FD00091A

  13. [21]

    Davis, Jeremy P

    Mindy I. Davis, Jeremy P. Hunt, Sanna Herrgard, Pietro Ciceri, Lisa M. Wodicka, Gabriel Pallares, Michael Hocker, Daniel K. Treiber, and Patrick P. Zarrinkar. Comprehensive analysis of kinase inhibitor selectivity. Nature Biotechnology, 29 0 (11): 0 1046--1051, Nov 2011. ISSN ...

  14. [22]

    Sequential monte carlo samplers

    Pierre Del Moral, Arnaud Doucet, and Ajay Jasra. Sequential monte carlo samplers. Journal of the Royal Statistical Society Series B: Statistical Methodology, 68 0 (3): 0 411--436, 05 2006. ISSN 1369-7412. doi:10.1111/j.1467-9868.2006.00553.x. URL https://doi.org/10.1111/j.1467...

  15. [23]

    Plinder: The protein-ligand interactions dataset and evaluation resource

    Janani Durairaj, Yusuf Adeshina, Zhonglin Cao, Xuejin Zhang, Vladas Oleinikovas, Thomas Duignan, Zachary McClure, Xavier Robin, Danny Kovtun, Emanuele Rossi, Guoqing Zhou, Srimukh Veccham, Clemens Isert, Yuxing Peng, Prabindh Sundareson, Mehmet Akdel, Gabriele Corso, Hannes St...

  16. [24]

    Tillack, and Stefano Forli

    Jerome Eberhardt, Diogo Santos-Martins, Andreas F. Tillack, and Stefano Forli. Autodock vina 1.2.0: New docking methods, expanded force field, and python bindings. Journal of Chemical Information and Modeling, 61 0 (8): 0 3891--3898, Aug 2021. ISSN 1549-9596. doi:10.1021/acs.j...

  17. [25]

    Robust and efficient implicit solvation model for fast semiempirical methods

    Sebastian Ehlert, Marcel Stahn, Sebastian Spicher, and Stefan Grimme. Robust and efficient implicit solvation model for fast semiempirical methods. Journal of Chemical Theory and Computation, 17 0 (7): 0 4250--4261, Jul 2021. ISSN 1549-9618. doi:10.1021/acs.jctc.1c00471. URL h...

  18. [26]

    Absolute binding free energy calculations improve virtual screening in structure-based drug discovery

    Mingyue Feng et al. Absolute binding free energy calculations improve virtual screening in structure-based drug discovery. Scientific Reports, 12 0 (1): 0 14355, 2022 a

  19. [27]

    Syntalinker-hybrid: A deep learning approach for target specific drug design

    Yu Feng, Yuyao Yang, Wenbin Deng, Hongming Chen, and Ting Ran. Syntalinker-hybrid: A deep learning approach for target specific drug design. Artificial Intelligence in the Life Sciences, 2: 0 100035, 2022 b . ISSN 2667-3185. doi:https://doi.org/10.1016/j.ailsci.2022.100035. UR...

  20. [28]

    Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B

    Paul G. Francoeur, Tomohide Masuda, Jocelyn Sunseri, Andrew Jia, Richard B. Iovanisci, Ian Snyder, and David R. Koes. Three-dimensional convolutional neural networks and a cross-docked data set for structure-based drug design. Journal of Chemical Information and Modeling, 60 0...

  21. [29]

    Friesner et al

    Richard A. Friesner et al. Glide: a new approach for rapid, accurate docking and scoring. 1. method and assessment of docking accuracy. Journal of Medicinal Chemistry, 47 0 (7): 0 1739--1749, 2004

  22. [30]

    Gumbart, Fran c ois Dehez, Beno \^i t Roux, Wensheng Cai, and Christophe Chipot

    Haohao Fu, Haochuan Chen, Marharyta Blazhynska, Emma Goulard Coderc de Lacam, Florence Szczepaniak, Anna Pavlova, Xueguang Shao, James C. Gumbart, Fran c ois Dehez, Beno \^i t Roux, Wensheng Cai, and Christophe Chipot. Accurate determination of protein:ligand standard binding ...

  23. [31]

    Niklas W. A. Gebauer, Michael Gastegger, and Kristof T. Schütt. Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules, 2020. URL https://arxiv.org/abs/1906.00957

  24. [32]

    u ller, and Kristof T. Sch \

    Niklas W. A. Gebauer, Michael Gastegger, Stefaan S. P. Hessmann, Klaus-Robert M \"u ller, and Kristof T. Sch \"u tt. Inverse design of 3d molecular structures with conditional generative neural networks. Nature Communications, 13 0 (1): 0 973, Feb 2022. ISSN 2041-1723. doi:10....

  25. [33]

    Mahdi Ghorbani, Leo Gendelev, Paul Beroza, and Michael J. Keiser. Autoregressive fragment-based diffusion for pocket-aware ligand design. arXiv preprint arXiv:2401.05370, 2023

  26. [34]

    Gilson and Huan-Xiang Zhou

    Michael K. Gilson and Huan-Xiang Zhou. Calculation of protein--ligand binding affinities. Annual Review of Biophysics and Biomolecular Structure, 36: 0 21--42, 2007. doi:10.1146/annurev.biophys.36.040306.132550

  27. [35]

    Gilson, James A

    Michael K. Gilson, James A. Given, Bruce L. Bush, and J. Andrew McCammon. The statistical‐thermodynamic basis for computation of binding affinities: A critical review. Biophysical Journal, 72 0 (3): 0 1047--1069, 1997. doi:10.1016/S0006-3495(97)78866-9

  28. [36]

    Holger Gohlke and David A. Case. Converging free energy estimates: Mm-pb(gb)sa studies on the protein–protein complex ras–raf. Journal of Computational Chemistry, 25 0 (2): 0 238--250, 2004. doi:10.1002/jcc.10379

  29. [37]

    Popowicz, Oliver Plettenburg, Grzegorz Dubin, Filipe Menezes, and Anna Czarna

    Przemyslaw Grygier, Katarzyna Pustelny, Jakub Nowak, Przemyslaw Golik, Grzegorz M. Popowicz, Oliver Plettenburg, Grzegorz Dubin, Filipe Menezes, and Anna Czarna. Silmitasertib (cx-4945), a clinically used ck2-kinase inhibitor with additional effects on gsk3b and dyrk1a kinases...

  30. [38]

    3d equivariant diffusion for target-aware molecule generation and affinity prediction

    Jiaqi Guan, Wesley Wei Qian, Xingang Peng, Yufeng Su, Jian Peng, and Jianzhu Ma. 3d equivariant diffusion for target-aware molecule generation and affinity prediction. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?...

  31. [39]

    Link-invent: generative linker design with reinforcement learning

    Jeff Guo, Franziska Knuth, Christian Margreitter, Jon Paul Janet, Kostas Papadopoulos, Ola Engkvist, and Atanas Patronov. Link-invent: generative linker design with reinforcement learning. Digital Discovery, 2: 0 392--408, 2023. doi:10.1039/D2DD00115B. URL http://dx.doi.org/10...

  32. [40]

    Hadfield et al

    Thomas E. Hadfield et al. Strife: Structure-informed fragment elaboration. Journal of Chemical Information and Modeling, 62 0 (10): 0 2343--2357, 2022

  33. [41]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33: 0 6840--6851, 2020 a

  34. [42]

    Denoising diffusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. NeurIPS, 2020 b . arXiv:2006.11239

  35. [43]

    Free energy calculations by the molecular mechanics poisson–boltzmann surface area method

    Nadine Homeyer and Holger Gohlke. Free energy calculations by the molecular mechanics poisson–boltzmann surface area method. Molecular Informatics, 31 0 (2): 0 114--122, 2012. doi:10.1002/minf.201100135

  36. [44]

    Equivariant diffusion for molecule generation in 3 D

    Emiel Hoogeboom, V\' ctor Garcia Satorras, Cl \'e ment Vignac, and Max Welling. Equivariant diffusion for molecule generation in 3 D . In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International...

  37. [45]

    Benson, Randy D

    Li Hu, Matthew L. Benson, Randy D. Smith, Matthew G. Lerner, and Heather A. Carlson. Binding moad (mother of all databases). Proteins: Structure, Function, and Bioinformatics, 60 0 (3): 0 333--340, 2005. doi:10.1002/prot.20512

  38. [46]

    Equivariant 3d-conditional diffusion model for molecular linker design

    Ilia Igashov, Hannes St \"a rk, Cl \'e ment Vignac, Arne Schneuing, Victor Garcia Satorras, Pascal Frossard, Max Welling, Michael Bronstein, and Bruno Correia. Equivariant 3d-conditional diffusion model for molecular linker design. Nature Machine Intelligence, 6 0 (4): 0 417--...

  39. [47]

    Deep generative models for 3d linker design

    Fergus Imrie et al. Deep generative models for 3d linker design. Journal of Chemical Information and Modeling, 60 0 (4): 0 1983--1995, 2020

  40. [48]

    Irwin, Khanh G

    John J. Irwin, Khanh G. Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R. Wong, Munkhzul Khurelbaatar, Yurii S. Moroz, John Mayfield, and Roger A. Sayle. Zinc20---a free ultralarge-scale chemical database for ligand discovery. Journal of Chemical Information and Model...

  41. [49]

    Efficient 3d molecular generation with flow matching and scale optimal transport, 2024

    Ross Irwin, Alessandro Tibo, Jon Paul Janet, and Simon Olsson. Efficient 3d molecular generation with flow matching and scale optimal transport, 2024. URL https://arxiv.org/abs/2406.07266

  42. [50]

    Interactiongraphnet: A novel and efficient deep graph representation learning framework for accurate protein--ligand interaction predictions

    Dejun Jiang, Chang-Yu Hsieh, Zhenxing Wang, Yue Kang, Jing Wang, Binbin Liao, Chun Shen, and Li Xu. Interactiongraphnet: A novel and efficient deep graph representation learning framework for accurate protein--ligand interaction predictions. Journal of Medicinal Chemistry, 64 ...

  43. [51]

    k\_ DEEP : Protein–ligand absolute binding affinity prediction via 3d-convolutional neural networks

    Jos \'e Jim \'e nez, Miha S kali c , Gerard Mart \' nez-Rosell, and Gianni De Fabritiis. k\_ DEEP : Protein–ligand absolute binding affinity prediction via 3d-convolutional neural networks. Journal of Chemical Information and Modeling, 58 0 (2): 0 287--296, 2018. doi:10.1021/a...

  44. [52]

    Torsional diffusion for molecular conformer generation, 2023

    Bowen Jing, Gabriele Corso, Jeffrey Chang, Regina Barzilay, and Tommi Jaakkola. Torsional diffusion for molecular conformer generation, 2023. URL https://arxiv.org/abs/2206.01729

  45. [53]

    Johnston, Kun Yao, Zachary Kaplan, Monica Chelliah, Karl Leswing, Sean Seekins, Shawn Watts, David Calkins, Jackson Chief Elk, Steven V

    Ryne C. Johnston, Kun Yao, Zachary Kaplan, Monica Chelliah, Karl Leswing, Sean Seekins, Shawn Watts, David Calkins, Jackson Chief Elk, Steven V. Jerome, Matthew P. Repasky, and John C. Shelley. Epik: pk _ a and protonation state prediction through machine learning. Journal of ...

  46. [54]

    Development and validation of a genetic algorithm for flexible docking

    Gareth Jones et al. Development and validation of a genetic algorithm for flexible docking. Journal of Molecular Biology, 267 0 (3): 0 727--748, 1997

  47. [55]

    Jorgensen

    William L. Jorgensen. The many roles of computation in drug discovery. Science, 303 0 (5665): 0 1813--1818, 2004. doi:10.1126/science.1096361

  48. [56]

    Karlov, Sergey Sosnin, Maxim V

    Dmitry S. Karlov, Sergey Sosnin, Maxim V. Fedorov, and Petr Popov. graphdelta: Mpnn scoring function for the affinity prediction of protein--ligand complexes. ACS Omega, 5 0 (10): 0 5150--5159, 2020. doi:10.1021/acsomega.9b04162

  49. [57]

    Inference-time scaling for flow models via stochastic generation and rollover budget forcing, 2025

    Jaihoon Kim, Taehoon Yoon, Jisung Hwang, and Minhyuk Sung. Inference-time scaling for flow models via stochastic generation and rollover budget forcing, 2025. URL https://arxiv.org/abs/2503.19385

  50. [58]

    Variational diffusion models

    Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  51. [59]

    Peter A. Kollman, Irina Massova, Carolina Reyes, Bernd Kuhn, Shuanghong Huo, Lillian Chong, Matthew Lee, Taisung Lee, Yong Duan, Wei Wang, Oreola Donini, Piotr Cieplak, Jayaraman Srinivasan, David A. Case, and Thomas E. Cheatham. Calculating structures and free energies of com...

  52. [60]

    A machine learning approach towards the prediction of protein--ligand binding affinity based on fundamental molecular properties

    Indrajit Kundu, Gopal Paul, and Ranjan Banerjee. A machine learning approach towards the prediction of protein--ligand binding affinity based on fundamental molecular properties. RSC Advances, 8 0 (22): 0 12127--12137, 2018. doi:10.1039/C8RA00003D

  53. [61]

    Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation, 2023

    Tuan Le, Julian Cremer, Frank Noé, Djork-Arné Clevert, and Kristof Schütt. Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation, 2023

  54. [62]

    Structural basis for the selective inhibition of Cdc2-Like kinases by CX-4945

    Joo Youn Lee, Ji-Sook Yun, Woo-Keun Kim, Hang-Suk Chun, Hyeonseok Jin, Sungchan Cho, and Jeong Ho Chang. Structural basis for the selective inhibition of Cdc2-Like kinases by CX-4945 . Biomed Res Int, 2019: 0 6125068, August 2019

  55. [63]

    Fragfm: Efficient fragment-based molecular generation via discrete flow matching

    Joongwon Lee, Seonghwan Kim, and Wou Youn Kim. Fragfm: Efficient fragment-based molecular generation via discrete flow matching. arXiv preprint, 2025

  56. [64]

    Shields, Thomas Merth, Punit K

    Pablo Lemos, Zane Beckwith, Sasaank Bandi, Maarten van Damme, Jordan Crivelli-Decker, Benjamin J. Shields, Thomas Merth, Punit K. Jha, Nicola De Mitri, Tiffany J. Callahan, AJ Nish, Paul Abruzzo, Romelia Salomon-Ferrer, and Martin Ganahl. Sair: Enabling deep learning for prote...

  57. [65]

    Levine, Muhammed Shuaibi, Evan Walter Clark Spotte-Smith, Michael G

    Daniel S. Levine, Muhammed Shuaibi, Evan Walter Clark Spotte-Smith, Michael G. Taylor, Muhammad R. Hasyim, Kyle Michel, Ilyes Batatia, Gábor Csányi, Misko Dzamba, Peter Eastman, Nathan C. Frey, Xiang Fu, Vahe Gharakhanyan, Aditi S. Krishnapriyan, Joshua A. Rackers, Sanjeev Raj...

  58. [66]

    Structure-aware interactive graph neural networks for the prediction of protein–ligand binding affinity

    Shuangli Li, Jingbo Zhou, Tong Xu, Liang Huang, Fan Wang, Hui Xiong, Weili Huang, and Dejing Dou. Structure-aware interactive graph neural networks for the prediction of protein–ligand binding affinity. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery an...

  59. [67]

    A high-quality data set of protein--ligand binding interactions via comparative complex structure modeling

    Xuelian Li, Cheng Shen, Hui Zhu, Yujian Yang, Qing Wang, Jincai Yang, and Niu Huang. A high-quality data set of protein--ligand binding interactions via comparative complex structure modeling. Journal of Chemical Information and Modeling, 64 0 (7): 0 2454--2466, Apr 2024. ISSN...

  60. [68]

    Assessing protein--ligand interaction scoring functions with the casf-2013 benchmark

    Yan Li, Minyi Su, Zhihai Liu, Jian Li, Jie Liu, Li Han, and Renxiao Wang. Assessing protein--ligand interaction scoring functions with the casf-2013 benchmark. Nature Protocols, 13 0 (4): 0 666--680, 2018. doi:10.1038/nprot.2017.114

  61. [69]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. ICLR, 2023. arXiv:2210.02747

  62. [70]

    Classification of current scoring functions

    Jie Liu and Renxiao Wang. Classification of current scoring functions. Journal of Chemical Information and Modeling, 55 0 (3): 0 475--482, 2015. doi:10.1021/ci500731a

  63. [71]

    Flow straight and fast: Learning to generate and transfer data with rectified flow

    Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint, 2022. arXiv:2209.03003

  64. [72]

    An autoregressive flow model for 3d molecular geometry generation from scratch

    Youzhi Luo and Shuiwang Ji. An autoregressive flow model for 3d molecular geometry generation from scratch. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=C03Ajc-NS5W

  65. [73]

    Inference-time scaling for diffusion models beyond scaling denoising steps, 2025

    Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, and Saining Xie. Inference-time scaling for diffusion models beyond scaling denoising steps, 2025. URL https://arxiv.org/abs/2501.09732

  66. [74]

    Manning, D

    G. Manning, D. B. Whyte, R. Martinez, T. Hunter, and S. Sudarsanam. The protein kinase complement of the human genome. Science, 298 0 (5600): 0 1912--1934, 2002. doi:10.1126/science.1075762. URL https://www.science.org/doi/abs/10.1126/science.1075762

  67. [75]

    Morris, and Philip C

    Rocco Meli, Garrett M. Morris, and Philip C. Biggin. Scoring functions for protein--ligand binding affinity prediction using structure-based deep learning: A review. Frontiers in Bioinformatics, 2: 0 885983, 2022. doi:10.3389/fbinf.2022.885983

  68. [76]

    Popowicz

    Filipe Menezes and Grzegorz M. Popowicz. Ulysses: An efficient and easy to use semiempirical library for c++. Journal of Chemical Information and Modeling, 62 0 (16): 0 3685--3694, 2022. doi:10.1021/acs.jcim.2c00757. URL https://doi.org/10.1021/acs.jcim.2c00757. PMID: 35930308

  69. [77]

    Antonia S. J. S. Mey et al. Best practices for alchemical free energy calculations. Living Journal of Computational Molecular Science, 2 0 (1): 0 18378, 2020

  70. [78]

    Mobley and Pavel V

    David L. Mobley and Pavel V. Klimovich. Perspective: Alchemical free energy calculations for drug discovery. The Journal of Chemical Physics, 137 0 (23): 0 230901, 2012

  71. [79]

    Accurate protein--ligand binding free energy estimation using qm/mm and the mining minima method

    Farzaneh Molani et al. Accurate protein--ligand binding free energy estimation using qm/mm and the mining minima method. Communications Chemistry, 7 0 (1): 0 49, 2024

  72. [80]

    Quinn, Trung Nguyen, and Svetha Venkatesh

    Thin Nguyen, Thao Minh Le, Thomas P. Quinn, Trung Nguyen, and Svetha Venkatesh. Graphdta: Predicting drug–target binding affinity with graph neural networks. Bioinformatics, 37 0 (8): 0 1140--1147, 2021. doi:10.1093/bioinformatics/btaa921

  73. [81]

    O'Boyle, Michael Banck, Craig A

    Noel M. O'Boyle, Michael Banck, Craig A. James, Chris Morley, Tim Vandermeersch, Pierre Hutchison, Geoffrey R. Del Moral, Arnaud Doucet, and Ajay Jasra. Open babel: An open chemical toolbox. Journal of Cheminformatics, 3 0 (33): 0 1758--2946, 2011. doi:10.1186/1758-2946-3-33

  74. [82]

    O zt \"u rk, Arzucan \

    Hakime \"O zt \"u rk, Arzucan \"O zg \"u r, and E. Ozkirimli. Deepdta: Deep drug–target binding affinity prediction. Bioinformatics, 34 0 (17): 0 i821--i829, 2018. doi:10.1093/bioinformatics/bty593

  75. [83]

    Review and analysis of synthetic dataset generation methods and techniques for application in computer vision

    Goran Paulin and Marina Iva s i \'c -Kos. Review and analysis of synthetic dataset generation methods and techniques for application in computer vision. Artificial Intelligence Review, 56: 0 9221--9265, 2023. doi:10.1007/s10462-022-10358-3

  76. [84]

    Sqm2.20: Semiempirical quantum-mechanical scoring function yields dft-quality protein--ligand binding affinity predictions in minutes

    Adam Pecina, Jind r ich Fanfrl \' k, Martin Lep s \' k, and Jan R ez \'a c . Sqm2.20: Semiempirical quantum-mechanical scoring function yields dft-quality protein--ligand binding affinity predictions in minutes. Nature Communications, 15: 0 1127, 2024

  77. [85]

    Data augmentation techniques in natural language processing

    Lucas Francisco Amaral Orosco Pellicer, Taynan Maier Ferreira, and Anna Helena Reali Costa. Data augmentation techniques in natural language processing. Applied Soft Computing, 132: 0 109803, 2023. doi:10.1016/j.asoc.2022.109803

  78. [86]

    Blum, and Lars Ruddigkeit

    Jean-Louis Reymond, Ruben van Deursen, Lorenz C. Blum, and Lars Ruddigkeit. Chemical space as a source for new drugs. MedChemComm, 1: 0 30--38, 2010. doi:10.1039/C0MD00020E

  79. [87]

    Biggin, and Aniket Magarkar

    Benjamin Ries, Irfan Alibay, Nithishwer Mouroug Anand, Philip C. Biggin, and Aniket Magarkar. Automated absolute binding free energy calculation workflow for drug discovery. Journal of Chemical Information and Modeling, 64 0 (14): 0 5357--5364, Jul 2024. ISSN 1549-9596. doi:10...

  80. [88]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. arXiv:2112.10752

  81. [90]

    Ross, Chao Lu, Guido Scarabelli, and Lingle Wang

    Gregory A. Ross, Chao Lu, Guido Scarabelli, and Lingle Wang. The maximal and current accuracy of rigorous protein--ligand binding free energy calculations. Communications Chemistry, 6: 0 222, 2023 b . doi:10.1038/s42004-023-01019-9

  82. [91]

    Madhavi Sastry, Matvey Adzhigirey, Tyler Day, Ramakrishna Annabhimoju, and Woody Sherman

    G. Madhavi Sastry, Matvey Adzhigirey, Tyler Day, Ramakrishna Annabhimoju, and Woody Sherman. Protein and ligand preparation: parameters, protocols, and influence on virtual screening enrichments. Journal of Computer-Aided Molecular Design, 27 0 (3): 0 221--234, 2013. doi:10.10...

  83. [92]

    Hadfield, Oliver M

    Jack Scantlebury, Lucy Vost, Adam Carbery, Tristan E. Hadfield, Oliver M. Turnbull, Nathan Brown, Vijil Chenthamarakshan, Payel Das, Hugues Grosjean, Frank von Delft, and Charlotte M. Deane. A small step toward generalizability: Training a machine learning scoring function for...

  84. [93]

    Structure-based drug design with equivariant diffusion models, 2023

    Arne Schneuing, Yuanqi Du, Charles Harris, Arian Jamasb, Ilia Igashov, Weitao Du, Tom Blundell, Pietro Lió, Carla Gomes, Max Welling, Michael Bronstein, and Bruno Correia. Structure-based drug design with equivariant diffusion models, 2023. URL https://arxiv.org/abs/2210.13695

  85. [94]

    Dobbelstein, Thomas Castiglione, Michael M

    Arne Schneuing, Ilia Igashov, Adrian W. Dobbelstein, Thomas Castiglione, Michael M. Bronstein, and Bruno Correia. Multi-domain distribution learning for de novo drug design. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.ne...

  86. [95]

    Schütt, Oliver T

    Kristof T. Schütt, Oliver T. Unke, and Michael Gastegger. Equivariant message passing for the prediction of tensorial properties and molecular spectra, 2021. URL https://arxiv.org/abs/2102.03150

  87. [96]

    Rewarding progress: Scaling automated process verifiers for llm reasoning, 2024

    Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, and Aviral Kumar. Rewarding progress: Scaling automated process verifiers for llm reasoning, 2024. URL https://arxiv.org/abs/2410.08146

  88. [97]

    A general framework for inference-time scaling and steering of diffusion models

    Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Zhou Yu, Kathleen McKeown, and Rajesh Ranganath. A general framework for inference-time scaling and steering of diffusion models. In Forty-second International Conference on Machine Learning, 2025. URL https://openrevie...

  89. [98]

    Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024

    Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024. URL https://arxiv.org/abs/2408.03314

  90. [99]

    Deep unsupervised learning using nonequilibrium thermodynamics

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning, pages 2256--2265. PMLR, 2015

  91. [101]

    Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole

    Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations, 2021 b . URL https://arxiv.org/abs/2011.13456

  92. [102]

    Clark Still, Anna Tempczyk, Ronald C

    W. Clark Still, Anna Tempczyk, Ronald C. Hawley, and Thomas Hendrickson. Semianalytical treatment of solvation for molecular mechanics and dynamics. Journal of the American Chemical Society, 112 0 (16): 0 6127--6129, 1990. doi:10.1021/ja00172a038

  93. [103]

    Comparative assessment of scoring functions: The casf-2016 update

    Minyi Su, Qifan Yang, Yun Du, Guoli Feng, Zhihai Liu, Yan Li, and Renxiao Wang. Comparative assessment of scoring functions: The casf-2016 update. Journal of Chemical Information and Modeling, 59 0 (2): 0 895--913, 2019. doi:10.1021/acs.jcim.8b00545

  94. [104]

    Drug–target interaction prediction with the kiba dataset

    Jian Tang, Andrzej Szwajda, Shweta Shakyawar, Tianyun Xu, Pekka Hintsanen, Krister Wennerberg, and Tero Aittokallio. Drug–target interaction prediction with the kiba dataset. PLoS Computational Biology, 10 0 (12): 0 e1003765, 2014 a . doi:10.1371/journal.pcbi.1003765

  95. [105]

    Making sense of large-scale kinase inhibitor bioactivity data sets: A comparative and integrative analysis

    Jian Tang, Andrzej Szwajda, Sumeet Shakyawar, Tianyun Xu, Perttu Hintsanen, Krister Wennerberg, and Tero Aittokallio. Making sense of large-scale kinase inhibitor bioactivity data sets: A comparative and integrative analysis. Journal of Chemical Information and Modeling, 54 0 ...

  96. [106]

    Improving and generalizing flow-based generative models with minibatch optimal transport

    Alexander Tong, Kilian FATRAS, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, and Yoshua Bengio. Improving and generalizing flow-based generative models with minibatch optimal transport. Transactions on Machine Learning Research, 2024. ISSN 283...

  97. [107]

    Mathis, and Pietro Li \`o

    Jos Torge, Charles Harris, Simon V. Mathis, and Pietro Li \`o . Diffhopp: A graph diffusion model for novel drug design via scaffold hopping. In ICML Workshop on Computational Biology, 2023

  98. [108]

    Oleg Trott and Arthur J. Olson. Autodock vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. Journal of Computational Chemistry, 31 0 (2): 0 455--461, 2010

  99. [109]

    Warren, Charlotte M

    \'I sak Valsson, Matthew T. Warren, Charlotte M. Deane, Aniket Magarkar, Garrett M. Morris, and Philip C. Biggin. Narrowing the gap between machine learning scoring functions and free energy perturbation using augmented data. Communications Chemistry, 8 0 (1): 0 41, Feb 2025. ...

  100. [110]

    Midi: Mixed graph and 3d denoising diffusion for molecule generation

    Cl \' e ment Vignac, Nagham Osman, Laura Toni, and Pascal Frossard. Midi: Mixed graph and 3d denoising diffusion for molecule generation. In Machine Learning and Knowledge Discovery in Databases: Research Track - European Conference, ECML PKDD 2023, Turin, Italy, September 18-...

  101. [111]

    Turk, Nicolas Drizard, Nicolas Martin, Bernd Hoffmann, Martine Stahl, Nikolay P

    Mikhail Volkov, James A. Turk, Nicolas Drizard, Nicolas Martin, Bernd Hoffmann, Martine Stahl, Nikolay P. Todorov, Gianni De Fabritiis, John D. Chodera, Gilles Marcou, and Didier Rognan. On the frustration to predict binding affinities from protein--ligand structures with deep...

  102. [112]

    A review on fragment-based de novo 2d molecule generation

    Sergei Voloboev. A review on fragment-based de novo 2d molecule generation. arXiv preprint, 2024

  103. [113]

    Molcraft: Structure-based drug design in continuous parameter space

    Chenyang Wang, Shiqi Xu, Yiheng Wang, et al. Molcraft: Structure-based drug design in continuous parameter space. arXiv preprint, 2024. arXiv:2404.12141

  104. [114]

    Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field

    Lingle Wang et al. Accurate and reliable prediction of relative ligand binding potency in prospective drug discovery by way of a modern free-energy calculation protocol and force field. Journal of the American Chemical Society, 137 0 (7): 0 2695--2703, 2015. doi:10.1021/ja512751q

  105. [115]

    The pdbbind database: methodologies and updates

    Renxiao Wang, Xueliang Fang, Yipin Lu, Chao-Yie Yang, and Shaomeng Wang. The pdbbind database: methodologies and updates. Journal of medicinal chemistry, 48 0 (12): 0 4111--4119, 2005

  106. [116]

    Carlson, and Teresa Head-Gordon

    Yingze Wang, Kunyang Sun, Jie Li, Xingyi Guan, Oufan Zhang, Dorian Bagni, Yang Zhang, Heather A. Carlson, and Teresa Head-Gordon. A workflow to create a high-quality protein–ligand binding dataset for training , validation , and prediction tasks. Digital Discovery, 4: 0 1209--...

  107. [117]

    Wohlwend et al

    J. Wohlwend et al. Boltz-2: Towards accurate and efficient binding affinity prediction. bioRxiv preprint, 2025

  108. [118]

    Boltz-1 democratizing biomolecular interaction modeling

    Jeremy Wohlwend, Gabriele Corso, Saro Passaro, Mateo Reveiz, Ken Leidal, Wojtek Swiderski, Tally Portnoi, Itamar Chinn, Jacob Silterra, Tommi Jaakkola, and Regina Barzilay. Boltz-1 democratizing biomolecular interaction modeling. bioRxiv, 2024. doi:10.1101/2024.11.19.624167. U...

  109. [119]

    Stepniewska-Dziubi \'n ska, and Pawe Siedlecki

    Maciej W \'o jcikowski, Micha Kukie ka, Marta M. Stepniewska-Dziubi \'n ska, and Pawe Siedlecki. Development of a protein--ligand extended connectivity (plec) fingerprint and its application for binding affinity predictions. Bioinformatics, 35 0 (8): 0 1334--1341, 2019. doi:10...

  110. [120]

    Trippe, Christian A Naesseth, John Patrick Cunningham, and David Blei

    Luhuan Wu, Brian L. Trippe, Christian A Naesseth, John Patrick Cunningham, and David Blei. Practical and asymptotically exact conditional sampling in diffusion models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum...

  111. [121]

    A folding-docking-affinity framework for protein-ligand binding affinity prediction

    Ming-Hsiu Wu, Ziqian Xie, and Degui Zhi. A folding-docking-affinity framework for protein-ligand binding affinity prediction. Communications Chemistry, 8 0 (1): 0 108, Apr 2025. ISSN 2399-3669. doi:10.1038/s42004-025-01506-1. URL https://doi.org/10.1038/s42004-025-01506-1

  112. [122]

    Geodiff: A geometric diffusion model for molecular conformation generation

    Minkai Xu, Shitong Luo, Yoshua Bengio, and et al. Geodiff: A geometric diffusion model for molecular conformation generation. In ICLR, 2022. arXiv:2203.02923

  113. [123]

    Predicting or pretending: Artificial intelligence for protein--ligand interactions lack of sufficiently large and unbiased datasets

    Jincai Yang, Cheng Shen, and Niu Huang. Predicting or pretending: Artificial intelligence for protein--ligand interactions lack of sufficiently large and unbiased datasets. Frontiers in Pharmacology, 11: 0 69, 2020 a . doi:10.3389/fphar.2020.00069

  114. [124]

    Repasky, Karl Leswing, Robert Abel, Brian K

    Yujia Yang, Kyle Yao, Matthew P. Repasky, Karl Leswing, Robert Abel, Brian K. Shoichet, and Spencer V. Jerome. Efficient exploration of chemical space with docking and deep learning. Journal of Chemical Theory and Computation, 17 0 (11): 0 7106--7119, 2021. doi:10.1021/acs.jct...

  115. [125]

    Syntalinker: automatic fragment linking with deep conditional transformer neural networks

    Yuyao Yang, Shuangjia Zheng, Shimin Su, Chao Zhao, Jun Xu, and Hongming Chen. Syntalinker: automatic fragment linking with deep conditional transformer neural networks. Chem. Sci., 11: 0 8312--8322, 2020 b . doi:10.1039/D0SC03126G. URL http://dx.doi.org/10.1039/D0SC03126G

  116. [126]

    Turbohopp: Accelerated molecule scaffold hopping with consistency models

    Kiwoong Yoo, Owen Oertell, Junhyun Lee, Sanghoon Lee, and Jaewoo Kang. Turbohopp: Accelerated molecule scaffold hopping with consistency models. In NeurIPS, 2024

  117. [127]

    Fraggen: towards 3d geometry reliable fragment-based molecular generation

    Odin Zhang, Yufei Huang, Shichen Cheng, Mengyao Yu, Xujun Zhang, Haitao Lin, Yundian Zeng, Mingyang Wang, Zhenxing Wu, Huifeng Zhao, Zaixi Zhang, Chenqing Hua, Yu Kang, Sunliang Cui, Peichen Pan, Chang-Yu Hsieh, and Tingjun Hou. Fraggen: towards 3d geometry reliable fragment-b...

  118. [128]

    Augmented bindingnet dataset for enhanced ligand binding pose predictions using deep learning

    Hui Zhu, Xuelian Li, Baoquan Chen, and Niu Huang. Augmented bindingnet dataset for enhanced ligand binding pose predictions using deep learning. npj Drug Discovery, 2 0 (1): 0 1, Jan 2025. ISSN 3005-1452. doi:10.1038/s44386-024-00003-0. URL https://doi.org/10.1038/s44386-024-00003-0

  119. [129]

    Robert W. Zwanzig. High-temperature equation of state by a perturbation method. i. nonpolar gases. The Journal of Chemical Physics, 22 0 (8): 0 1420--1426, 1954. doi:10.1063/1.1740409

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.