Pith. sign in

REVIEW 5 major objections 4 minor 92 references

Unraveling the Potential of Diffusion Models in Small Molecule Generation

T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This review organizes diffusion-based small-molecule generation with a new taxonomy and reports a benchmark in which MiDi leads unconditional generation, KGDiff and PMDM lead target-aware generation, and all models still need geometric…

desk verdict A solid survey of diffusion models for small molecule generation, but the benchmark section's headline conclusion does not follow from the reported table. read the letter →

arxiv 2507.08005 v1 pith:YUFO4CZW submitted 2025-06-25 q-bio.BM cs.AIcs.LG

classification q-bio.BMcs.AIcs.LG
keywords diffusionmodelssmallmoleculegenerationdenovodrugdesign3Dmolecularstructuresequivarianttarget-awarebenchmarkingvalidity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a systematic review that tries to establish how diffusion models should be categorized for small-molecule generation and which of the existing methods actually work. It proposes a five-level taxonomy that separates target-free from target-aware generation, then splits by task, representation, diffusion formulation, and equivariance. To back the taxonomy with evidence, the authors run their own benchmark of eighteen models across QM9, GEOM-Drugs, and CrossDocked2020. Their empirical conclusion is that MiDi performs best on stability and validity for unconditional generation, while KGDiff and PMDM achieve the best target-aware performance, yet every tested model still depends on post-hoc geometric relaxation to reach acceptable validity. A reader should care because the paper converts a crowded literature into a structured map and states, in concrete metric terms, where the field stands.

What carries the argument

The organizing device is a five-level taxonomy that classifies each diffusion model by target specificity, generation configuration, molecular modality, diffusion formulation (DDPM versus score-based), and symmetry treatment (SE(3)-equivariant, permutation-equivariant, or neither). The evidentiary mechanism is the benchmark protocol: 1000 molecules per model across five protein targets from CrossDocked2020, evaluated with GenBench3D metrics before and after MMFF relaxation. The diffusion process itself—forward noise addition and learned reverse denoising over atom coordinates and categorical features—is the underlying generative machinery all surveyed models share.

What would settle it

Recompute Table 2 from the generated samples with explicitly labeled and validated QED, SA, Vina, and Validity3D columns for the same five proteins; if any relaxed QED value exceeds 1.0 or the column alignment is off, the reported superiority of KGDiff and PMDM over the other target-aware models collapses and the benchmark must be rerun.

Watch

Extended reading notes

Core claim

The central claim, stated in the conclusion, is a comparative verdict: in the authors' evaluation, MiDi demonstrates superior performance in stability and validity for unconditional generation, while KGDiff and PMDM achieve the best target-aware performance, and all models still require geometric correction. The paper supports this with its proposed taxonomy, organizing models by target specificity, generation configuration, molecular modality, diffusion formulation, and symmetry properties, and with a benchmark that reports validity, uniqueness, stability, novelty, QED, synthetic accessibility, Vina score, and strain energy. The same benchmark shows that Merck Molecular Force Field relaxation raises mean geometric validity from 1.42 percent to 37.0 percent, which the paper reads as evidence that current training objectives do not enforce sufficient chemical constraints.

Load-bearing premise

The benchmark verdict rests on the correctness of Table 2, and that table contains an apparent error: relaxed QED values are reported above 1.0, which a 0-to-1 drug-likeness score cannot reach, so the model rankings are only as reliable as the table's column alignment.

Editorial extensions

If this is right

  • If MiDi's ranking holds, unconditional 3D generation should move toward joint graph-and-geometry diffusion rather than coordinate-only or voxel approaches.
  • If KGDiff and PMDM are the target-aware leaders, then binding-affinity guidance and dual-encoder pocket conditioning are the productive directions to build on.
  • The jump in geometric validity from 1.42% to 37.0% after MMFF relaxation implies that training objectives currently lack chemical constraints; physics-informed diffusion is the natural next step.
  • Fragment-based models such as AutoFragDiff and DecompOPT score poorly across all metrics and should be redesigned rather than incrementally tuned.
  • Because every model needs geometric correction, 3D diffusion generation is not yet ready to feed directly into virtual-screening pipelines.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Our reading: Table 2 reports relaxed QED values above 1.0, which is impossible for a 0-to-1 drug-likeness score; if the columns are misaligned, the claimed model rankings are not yet supported.
  • Our reading: five protein targets and one 1000-molecule sample per model is a thin basis for ranking models; replicating the benchmark on more targets and multiple seeds could change the ordering.
  • Our reading: the paper's conclusion implies that any future benchmark should report post-relaxation validity and energy-based metrics, not just raw generation scores, before claiming a model is usable.
  • Our reading: the taxonomy's target-free/target-aware split could also serve as a checklist for choosing a model, but cross-dataset comparisons remain unreliable because 1D, 2D, and 3D models are trained on different benchmarks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. This manuscript is a review of diffusion models for small-molecule generation, organized around a proposed taxonomy (target-free vs target-aware, conformation vs de novo generation, molecular representation, diffusion formulation, and equivariance/network architecture), and it includes a new benchmark of nine target-free and nine target-aware models. Target-free models are evaluated on QM9 and GEOM-Drugs; target-aware models are evaluated on five CrossDocked2020 protein targets using GenBench3D metrics with and without MMFF relaxation. The paper's concluding empirical claim (Section 8) is that MiDi is superior for unconditional generation, KGDiff and PMDM are the best target-aware models, and all models still require geometric correction.

Significance. A current, structured survey of this fast-moving field is valuable, and the taxonomy in Table 1 plus the side-by-side benchmark table will be a useful entry point for practitioners. The inclusion of MMFF relaxation and the observed Validity3D/Vina trade-off are useful observations. However, the empirical ranking claims are under-supported as presented: the paper does not state a multi-metric aggregation rule, does not report per-target variance or statistical significance, and contains at least one statement about fragment-based models that is contradicted by its own table. These issues are fixable but are load-bearing for the paper's empirical contribution. The review portion is generally informative, though it contains a technical inaccuracy about classifier-free guidance and several citation errors.

major comments (5)
  1. [Section 8 / Table 2] The conclusion that 'KGDiff and PMDM achieved the best target-aware performance' is not derivable from Table 2 because no metric-aggregation rule is stated. Section 5 names only KGDiff as best, and Table 2 shows that PMDM has the lowest relaxed Validity3D (2.9%) and the worst SA Score (7.285) among the nine models, while KGDiff has the best relaxed Vina (-5.557) but middle-ranked validity (42.9%). A mean-rank, Pareto-dominance, or explicit weighting rule is needed before any co-best statement can be made.
  2. [Section 5, Benchmarking Representative DMs] The benchmark compares averages of per-protein medians over only five protein targets but provides no per-protein breakdowns, error bars, confidence intervals, or significance tests. With n=5, a five-percentage-point difference in relaxed Validity3D or a ~1.2 kcal/mol difference in Vina (e.g., KGDiff -5.557 vs DecompDiff -4.337) may be within inter-target noise. The authors should report per-target results with variance or downgrade the ranking language to descriptive observations; releasing the evaluation code and data would also allow readers to assess the ranking.
  3. [Section 5 / Table 2, fragment-based models] The statement that fragment-based approaches 'performed poorly across all metrics' is contradicted by Table 2: DecompOPT has a relaxed Vina score of -4.834, which is better than MolSnapper (-2.728), DiffSBDD (-4.324), and AutoFragDiff (49.145), and its relaxed strain energy of 42.4 is not the worst among the nine models. This claim should be qualified by metric and by model.
  4. [Section 5 / Table 2, target-free rows] The text says 'MiDi has the best performance' and Section 8 says MiDi is superior in stability and validity, but Table 2 shows that HierDiff has a higher QM9 molecule stability (100 vs 97.5) and VoxMol has a higher QM9 validity (98.6 vs 97.9), while MiDi's GEOM-Drugs validity (77.8) is well below VoxMol (94.7) and MDM (99.5). If MiDi is preferred by a composite criterion, that criterion should be stated explicitly.
  5. [Section 4.2, Conditioning Methods] The description of classifier-free guidance is technically incorrect. The formula p_theta(x_L^t | x_L^{t-1}, x_P) proportional to p_theta(x_L^t | x_L^{t-1}) * p_phi(x_P | x_L^t) describes classifier guidance with an auxiliary model p_phi. Classifier-free guidance uses a single conditional model and interpolates the conditional and unconditional score estimates, without an auxiliary classifier p_phi. Please correct this formula and the surrounding text.
minor comments (4)
  1. [Table 2 header] The header layout makes it easy to misread QED and SA as having Normal/Relaxed subcolumns. I agree with the stress-test that the QED values (0.266-0.628) are plausible single-column values and that no impossible relaxed-QED value appears; nevertheless, the table should be reformatted with visually grouped subheaders so that only Validity3D, Vina, and Strain have Normal/Relaxed columns.
  2. [Sections 1, 2, and 5] There are unresolved cross-references: Section 1 and Section 2 refer to 'Figure ??' for the taxonomy, and Section 5 refers to 'Figure ??' for the Vina-score comparison, while a Figure 2 caption is present but not connected to the text. These cross-references should be resolved.
  3. [References in Section 4.2] Several citations are incorrect: 'AlphaFold31' in Section 4.2 should cite [1] (AlphaFold3) rather than [31] (JODO); 'InterDiff,36' should cite [79] (InterDiff) rather than [36] (IPDiff); and the metric definitions attributed to 'Pocket2Mol.21' point to a Pocket PC molecular visualization tool rather than to the Pocket2Mol molecular sampling paper.
  4. [Section 5, execution-time discussion] The claim that voxel-based approaches VoxMol and FMG 'require significantly more training time compared to other methods' is not supported by Figure 3, which reports only inference/batch execution time; either remove the training-time claim or back it with a citation or data.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark conclusions rest on externally defined metrics and an original evaluation, not on fitted inputs or self-citation.

full rationale

This paper is a literature review with an original benchmark, not a derivation of a model or theory. It introduces no model, fits no free parameters, and makes no prediction that is defined in terms of its own conclusion. The reported metrics come from pre-existing, externally defined tools and datasets: QM9, GEOM-Drugs, CrossDocked2020, Genbench3D, AutoDock Vina, RDKit, and MMFF relaxation, and the only load-bearing relation between Table 2 and the Section 8 conclusion is that the conclusion verbally restates the table. Restating one's own empirical results is not circular. The reader's relaxed-QED objection misreads Table 2: QED and SA are single columns with values in the plausible ranges 0.266-0.628 and 4.223-7.285, while Normal/Relaxed subcolumns appear only under Validity3D, Vina Score, and Strain Energy. The KGDiff/PMDM 'best target-aware performance' claim is under-determined by the stated evidence (Section 5 names only KGDiff, and PMDM has the lowest relaxed Validity3D at 2.9%), but under-determination is a correctness and rigor concern, not circularity. No self-citations appear in the reference list; the citations used for the diffusion equations and equivariance concepts are external foundational works. There is consequently no fitted input renamed as a prediction, no self-citation chain, no imported uniqueness theorem, and no ansatz smuggled in via citation to the authors' own work.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities; it is a survey and benchmark. The claims rest on domain assumptions about equivariance and the validity of evaluation metrics, one of which (metric correctness) is contradicted by the table.

assumptions (3)
  • domain assumption SE(3) equivariance is a beneficial inductive bias for 3D molecule generation.
    The paper states this repeatedly (Section 3.2) and uses it to categorize models, but it is an inductive bias assumption, not a proven requirement for the surveyed results.
  • domain assumption Established benchmark metrics such as QED, SA Score, Vina, and Validity3D are correctly implemented in the reported evaluation.
    The empirical conclusions in Section 5 rely on these metrics; the impossible relaxed QED values in Table 2 show this assumption is violated, at least in the reporting.
  • domain assumption MMFF relaxation is an appropriate post-processing step for evaluating generated molecular geometries.
    The paper uses relaxation in Section 5 to compute relaxed Validity3D and Vina scores, and bases model rankings on relaxed values, assuming this correction is chemically sound.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Unraveling the Potential of Diffusion Models in Small Molecule Generation." pith.science (2026). https://pith.science/paper/YUFO4CZW

@misc{pith2026250708005,
  author       = {Pith},
  title        = {Pith review of: Unraveling the Potential of Diffusion Models in Small Molecule Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YUFO4CZW}},
  note         = {Machine review of arXiv:2507.08005}
}
read the original abstract

Generative AI presents chemists with novel ideas for drug design and facilitates the exploration of vast chemical spaces. Diffusion models (DMs), an emerging tool, have recently attracted great attention in drug R\&D. This paper comprehensively reviews the latest advancements and applications of DMs in molecular generation. It begins by introducing the theoretical principles of DMs. Subsequently, it categorizes various DM-based molecular generation methods according to their mathematical and chemical applications. The review further examines the performance of these models on benchmark datasets, with a particular focus on comparing the generation performance of existing 3D methods. Finally, it concludes by emphasizing current challenges and suggesting future research directions to fully exploit the potential of DMs in drug discovery.

Figures

Figures reproduced from arXiv: 2507.08005 by the authors.

Figure 1
Figure 1. Left: 1D/2D/3D molecular representations; Right: the diffusion model generative process, consists of steps of forward diffusion process (from left to right) and backward diffusion process (from right to left), described by Gaussian process (top), and how DM generates a new 3D molecule (middle) and generates 3D molecule based on protein target (bottom). molecular geometry generation methods that rely on rule￾based sy… view at source ↗
Figure 2
Figure 2. Comparison of Vina scores between raw and relaxed molecules. A decreased score indicates relaxation improves molecular validity (lower scores are better), while an increased score suggests the model overfits to Vina scores. 5. Benchmarking Representative DMs We conducted extensive experiments to evaluate eigh￾teen representative DMs, nine for target-free generation and nine for target-aware generation across three b… view at source ↗
Figure 3
Figure 3. Star plot showing the averaged execution time (in seconds) for a batch of 10 molecules across five different proteins. For each protein and each model, we generate 10 batches of 10 molecules and take the average execution time. fosters trust and validation, which is vital for advancing their use in drug discovery. Furthermore, the gap between theoretical molecule gen￾eration and the practical application of pharmace… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

92 extracted references · 75 canonical work pages

  1. [1]

    Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A.J., Bambrick, J., et al.,

  2. [2]

    The process of structure-based drug design

    Anderson, A.C., 2003. The process of structure-based drug design. Chemistry & biology 10, 787–797

  3. [3]

    Geom, energy-annotated molecular conformations for property prediction and molecular gen- eration

    Axelrod, S., Gomez-Bombarelli, R., 2022. Geom, energy-annotated molecular conformations for property prediction and molecular gen- eration. Scientific Data 9, 185

  4. [4]

    Benchmarking structure-based three-dimensional molecular generative models using GenBench3D: ligand conformation quality matters

    Baillif, B., Cole, J., McCabe, P., Bender, A., 2024. Benchmarking structure-basedthree-dimensionalmoleculargenerativemodelsusing genbench3d: ligand conformation quality matters. URL: https:// arxiv.org/abs/2407.04424, arXiv:2407.04424

  5. [5]

    Quantifying the chemical beauty of drugs

    Bickerton, G.R., Paolini, G.V., Besnard, J., Muresan, S., Hopkins, A.L., 2012. Quantifying the chemical beauty of drugs. Nature chemistry 4, 90–98

  6. [6]

    Gua- camol: benchmarking models for de novo molecular design

    Brown, N., Fiscato, M., Segler, M.H., Vaucher, A.C., 2019. Gua- camol: benchmarking models for de novo molecular design. Journal of chemical information and modeling 59, 1096–1108

  7. [7]

    Analog bits: Generating discrete data using diffusion models with self-conditioning, in: The Eleventh International Conference on Learning Representations

    Chen, T., ZHANG, R., Hinton, G., 2023a. Analog bits: Generating discrete data using diffusion models with self-conditioning, in: The Eleventh International Conference on Learning Representations

  8. [8]

    Shape- conditioned3dmoleculegenerationviaequivariantdiffusionmodels

    Chen, Z., Peng, B., Parthasarathy, S., Ning, X., 2023b. Shape- conditioned3dmoleculegenerationviaequivariantdiffusionmodels. CoRR

Show all 92 references
  1. [9]

    Ilvr: Conditioning method for denoising diffusion probabilistic models, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE

    Choi, J., Kim, S., Jeong, Y., Gwon, Y., Yoon, S., 2021. Ilvr: Conditioning method for denoising diffusion probabilistic models, in: 2021 IEEE/CVF International Conference on Computer Vision (ICCV), IEEE. pp. 14347–14356

  2. [10]

    Diffdock:Diffusionsteps,twists,andturnsformoleculardocking,in: International Conference on Learning Representations (ICLR 2023)

    Corso, G., StÃ, H., Jing, B., Barzilay, R., Jaakkola, T., et al., 2023. Diffdock:Diffusionsteps,twists,andturnsformoleculardocking,in: International Conference on Learning Representations (ICLR 2023)

  3. [11]

    Molgan:Animplicitgenerativemodelfor small molecular graphs

    DeCao,N.,Kipf,T.,2018. Molgan:Animplicitgenerativemodelfor small molecular graphs. arXiv preprint arXiv:1805.11973

  4. [12]

    Diffusion models beat gans on image synthesis

    Dhariwal, P., Nichol, A., 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34, 8780–8794

  5. [13]

    Field-based molecule generation

    Dumitrescu, A., Korpela, D., Heinonen, M., Verma, Y., Iakovlev, V., Garg, V., Lähdesmäki, H., 2024. Field-based molecule generation. arXiv preprint arXiv:2402.15864 . P. Zhang et al.:Preprint submitted to Elsevier Page 11 of 13 Diffusion-based Small Molecule Generation

  6. [14]

    Autodock vina 1.2

    Eberhardt, J., Santos-Martins, D., Tillack, A.F., Forli, S., 2021. Autodock vina 1.2. 0: New docking methods, expanded force field, and python bindings. Journal of chemical information and modeling 61, 3891–3898

  7. [15]

    Translationbetweenmoleculesandnaturallanguage,in:Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp

    Edwards, C., Lai, T., Ros, K., Honke, G., Cho, K., Ji, H., 2022. Translationbetweenmoleculesandnaturallanguage,in:Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 375–413

  8. [16]

    Estimationofsyntheticaccessibility score of drug-like molecules based on molecular complexity and fragment contributions

    Ertl,P.,Schuffenhauer,A.,2009. Estimationofsyntheticaccessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics 1, 1–11

  9. [17]

    Three-dimensionalconvolutionalneural networksandacross-dockeddatasetforstructure-baseddrugdesign

    Francoeur, P.G., Masuda, T., Sunseri, J., Jia, A., Iovanisci, R.B., Snyder,I.,Koes,D.R.,2020. Three-dimensionalconvolutionalneural networksandacross-dockeddatasetforstructure-baseddrugdesign. Journal of chemical information and modeling 60, 4200–4215

  10. [18]

    Se (3)- transformers: 3d roto-translation equivariant attention networks

    Fuchs, F., Worrall, D., Fischer, V., Welling, M., 2020. Se (3)- transformers: 3d roto-translation equivariant attention networks. Ad- vances in neural information processing systems 33, 1970–1981

  11. [19]

    E (n) equivariant normalizing flows

    Garcia Satorras, V., Hoogeboom, E., Fuchs, F., Posner, I., Welling, M., 2021. E (n) equivariant normalizing flows. Advances in Neural Information Processing Systems 34, 4181–4192

  12. [20]

    Autoregressive fragment-baseddiffusionforpocket-awareliganddesign,in:NeurIPS 2023 Generative AI and Biology (GenBio) Workshop

    Ghorbani, M., Gendelev, L., Beroza, P., Keiser, M., . Autoregressive fragment-baseddiffusionforpocket-awareliganddesign,in:NeurIPS 2023 Generative AI and Biology (GenBio) Workshop

  13. [21]

    Pocketmol: a molecular visualization tool for the pocket pc, in: Proceedings 2nd Annual IEEE International Symposium on Bioinformatics and Bioengineer- ing (BIBE 2001), IEEE

    Gilder, J.R., Raymer, M., Doom, T., 2001. Pocketmol: a molecular visualization tool for the pocket pc, in: Proceedings 2nd Annual IEEE International Symposium on Bioinformatics and Bioengineer- ing (BIBE 2001), IEEE. pp. 11–14

  14. [22]

    Text-guided molecule generation with diffusion language model, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

    Gong, H., Liu, Q., Wu, S., Wang, L., 2024. Text-guided molecule generation with diffusion language model, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 109–117

  15. [23]

    Aligning target-aware molecule diffusionmodelswithexactenergyoptimization

    Gu,S.,Xu,M.,Powers,A.,Nie,W.,Geffner,T.,Kreis,K.,Leskovec, J., Vahdat, A., Ermon, S., 2024. Aligning target-aware molecule diffusionmodelswithexactenergyoptimization. AdvancesinNeural Information Processing Systems 37, 44040–44063

  16. [24]

    Linkernet: Fragment poses and linker co-design with 3d equivariant diffusion

    Guan,J.,Peng,X.,Jiang,P.,Luo,Y.,Peng,J.,Ma,J.,2024. Linkernet: Fragment poses and linker co-design with 3d equivariant diffusion. Advances in Neural Information Processing Systems 36

  17. [25]

    3d equivariantdiffusionfortarget-awaremoleculegenerationandaffinity prediction, in: The Eleventh International Conference on Learning Representations

    Guan, J., Qian, W.W., Peng, X., Su, Y., Peng, J., Ma, J., 2023a. 3d equivariantdiffusionfortarget-awaremoleculegenerationandaffinity prediction, in: The Eleventh International Conference on Learning Representations

  18. [26]

    Decompdiff: diffusion models with decomposed priors for structure-based drug design, in: Proceedings of the 40th International Conference on Machine Learning, pp

    Guan,J.,Zhou,X.,Yang,Y.,Bao,Y.,Peng,J.,Ma,J.,Liu,Q.,Wang, L., Gu, Q., 2023b. Decompdiff: diffusion models with decomposed priors for structure-based drug design, in: Proceedings of the 40th International Conference on Machine Learning, pp. 11827–11846

  19. [27]

    Denoising diffusion probabilistic models

    Ho, J., Jain, A., Abbeel, P., 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33, 6840–6851

  20. [28]

    Classifier-free diffusion guidance, in: NeurIPS 2021WorkshoponDeepGenerativeModelsandDownstreamAppli- cations

    Ho, J., Salimans, T., . Classifier-free diffusion guidance, in: NeurIPS 2021WorkshoponDeepGenerativeModelsandDownstreamAppli- cations

  21. [29]

    Equivariant diffusion for molecule generation in 3d, in: International conference on machine learning, PMLR

    Hoogeboom, E., Satorras, V.G., Vignac, C., Welling, M., 2022. Equivariant diffusion for molecule generation in 3d, in: International conference on machine learning, PMLR. pp. 8867–8887

  22. [30]

    Hua, C., Luan, S., Xu, M., Ying, Z., Fu, J., Ermon, S., Precup, D.,

  23. [31]

    Learning joint 2-d and 3-d graph diffusion models for complete molecule generation

    Huang, H., Sun, L., Du, B., Lv, W., 2024a. Learning joint 2-d and 3-d graph diffusion models for complete molecule generation. IEEE Transactions on Neural Networks and Learning Systems

  24. [32]

    Mudiff: Unified diffusion for complete molecule generation, in: Learning on Graphs Conference, PMLR. pp. 33–1

  25. [33]

    Mdm: Molecular diffusion model for 3d molecule generation, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

    Huang, L., Zhang, H., Xu, T., Wong, K.C., 2023a. Mdm: Molecular diffusion model for 3d molecule generation, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 5105–5112

  26. [34]

    Adualdiffusionmodelenables 3dmoleculegenerationandleadoptimizationbasedontargetpockets

    Huang, L., Xu, T., Yu, Y., Zhao, P., Chen, X., Han, J., Xie, Z., Li, H., Zhong,W.,Wong,K.C.,etal.,2024b. Adualdiffusionmodelenables 3dmoleculegenerationandleadoptimizationbasedontargetpockets. Nature Communications 15, 2657

  27. [35]

    Binding-adaptivediffusionmodelsfor structure-baseddrugdesign,in:ProceedingsoftheAAAIConference on Artificial Intelligence, pp

    Huang, Z., Yang, L., Zhang, Z., Zhou, X., Bao, Y., Zheng, X., Yang, Y.,Wang,Y.,Yang,W.,2024d. Binding-adaptivediffusionmodelsfor structure-baseddrugdesign,in:ProceedingsoftheAAAIConference on Artificial Intelligence, pp. 12671–12679

  28. [36]

    Re-dock: Towards flexible and realistic molecular dock- ing with diffusion bridge, in: International Conference on Machine Learning, PMLR

    Huang, Y., Zhang, O., Wu, L., Tan, C., Lin, H., Gao, Z., Li, S., Li, S.Z., 2024c. Re-dock: Towards flexible and realistic molecular dock- ing with diffusion bridge, in: International Conference on Machine Learning, PMLR. pp. 20474–20489

  29. [37]

    Estimation of non-normalized statistical models by score matching

    Hyvärinen, A., Dayan, P., 2005. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research 6

  30. [38]

    Protein-ligand inter- action prior for binding-aware 3d molecule diffusion models, in: The Twelfth International Conference on Learning Representations

    Huang, Z., Yang, L., Zhou, X., Zhang, Z., Zhang, W., Zheng, X., Chen, J., Wang, Y., Bin, C., Yang, W., 2023b. Protein-ligand inter- action prior for binding-aware 3d molecule diffusion models, in: The Twelfth International Conference on Learning Representations

  31. [39]

    Junction tree variational autoencoder for molecular graph generation, in: International confer- ence on machine learning, PMLR

    Jin, W., Barzilay, R., Jaakkola, T., 2018. Junction tree variational autoencoder for molecular graph generation, in: International confer- ence on machine learning, PMLR. pp. 2323–2332

  32. [40]

    Equiv- ariant 3d-conditional diffusion model for molecular linker design

    Igashov, I., Stärk, H., Vignac, C., Schneuing, A., Satorras, V.G., Frossard, P., Welling, M., Bronstein, M., Correia, B., 2024. Equiv- ariant 3d-conditional diffusion model for molecular linker design. Nature Machine Intelligence , 1–11

  33. [41]

    Auto-encoding variational{Bayes}, in: Int

    Kingma, D.P., Welling, M., . Auto-encoding variational{Bayes}, in: Int. Conf. on Learning Representations

  34. [42]

    Torsional diffusion for molecular conformer generation

    Jing, B., Corso, G., Chang, J., Barzilay, R., Jaakkola, T., 2022. Torsional diffusion for molecular conformer generation. Advances in Neural Information Processing Systems 35, 24240–24253

  35. [43]

    Diffusion-lm improves controllable text generation

    Li,X.,Thickstun,J.,Gulrajani,I.,Liang,P.S.,Hashimoto,T.B.,2022. Diffusion-lm improves controllable text generation. Advances in Neural Information Processing Systems 35, 4328–4343

  36. [44]

    Le, T., Cremer, J., Noe, F., Clevert, D.A., Schütt, K.T., . Navigating the design space of equivariant diffusion-based generative models for de novo 3d molecule generation, in: The Twelfth International Conference on Learning Representations

  37. [45]

    Functional-group-based diffusion for pocket-specific moleculegenerationandelaboration.AdvancesinNeuralInformation Processing Systems 36

    Lin, H., Huang, Y., Zhang, O., Liu, Y., Wu, L., Li, S., Chen, Z., Li, S.Z., 2024. Functional-group-based diffusion for pocket-specific moleculegenerationandelaboration.AdvancesinNeuralInformation Processing Systems 36

  38. [46]

    Autodiff:Autoregressivediffusionmodelingforstructure-baseddrug design

    Li, X., Wang, P., Fu, T., Gao, W., Li, C., Shi, L., Liu, J., 2024. Autodiff:Autoregressivediffusionmodelingforstructure-baseddrug design. arXiv preprint arXiv:2404.02003

  39. [47]

    Generating 3d molecules for target protein binding, in: International Conference on Machine Learning (ICML)

    Liu,M.,Luo,Y.,Uchino,K.,Maruhashi,K.,Ji,S.,2022. Generating 3d molecules for target protein binding, in: International Conference on Machine Learning (ICML)

  40. [48]

    Flow matching for generative modeling, in: The Eleventh International Conference on Learning Representations

    Lipman, Y., Chen, R.T., Ben-Hamu, H., Nickel, M., Le, M., . Flow matching for generative modeling, in: The Eleventh International Conference on Learning Representations

  41. [49]

    Classifier-free graph diffusionformolecularpropertytargeting,in:JointEuropeanConfer- ence on Machine Learning and Knowledge Discovery in Databases, Springer

    Ninniri, M., Podda, M., Bacciu, D., 2024. Classifier-free graph diffusionformolecularpropertytargeting,in:JointEuropeanConfer- ence on Machine Learning and Knowledge Discovery in Databases, Springer. pp. 318–335

  42. [50]

    Repaint: Inpainting using denoising diffusion probabilisticmodels,in:ProceedingsoftheIEEE/CVFconferenceon computer vision and pattern recognition, pp

    Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., Van Gool, L., 2022. Repaint: Inpainting using denoising diffusion probabilisticmodels,in:ProceedingsoftheIEEE/CVFconferenceon computer vision and pattern recognition, pp. 11461–11471

  43. [51]

    Open babel: An open chemical toolbox

    O’Boyle,N.M.,Banck,M.,James,C.A.,Morley,C.,Vandermeersch, T., Hutchison, G.R., 2011. Open babel: An open chemical toolbox. Journal of cheminformatics 3, 1–14

  44. [52]

    3d molecule generationbydenoisingvoxelgrids

    O Pinheiro, P.O., Rackers, J., Kleinhenz, J., Maser, M., Mahmood, O., Watkins, A., Ra, S., Sresht, V., Saremi, S., 2024. 3d molecule generationbydenoisingvoxelgrids. AdvancesinNeuralInformation Processing Systems 36

  45. [53]

    Scalablediffusionmodelswithtransform- ers, in: Proceedings of the IEEE/CVF International Conference on P

    Peebles,W.,Xie,S.,2023. Scalablediffusionmodelswithtransform- ers, in: Proceedings of the IEEE/CVF International Conference on P. Zhang et al.:Preprint submitted to Elsevier Page 12 of 13 Diffusion-based Small Molecule Generation Computer Vision, pp. 4195–4205

  46. [54]

    Training language models to follow instructions with human feedback

    Ouyang,L.,Wu,J.,Jiang,X.,Almeida,D.,Wainwright,C.,Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al., 2022. Training language models to follow instructions with human feedback. Ad- vances in neural information processing systems 35, 27730–27744

  47. [55]

    Hitting stride by degrees: Fine grained molecular generation via diffusion model

    Peng, X., Zhu, F., 2024. Hitting stride by degrees: Fine grained molecular generation via diffusion model. Expert Systems with Applications 244, 122949

  48. [56]

    Moldiff: Addressing the atom-bond inconsistency problem in 3d molecule diffusion genera- tion, in: International Conference on Machine Learning, PMLR

    Peng, X., Guan, J., Liu, Q., Ma, J., 2023. Moldiff: Addressing the atom-bond inconsistency problem in 3d molecule diffusion genera- tion, in: International Conference on Machine Learning, PMLR. pp. 27611–27629

  49. [57]

    Qian,H.,Huang,W.,Tu,S.,Xu,L.,2024.Kgdiff:towardsexplainable target-awaremoleculegenerationwithknowledgeguidance.Briefings in Bioinformatics 25, bbad435

  50. [58]

    Rcsb protein data bank

    Protein Data Bank, n.d. Rcsb protein data bank. URL: https: //www.rcsb.org/. accessed: [01.23.2025]

  51. [59]

    Quantum chemistry structures and properties of 134 kilo molecules

    Ramakrishnan, R., Dral, P.O., Rupp, M., Von Lilienfeld, O.A., 2014. Quantum chemistry structures and properties of 134 kilo molecules. Scientific data 1, 1–7

  52. [60]

    Coarse-to-fine: a hierarchical diffusion model for molecule generation in 3d, in: International Conference on Machine Learning, PMLR

    Qiang, B., Song, Y., Xu, M., Gong, J., Gao, B., Zhou, H., Ma, W.Y., Lan, Y., 2023. Coarse-to-fine: a hierarchical diffusion model for molecule generation in 3d, in: International Conference on Machine Learning, PMLR. pp. 28277–28299

  53. [61]

    A small-molecule tnikinhibitortargetsfibrosisinpreclinicalandclinicalmodels.Nature Biotechnology 43, 63–75

    Ren, F., Aliper, A., Chen, J., Zhao, H., Rao, S., Kuppe, C., Ozerov, I.V., Zhang, M., Witte, K., Kruse, C., et al., 2025. A small-molecule tnikinhibitortargetsfibrosisinpreclinicalandclinicalmodels.Nature Biotechnology 43, 63–75

  54. [62]

    RDKit: Open-source cheminformatics.http://www

    RDKit, online, . RDKit: Open-source cheminformatics.http://www. rdkit.org. [Online; accessed 11-April-2013]

  55. [63]

    Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation, in: Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, pro- ceedings, part III...

  56. [64]

    High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp

    Rombach,R.,Blattmann,A.,Lorenz,D.,Esser,P.,Ommer,B.,2022. High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695

  57. [65]

    E (n) equivari- ant graph neural networks, in: International conference on machine learning, PMLR

    Satorras, V.G., Hoogeboom, E., Welling, M., 2021. E (n) equivari- ant graph neural networks, in: International conference on machine learning, PMLR. pp. 9323–9332

  58. [66]

    Silvr: Guided diffusion for molecule generation

    Runcie, N.T., Mey, A.S., 2023. Silvr: Guided diffusion for molecule generation. JournalofChemicalInformationandModeling63,5996– 6005

  59. [67]

    Structure- baseddrugdesignwithequivariantdiffusionmodels

    Schneuing, A., Harris, C., Du, Y., Didi, K., Jamasb, A., Igashov, I., Du, W., Gomes, C., Blundell, T.L., Lio, P., et al., 2024. Structure- baseddrugdesignwithequivariantdiffusionmodels. NatureCompu- tational Science 4, 899–909

  60. [68]

    Fast high-resolution image synthesis with latent ad- versarialdiffusiondistillation,in:SIGGRAPHAsia2024Conference Papers, pp

    Sauer, A., Boesel, F., Dockhorn, T., Blattmann, A., Esser, P., Rom- bach, R., 2024. Fast high-resolution image synthesis with latent ad- versarialdiffusiondistillation,in:SIGGRAPHAsia2024Conference Papers, pp. 1–11

  61. [69]

    Consistency models, in: International Conference on Machine Learning, PMLR

    Song, Y., Dhariwal, P., Chen, M., Sutskever, I., 2023. Consistency models, in: International Conference on Machine Learning, PMLR. pp. 32211–32252

  62. [70]

    Simonovsky, M., Komodakis, N., 2018. Graphvae: Towards genera- tionofsmallgraphsusingvariationalautoencoders,in:ArtificialNeu- ralNetworksandMachineLearning–ICANN2018:27thInternational Conference on Artificial Neural Networks, Rhodes, Greece, October 4-7, 2018, Proceedings, Pa...

  63. [71]

    Score-based generative modeling through stochas- tic differential equations

    Song, Y., Sohl-dickstein, J., Kingma, D.P., Kumar, A., Ermon, S., Poole, B., 2021. Score-based generative modeling through stochas- tic differential equations. URL: https://openreview.net/forum?id= PxTIG12RRHS

  64. [72]

    Generative modeling by estimating gradients of the data distribution

    Song, Y., Ermon, S., 2019. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems 32

  65. [73]

    Digress: Discrete denoising diffusion for graph generation, in: ICLR

    Vignac, C., Krawczuk, I., Siraudin, A., Wang, B., Cevher, V., Frossard, P., 2023a. Digress: Discrete denoising diffusion for graph generation, in: ICLR

  66. [74]

    A deep learning approach to antibiotic discovery

    Stokes, J.M., Yang, K., Swanson, K., Jin, W., Cubillos-Ruiz, A., Donghia, N.M., MacNair, C.R., French, S., Carfrae, L.A., Bloom- Ackermann, Z., et al., 2020. A deep learning approach to antibiotic discovery. Cell 180, 688–702

  67. [75]

    Boosting performance of generative diffusion model for molecular docking by training on artificial binding pockets

    Voitsitskyi, T., Bdzhola, V., Stratiichuk, R., Koleiev, I., Ostrovsky, Z.,Vozniak,V.,Khropachov,I.,Henitsoi,P.,Popryho,L.,Zhytar,R., et al., 2023. Boosting performance of generative diffusion model for molecular docking by training on artificial binding pockets. bioRxiv , 2023–11

  68. [76]

    Midi: Mixed graph and 3d denoising diffusion for molecule generation, in: Joint European Conference on Machine Learning and Knowledge Discov- ery in Databases, Springer

    Vignac, C., Osman, N., Toni, L., Frossard, P., 2023b. Midi: Mixed graph and 3d denoising diffusion for molecule generation, in: Joint European Conference on Machine Learning and Knowledge Discov- ery in Databases, Springer. pp. 560–576

  69. [77]

    Swallowing the bitter pill: Simplified scalable conformer generation, in: ICML 2024 AI for Science Workshop

    Wang, Y., Elhag, A.A., Jaitly, N., Susskind, J.M., Bautista, M.Á., . Swallowing the bitter pill: Simplified scalable conformer generation, in: ICML 2024 AI for Science Workshop

  70. [78]

    Gldm: hit molecule generation with constrained graph latent diffusion model

    Wang, C., Ong, H.H., Chiba, S., Rajapakse, J.C., 2024. Gldm: hit molecule generation with constrained graph latent diffusion model. Briefings in Bioinformatics 25, bbae142

  71. [79]

    Guided diffusionformoleculargenerationwithinteractionprompt

    Wu, P., Du, H., Yan, Y., Lee, T.Y., Bai, C., Wu, S., 2024. Guided diffusionformoleculargenerationwithinteractionprompt. Briefings in Bioinformatics 25, bbae174

  72. [80]

    Smiles, a chemical language and information system

    Weininger, D., 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of chemical information and computer sciences 28, 31–36

  73. [81]

    Geometric- facilitated denoising diffusion model for 3d molecule generation, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp

    Xu, C., Wang, H., Wang, W., Zheng, P., Chen, H., 2024. Geometric- facilitated denoising diffusion model for 3d molecule generation, in: Proceedings of the AAAI Conference on Artificial Intelligence, pp. 338–346

  74. [82]

    Diffdec: Structure-aware scaffold decoration with an end-to-end diffusion model

    Xie, J., Chen, S., Lei, J., Yang, Y., 2024. Diffdec: Structure-aware scaffold decoration with an end-to-end diffusion model. Journal of Chemical Information and Modeling

  75. [83]

    Geodiff: A geometric diffusion model for molecular conformation generation, in: International Conference on Learning Representations

    Xu,M.,Yu,L.,Song,Y.,Shi,C.,Ermon,S.,Tang,J.,2022. Geodiff: A geometric diffusion model for molecular conformation generation, in: International Conference on Learning Representations

  76. [84]

    Geometric latent diffusion models for 3d molecule generation, in: International Conference on Machine Learning, PMLR

    Xu, M., Powers, A.S., Dror, R.O., Ermon, S., Leskovec, J., 2023. Geometric latent diffusion models for 3d molecule generation, in: International Conference on Machine Learning, PMLR. pp. 38592– 38610

  77. [85]

    Yim, J., Stärk, H., Corso, G., Jing, B., Barzilay, R., Jaakkola, T.S.,

  78. [86]

    Prompt-based 3d molecular diffusion models for structure-based drug design

    Yang, L., Huang, Z., Zhou, X., Xu, M., Zhang, W., Wang, Y., Zheng, X., Yang, W., Dror, R.O., Hong, S., CUI, B., 2024. Prompt-based 3d molecular diffusion models for structure-based drug design. URL: https://openreview.net/forum?id=FWsGuAFn3n

  79. [87]

    Graph neural networks: A review of methods and applications

    Zhou, J., Cui, G., Hu, S., Zhang, Z., Yang, C., Liu, Z., Wang, L., Li, C., Sun, M., 2020. Graph neural networks: A review of methods and applications. AI open 1, 57–81

  80. [88]

    WileyInter- disciplinary Reviews: Computational Molecular Science 14, e1711

    Diffusionmodelsinproteinstructureanddocking. WileyInter- disciplinary Reviews: Computational Molecular Science 14, e1711

  81. [89]

    Stride: Structure-guided generation for inverse design of molecules, in: NeurIPS 2023 AI for Science Workshop

    Zaman, S., Akhiyarov, D., Araya-Polo, M., Chiu, K., 2023. Stride: Structure-guided generation for inverse design of molecules, in: NeurIPS 2023 AI for Science Workshop

  82. [91]

    Decompopt: Controllable and decomposed diffusion models for structure-basedmolecularoptimization,in:TheTwelfthInternational Conference on Learning Representations

    Zhou, X., Cheng, X., Yang, Y., Bao, Y., Wang, L., Gu, Q., 2024. Decompopt: Controllable and decomposed diffusion models for structure-basedmolecularoptimization,in:TheTwelfthInternational Conference on Learning Representations

  83. [92]

    Molsnapper: Conditioning diffusion for structure based drug design

    Ziv, Y., Marsden, B., Deane, C., 2024. Molsnapper: Conditioning diffusion for structure based drug design. bioRxiv , 2024–03. P. Zhang et al.:Preprint submitted to Elsevier Page 13 of 13

  84. [2024]

    Nature , 1–3

    Accuratestructurepredictionofbiomolecularinteractionswith alphafold 3. Nature , 1–3

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.