REVIEW 4 major objections 4 minor 1 cited by
An evaluation of unconditional 3D molecular generation methods
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This paper argues that the near-saturation of standard 3D molecule-generation benchmarks is misleading: once chemical and physical validity are required, the best raw generator reaches 87.0% valid-unique-novel molecules, and its own…
desk verdict Useful, mostly reproducible validity-aware benchmark of five 3D generators, undercut by an abstract that contradicts its own post-processing table and a raw ranking confounded by hydrogen representation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is a two-stage validity filter applied to every generated molecule. The first stage is graph-level: the molecule must parse, pass chemical sanitisation, have all explicit hydrogens, and be connected. The second stage is conformation-level: the 3D structure must pass six geometry and energy checks, including bond lengths and angles within 25% of experimental bounds, planar aromatic rings, planar double bonds, no internal steric clash, and a force-field strain-energy ratio below 100. The paper applies this filter, together with uniqueness and novelty checks based on canonical string identifiers, to 100,000 unconditional samples per model, with and without post-processing. Because some thresholds, notably the 30% van der Waals overlap allowance and the strain ratio of 100, are deliberately generous, passing them is a meaningful but not strict test of physical plausibility.
What would settle it
Re-evaluate the same five models' 100,000 samples under stricter physicality cutoffs, for example 10% bond-geometry deviation, 10% van der Waals overlap allowance, and a strain-energy ratio of 20. If the best methods still score above 90% valid-unique-novel, the paper's conclusion that unconditional 3D generation is not saturated would be called into question; if their scores drop sharply, the conclusion stands but is threshold-dependent.
Extended reading notes
Core claim
Put in the authors' terms, the discovery is that standard benchmarks are saturated but the task is not. More than 99% of the valid molecules generated by the five methods are unique and novel, so diversity is not the bottleneck. The bottleneck is validity once it is defined to include the physical conformation: raw success rates are 59.7% for EQGAT-diff, 59.8% for FlowMol, 0.2% for GCDM, 2.9% for GeoLDM, and 87.5% for SemlaFlow. The two low models fail mainly because they do not generate explicit hydrogens; the other failures are split between chemically invalid graphs and physically invalid conformations. Post-processing, meaning keeping the largest fragment, adding hydrogens, and relaxing the structure with a universal force field, changes the ordering, putting GCDM at 95.2%, SemlaFlow at 92.4%, EQGAT-diff and FlowMol at 84.1%, and GeoLDM at 69.6%. Even the best results remain below the 94.2% validity of the training data and the 98.8% of an approved-drug reference set.
Load-bearing premise
The load-bearing premise is that the chosen thresholds, a 25% tolerance on bond geometry, a 30% allowance on van der Waals overlap, and a strain-energy ratio of 100, define what counts as physically valid; if a stricter definition is used, the reported success rates and the ranking of methods could change.
Editorial extensions
If this is right
- The reported valid-unique-novel shares are upper bounds on useful output, since several of the physicality thresholds are acknowledged to be generous.
- Uniqueness and novelty are effectively solved for these models, with over 99% of valid molecules being novel and unique, so future work should target chemical and physical validity rather than diversity.
- Hydrogen handling is a first-order design decision: models that omit explicit hydrogens jump from near-zero raw validity to high validity once hydrogens are added in post-processing.
- Because the training set itself scores 94.2% on the same checks, even the best method leaves room for improvement, contradicting the saturated-benchmark narrative.
- Distribution metrics such as drug-likeness and synthetic accessibility can match the training distribution while chemical-space coverage remains incomplete, so those metrics alone cannot certify coverage.
Reading between the lines
- Using the geometry and strain checks as a rejection filter or scoring term at sampling time could push several methods above 90% valid-unique-novel, an option the paper does not test.
- The hydrogen failure pattern implies that changing the generative target from heavy-atom-only structures to full explicit-hydrogen structures could remove the largest chemical-invalidity source at the root.
- Since the paper frames unconditional generation as a stepping stone to conditional tasks, the same physical-validity failures would likely appear in conditional generation; auditing conditional outputs with the same conformer checks would test that transfer.
- Rankings by a single aggregate rate hide different failure modes: models fail at different stages, so future evaluations should report graph and conformation validity separately, as the paper's own tables do.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates five recent unconditional 3D molecular generation methods (EQGAT-diff, FlowMol, GCDM, GeoLDM, SemlaFlow) on the GEOM Drugs benchmark. The authors generate 100,000 molecules per method and report standard validity, uniqueness, and novelty metrics together with chemical and physical validity checks based on RDKit and the PoseBusters suite, both with and without post-processing (largest fragment, hydrogen addition, UFF minimization). They report that SemlaFlow achieves 87.0% valid-unique-novel molecules without post-processing and GCDM achieves 95.2% with post-processing, and they conclude that the widely reported saturated benchmarks are misleading because physical and chemical validity are far from perfect.
Significance. If the findings are correct, the paper makes a useful contribution by showing that standard benchmarks may saturate while geometric and energetic validity remains incomplete, and it provides detailed per-test failure analyses (chemical components in Table 4, physical components in Table 5) and a transparent computational protocol. The use of publicly available model weights and the fixed published thresholds of PoseBusters are strengths. However, internal inconsistencies in the headline claims and a methodological confound between hydrogen representation and chemical validity weaken the contribution as written and require correction before the conclusions can be fully credited.
major comments (4)
- [Abstract and Section 3 (Table 3)] The abstract states: "Overall, the best method, SemlaFlow, has a success rate of 87% in generating valid, unique, and novel molecules without post-processing and 92.4% with post-processing." This is directly contradicted by Section 3, which says "With post-processing, the best method according to these metrics becomes GCDM," and by Table 3, where GCDM+PP achieves 95.2% versus SemlaFlow+PP at 92.4%. The abstract must be corrected to report GCDM as the best post-processed method or to clearly qualify SemlaFlow as best only without post-processing.
- [Section 3 and Section 2.2 (Tables 1, 3, 4)] The no-post-processing ranking is confounded by hydrogen representation. Section 2.2 includes "the molecule has all of its hydrogens added explicitly" as a chemical validity test, but GCDM and GeoLDM output heavy-atom-only structures, with explicit-hydrogen pass rates of 0.2% and 5.3% in Table 4. Consequently, the sentence "All five 3D molecular generation methods generated large sets of valid, unique, and novel molecules" is false for these two methods, which achieve only 0.2% and 2.9% valid-unique-novel in Table 1. The authors acknowledge the cause, but a fair method comparison requires either identical hydrogen addition for all methods before the validity assessment or a heavy-atom-only validity definition; without one of these, the 87.0% headline for SemlaFlow conflates model quality with output format.
- [Section 3] The sentence "Given that these training data scores are higher than the best model, there is still room for improvement" is contradicted by Table 3: GEOM Drugs achieves 99.8% chemical and 94.2% physical validity for an overall 94.2%, while GCDM+PP achieves 95.5% chemical and 95.2% physical validity for an overall 95.2%. The best post-processed model therefore exceeds the training-data aggregate validity. The claim should be revised to refer to the per-test failure rates (for example, connectedness or steric clashes) rather than the aggregate training-data scores.
- [Section 2.2 and Tables 1/3] All success rates are reported as point estimates without variance, confidence intervals, or repeated-seed statistics, and the physical-validity thresholds are acknowledged as generous (30% van der Waals overlap tolerance and UFF energy ratio of 100). Since the post-processing ranking changes by only a few percentage points (SemlaFlow+PP 92.4% versus GCDM+PP 95.2%), the authors should report variability across runs and a sensitivity analysis over the validity thresholds to establish that the ranking is robust rather than an artifact of threshold choice.
minor comments (4)
- [Section 3] The sentence "The failure modes are in terms of both the generated molecular graphs and the 3D conformations are shown in Table 3" is grammatically broken and should be rephrased, for example as "The failure modes, for both the generated molecular graphs and the 3D conformations, are shown in Table 3."
- [Appendix C] The phrase "CA approved" appears to be a typo; if the intended status is FDA-approved or simply "approved," the text should be corrected to avoid ambiguity.
- [Table 6 and Appendix A] The term "spacial score" is likely a misspelling of "spatial score," unless the cited source (Krzyzanowski et al., 2023) intentionally uses that spelling; please verify and align the usage with the original publication.
- [Figure 3 caption] The caption contains a formatting error ("Fr ´echet") and the sentence about the recommended data set size "(5'000)" is unclear; please state the source of the recommended size and use consistent thousands separators.
Circularity Check
No circularity found: the benchmark is self-contained and no predicted quantity is identical to an input by construction.
full rationale
This paper is an empirical evaluation rather than a derivation, and no step in its argument reduces to its own inputs. The validity definition is stated upfront: a molecule must pass four chemical tests and six physical PoseBusters tests, and the authors apply these fixed tests uniformly to all five models and to the GEOM Drugs and DrugBank reference sets. No parameter is fitted to any subset of the generated data and then renamed as a prediction; the reported percentages are direct counts of molecules passing the predefined tests. The only author-overlapping input is the PoseBusters tool (Buttenschoen et al., 2024), but it is an externally published, fixed-threshold benchmark, and the thresholds are stated openly rather than tuned to the reported ranking, so its use does not make the central claim circular. The paper also discloses the hydrogen-format confound explicitly: 'Note that GCDM and GeoLDM do not add all hydrogens without post-processing,' and later states, 'The improvements of GCDM and GeoLDM are mostly due to the addition of hydrogens since these two methods do not generate all hydrogens explicitly.' That is a benchmark-design limitation, not a circularity, because the paper does not claim these models pass a heavy-atom-only validity test and it separately reports post-processed results. No uniqueness theorem, ansatz, or known result is imported from the authors' prior work to force a conclusion, and no equation in the paper has an output that is equivalent to its input by definition. The comparison of model scores to training-data scores is a standard benchmark practice even if the no-post-processing ranking is debatable; the debate concerns metric choice and interpretation, not circular reasoning.
Assumptions & free parameters
free parameters (3)
- Bond length/angle validity tolerance =
25% deviation threshold
- Steric clash tolerance =
30% van der Waals overlap allowed
- Strain energy ratio cutoff =
100 (UFF energy ratio)
assumptions (4)
- domain assumption RDKit sanitization and canonical SMILES are valid operational definitions of chemical validity, uniqueness, and novelty.
- domain assumption PoseBusters geometric thresholds and the UFF energy ratio are valid proxies for physical validity of conformations.
- domain assumption The published model weights are faithful implementations of EQGAT-diff, FlowMol, GCDM, GeoLDM, and SemlaFlow.
- domain assumption GEOM Drugs is an appropriate novelty reference and distribution target for all five methods despite differing train/validation splits.
Cite this review
Pith. "Pith review of An evaluation of unconditional 3D molecular generation methods." pith.science (2026). https://pith.science/paper/NIPDYBT4
@misc{pith2026250500518,
author = {Pith},
title = {Pith review of: An evaluation of unconditional 3D molecular generation methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/NIPDYBT4}},
note = {Machine review of arXiv:2505.00518}
}
read the original abstract
Unconditional molecular generation is a stepping stone for conditional molecular generation, which is important in \emph{de novo} drug design. Recent unconditional 3D molecular generation methods report saturated benchmarks, suggesting it is time to re-evaluate our benchmarks and compare the latest models. We assess five recent high-performing 3D molecular generation methods (EQGAT-diff, FlowMol, GCDM, GeoLDM, and SemlaFlow), in terms of both standard benchmarks and chemical and physical validity. Overall, the best method, SemlaFlow, has a success rate of 87% in generating valid, unique, and novel molecules without post-processing and 92.4% with post-processing.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
FlowMol3: Flow Matching for 3D De Novo Small-Molecule Generation
Combining self-conditioning, fake atoms, and late-stage geometry distortion lets a compact flow-matching model generate nearly always valid 3D drug-like molecules and match training-data chemistry better than existing...
Reference graph
Works this paper leans on
-
[1]
Cormorant: Covariant molecular neural networks
Brandon Anderson, Truong Son Hy, and Risi Kondor. Cormorant: Covariant molecular neural networks. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019
work page 2019
-
[2]
GEOM , energy-annotated molecular conformations for property prediction and molecular generation
Simon Axelrod and Rafael G \'o mez-Bombarelli . GEOM , energy-annotated molecular conformations for property prediction and molecular generation. Scientific Data, 9 0 (1): 0 185, 2022
work page 2022
-
[3]
Deep generative models for 3D molecular structure
Benoit Baillif, Jason Cole, Patrick McCabe, and Andreas Bender. Deep generative models for 3D molecular structure. Current Opinion in Structural Biology, 80: 0 102566, 2023
work page 2023
-
[4]
G. Richard Bickerton, Gaia V. Paolini, J \'e r \'e my Besnard, Sorel Muresan, and Andrew L. Hopkins. Quantifying the chemical beauty of drugs. Nature Chemistry, 4 0 (2): 0 90--98, 2012
work page 2012
-
[5]
Nathan Brown, Marco Fiscato, Marwin H.S. Segler, and Alain C. Vaucher. GuacaMol : Benchmarking models for de novo molecular design. Journal of Chemical Information and Modeling, 59 0 (3): 0 1096--1108, 2019
work page 2019
-
[6]
Martin Buttenschoen, Garrett M. Morris, and Charlotte M. Deane. PoseBusters : AI-based docking methods fail to generate physically valid poses or generalise to novel sequences. Chemical Science, 15 0 (9): 0 3130--3139, 2024
work page 2024
-
[7]
Markus Dablander, Thierry Hanser, Renaud Lambiotte, and Garrett M. Morris. Sort & Slice : A simple and superior alternative to hash-based folding for extended-connectivity fingerprints. Journal of Cheminformatics, 16 0 (1): 0 135, 2024
work page 2024
-
[8]
Ian Dunn and David R. Koes. Exploring discrete flow matching for 3D de novo molecule generation, 2024. arXiv:2411.16644
arXiv 2024
Show all 35 references
-
[9]
Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions
Peter Ertl and Ansgar Schuffenhauer. Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of Cheminformatics, 1 0 (1): 0 8, 2009
2009
-
[10]
Equivariant diffusion for molecule generation in 3D
Emiel Hoogeboom, V \'i ctor Garcia Satorras, Cl \'e ment Vignac, and Max Welling. Equivariant diffusion for molecule generation in 3D . In Proceedings of the 39th International Conference on Machine Learning , volume 162 of Proceedings of Machine Learning Research , pp.\ 8867-...
2022
-
[11]
Efficient 3D molecular generation with flow matching and scale optimal transport
Ross Irwin, Alessandro Tibo, Jon Paul Janet, and Simon Olsson. Efficient 3D molecular generation with flow matching and scale optimal transport. In ICML 2024 AI for Science Workshop , 2024
2024
-
[12]
DrugBank 6.0: The DrugBank knowledgebase for 2024
Craig Knox, Mike Wilson, Christen M Klinger, Mark Franklin, Eponine Oler, Alex Wilson, Allison Pon, Jordan Cox, Na Eun (Lucy) Chin, Seth A Strawbridge, Marysol Garcia-Patino , Ray Kruger, Aadhavya Sivakumaran, Selena Sanford, Rahil Doshi, Nitya Khetarpal, Omolola Fatokun, Daph...
2024
-
[13]
Spacial score-a comprehensive topological indicator for small-molecule complexity
Adrian Krzyzanowski, Axel Pahl, Michael Grigalunas, and Herbert Waldmann. Spacial score-a comprehensive topological indicator for small-molecule complexity. Journal of Medicinal Chemistry, 66 0 (18): 0 12739--12750, 2023
2023
-
[14]
Scalfani, Rachel Walker, Kazuya Ujihara, Daniel Probst, Juuso Lehtivarjo, Guillaume Godin, Axel Pahl, Fran c ois Fran c ois B \'e renger, and Hussein Faara
Greg Landrum, Paolo Tosco, Brian Kelley, Ricardo Rodriguez, David Cosgrove, Riccardo Vianello, Sereina Riniker, Peter Gedeck, Gareth Jones, Eisuke Kawashima, Andrew Dalke, Matt Swain, Brian Cole, Samo Turk, Aleksandr Savelev, Alain Vaucher, Maciej W \'o jcikowski, Ichiru Take,...
2024
-
[15]
Navigating the design space of equivariant diffusion-based generative models for de novo 3D molecule generation
Tuan Le, Julian Cremer, Frank Noe, Djork-Arn \'e Clevert, and Kristof T Sch \"u tt. Navigating the design space of equivariant diffusion-based generative models for de novo 3D molecule generation. In The Twelfth International Conference on Learning Representations, 2024
2024
-
[16]
UMAP : Uniform manifold approximation and projection for dimension reduction, 2020
Leland McInnes, John Healy, and James Melville. UMAP : Uniform manifold approximation and projection for dimension reduction, 2020. arXiv:1802.03426
2020 arXiv
-
[17]
Geometry-complete diffusion for 3D molecule generation and optimization
Alex Morehead and Jianlin Cheng. Geometry-complete diffusion for 3D molecule generation and optimization. Communications Chemistry, 7 0 (1): 0 150, 2024
2024
-
[18]
H. L. Morgan. The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. Journal of Chemical Documentation, 5 0 (2): 0 107--113, 1965
1965
-
[19]
Open Babel : An open chemical toolbox
Noel M O'Boyle, Michael Banck, Craig A James, Chris Morley, Tim Vandermeersch, and Geoffrey R Hutchison. Open Babel : An open chemical toolbox. Journal of Cheminformatics, 3 0 (1): 0 33, 2011
2011
-
[20]
MolDiff : Addressing the atom-bond inconsistency problem in 3D molecule diffusion generation
Xingang Peng, Jiaqi Guan, Qiang Liu, and Jianzhu Ma. MolDiff : Addressing the atom-bond inconsistency problem in 3D molecule diffusion generation. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, p...
2023
-
[21]
Molecular sets ( MOSES ): A benchmarking platform for molecular generation models
Daniil Polykovskiy, Alexander Zhebrak, Benjamin Sanchez-Lengeling , Sergey Golovanov, Oktai Tatanov, Stanislav Belyaev, Rauf Kurbanov, Aleksey Artamonov, Vladimir Aladinskiy, Mark Veselov, Artur Kadurin, Simon Johansson, Hongming Chen, Sergey Nikolenko, Al \'a n Aspuru-Guzik ,...
2020
-
[22]
Fr \'e chet ChemNet distance: A metric for generative models for molecules in drug discovery
Kristina Preuer, Philipp Renz, Thomas Unterthiner, Sepp Hochreiter, and G \"u nter Klambauer. Fr \'e chet ChemNet distance: A metric for generative models for molecules in drug discovery. Journal of Chemical Information and Modeling, 58 0 (9): 0 1736--1741, 2018
2018
-
[23]
A. K. Rappe, C. J. Casewit, K. S. Colwell, W. A. Goddard, and W. M. Skiff. UFF , a full periodic table force field for molecular mechanics and molecular dynamics simulations. Journal of the American Chemical Society, 114 0 (25): 0 10024--10035, 1992
1992
-
[24]
Sereina Riniker and Gregory A. Landrum. Better informed distance geometry: Using what we know to improve conformation generation. Journal of Chemical Information and Modeling, 55 0 (12): 0 2562--2574, 2015
2015
-
[25]
E(n) equivariant graph neural networks, 2021
Victor Garcia Satorras, Emiel Hoogeboom, and Max Welling. E(n) equivariant graph neural networks, 2021. arXiv:2102.09844
2021 arXiv
-
[26]
Equivariant flow matching with hybrid probability transport for 3D molecule generation
Yuxuan Song, Jingjing Gong, Minkai Xu, Ziyao Cao, Yanyan Lan, Stefano Ermon, Hao Zhou, and Wei-Ying Ma. Equivariant flow matching with hybrid probability transport for 3D molecule generation. In Thirty-Seventh Conference on Neural Information Processing Systems, 2023
2023
-
[27]
MiDi : Mixed graph and 3D denoising diffusion for molecule generation
Cl \'e ment Vignac, Nagham Osman, Laura Toni, and Pascal Frossard. MiDi : Mixed graph and 3D denoising diffusion for molecule generation. In Machine Learning and Knowledge Discovery in Databases : Research Track , volume 14170, pp.\ 560--576. Springer Nature Switzerland, Cham, 2023
2023
-
[28]
Michael L. Waskom. Seaborn: Statistical data visualization. Journal of Open Source Software, 6 0 (60): 0 3021, 2021
2021
-
[29]
Wildman and Gordon M
Scott A. Wildman and Gordon M. Crippen. Prediction of physicochemical parameters by atomic contributions. Journal of Chemical Information and Computer Sciences, 39 0 (5): 0 868--873, 1999
1999
-
[30]
Dror, Stefano Ermon, and Jure Leskovec
Minkai Xu, Alexander S Powers, Ron O. Dror, Stefano Ermon, and Jure Leskovec. Geometric latent diffusion models for 3D molecule generation. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pp.\ 385...
-
[31]
Yael Ziv, Brian Marsden, and Charlotte M. Deane. MolSnapper : Conditioning diffusion for structure based drug design, 2024. bioRxiv:2024.03.28.586278
2024
-
[32]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[33]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[34]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[35]
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.