{"id":"7da647a8-2bfb-4a5d-a093-58eed5ee76cd","arxiv_id":"2505.00518","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An evaluation of five unconditional 3D molecular generators finds that standard saturated benchmarks overstate quality, and that post-processing changes the model ranking.","lead":"Five recent 3D molecular generation models were tested on 100,000 samples each, with and without post-processing, using standard benchmarks and PoseBusters geometry checks. The comparison shows benchmark saturation is misleading: method rankings and validity rates change substantially once chemical and physical validity is required, and post-processing changes the winner from SemlaFlow to GCDM.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-post-processing ranking is confounded by hydrogen representation: GCDM and GeoLDM output heavy atoms only (explicit-H pass 0.2% and 5.3%), so SemlaFlow's 87% headline partly reflects output format, not model quality, and the paper's own post-processing results give GCDM 95.2%.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the identified threshold sensitivity is real. However, the most load-bearing weakness is not primarily the generosity of the clash or strain thresholds; it is the hydrogen-representation mismatch between models. GCDM and GeoLDM are penalized in the no-post-processing comparison for a formatting choice, and the paper's own Table 4 shows that their explicit-H failure is the dominant cause of their near-zero raw validity. This undermines the abstract's best-method claim and the claim that all five methods produced large valid sets. The concern does not invalidate the overall evaluation, because the post-processed results and the acknowledgment in Section 3 partly correct for it, but the comparison should be made representation-fair before the results are used as a canonical benchmark. I therefore keep the reader's CONDITIONAL verdict unchanged.","tokens_in":11077,"tokens_out":8395,"duration_ms":89987,"concrete_test":"Recompute Table 1 in a heavy-atom-only mode: strip all hydrogens from every generated molecule and from the GEOM Drugs and DrugBank references before applying the chemical validity tests, so the explicit-hydrogen check is neutralized. If GCDM and GeoLDM raw valid-unique-novel rates rise from 0.2% and 2.9% to levels comparable to the other methods, or if SemlaFlow's no-post-processing lead materially shrinks, then the no-PP headline is an output-format artifact rather than a model-capability result. This single rerun directly tests whether the hydrogen-format mismatch is load-bearing for the paper's ranking and saturation claims.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 makes 'explicit hydrogens' a required chemical-validity test, but two of the five evaluated models, GCDM and GeoLDM, generate heavy-atom-only structures. The paper itself notes this in Section 3, and Table 4 shows the effect: GCDM passes sanitization at 97.2% and connectedness at 86.6%, yet passes the explicit-hydrogen test at only 0.2%; GeoLDM passes sanitization at 72.5%, connectedness at 37.4%, and explicit-H at 5.3%. Consequently, the raw validity rates of 0.2% and 2.9% in Table 1 are dominated by a hydrogen-format requirement rather than by chemical invalidity. The prose claim that all five methods generated large sets of valid, unique, and novel molecules is therefore false for these two rows, and the abstract's headline that SemlaFlow is best without post-processing conflates model quality with output representation. This is not a hidden flaw: Section 3 attributes most of the GCDM and GeoLDM post-processing gains to adding hydrogens. But the unqualified no-post-processing ranking and the sentence 'Given that these training data scores are higher than the best model, there is still room for improvement' remain misleading, because Table 3 gives GCDM+PP 95.2% overall validity, higher than GEOM Drugs' 94.2%. A fair method comparison requires either identical hydrogen addition for all models before ranking or a heavy-atom-only validity definition.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript evaluates five recent unconditional 3D molecular generation methods (EQGAT-diff, FlowMol, GCDM, GeoLDM, SemlaFlow) on the GEOM Drugs benchmark. The authors generate 100,000 molecules per method and report standard validity, uniqueness, and novelty metrics together with chemical and physical validity checks based on RDKit and the PoseBusters suite, both with and without post-processing (largest fragment, hydrogen addition, UFF minimization). They report that SemlaFlow achieves 87.0% valid-unique-novel molecules without post-processing and GCDM achieves 95.2% with post-processing, and they conclude that the widely reported saturated benchmarks are misleading because physical and chemical validity are far from perfect.","tokens_in":11416,"tokens_out":5651,"duration_ms":53187,"significance":"If the findings are correct, the paper makes a useful contribution by showing that standard benchmarks may saturate while geometric and energetic validity remains incomplete, and it provides detailed per-test failure analyses (chemical components in Table 4, physical components in Table 5) and a transparent computational protocol. The use of publicly available model weights and the fixed published thresholds of PoseBusters are strengths. However, internal inconsistencies in the headline claims and a methodological confound between hydrogen representation and chemical validity weaken the contribution as written and require correction before the conclusions can be fully credited.","major_comments":[{"comment":"The abstract states: \"Overall, the best method, SemlaFlow, has a success rate of 87% in generating valid, unique, and novel molecules without post-processing and 92.4% with post-processing.\" This is directly contradicted by Section 3, which says \"With post-processing, the best method according to these metrics becomes GCDM,\" and by Table 3, where GCDM+PP achieves 95.2% versus SemlaFlow+PP at 92.4%. The abstract must be corrected to report GCDM as the best post-processed method or to clearly qualify SemlaFlow as best only without post-processing.","section":"Abstract and Section 3 (Table 3)"},{"comment":"The no-post-processing ranking is confounded by hydrogen representation. Section 2.2 includes \"the molecule has all of its hydrogens added explicitly\" as a chemical validity test, but GCDM and GeoLDM output heavy-atom-only structures, with explicit-hydrogen pass rates of 0.2% and 5.3% in Table 4. Consequently, the sentence \"All five 3D molecular generation methods generated large sets of valid, unique, and novel molecules\" is false for these two methods, which achieve only 0.2% and 2.9% valid-unique-novel in Table 1. The authors acknowledge the cause, but a fair method comparison requires either identical hydrogen addition for all methods before the validity assessment or a heavy-atom-only validity definition; without one of these, the 87.0% headline for SemlaFlow conflates model quality with output format.","section":"Section 3 and Section 2.2 (Tables 1, 3, 4)"},{"comment":"The sentence \"Given that these training data scores are higher than the best model, there is still room for improvement\" is contradicted by Table 3: GEOM Drugs achieves 99.8% chemical and 94.2% physical validity for an overall 94.2%, while GCDM+PP achieves 95.5% chemical and 95.2% physical validity for an overall 95.2%. The best post-processed model therefore exceeds the training-data aggregate validity. The claim should be revised to refer to the per-test failure rates (for example, connectedness or steric clashes) rather than the aggregate training-data scores.","section":"Section 3"},{"comment":"All success rates are reported as point estimates without variance, confidence intervals, or repeated-seed statistics, and the physical-validity thresholds are acknowledged as generous (30% van der Waals overlap tolerance and UFF energy ratio of 100). Since the post-processing ranking changes by only a few percentage points (SemlaFlow+PP 92.4% versus GCDM+PP 95.2%), the authors should report variability across runs and a sensitivity analysis over the validity thresholds to establish that the ranking is robust rather than an artifact of threshold choice.","section":"Section 2.2 and Tables 1/3"}],"minor_comments":[{"comment":"The sentence \"The failure modes are in terms of both the generated molecular graphs and the 3D conformations are shown in Table 3\" is grammatically broken and should be rephrased, for example as \"The failure modes, for both the generated molecular graphs and the 3D conformations, are shown in Table 3.\"","section":"Section 3"},{"comment":"The phrase \"CA approved\" appears to be a typo; if the intended status is FDA-approved or simply \"approved,\" the text should be corrected to avoid ambiguity.","section":"Appendix C"},{"comment":"The term \"spacial score\" is likely a misspelling of \"spatial score,\" unless the cited source (Krzyzanowski et al., 2023) intentionally uses that spelling; please verify and align the usage with the original publication.","section":"Table 6 and Appendix A"},{"comment":"The caption contains a formatting error (\"Fr ´echet\") and the sentence about the recommended data set size \"(5'000)\" is unclear; please state the source of the recommended size and use consistent thousands separators.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The authors are also authors of the PoseBusters benchmark used in the evaluation, which is a potential conflict of interest. Although the thresholds are fixed and public, this relationship should be disclosed explicitly in the manuscript. The paper is a workshop paper (GEM workshop at ICLR 2025) and its current depth is reasonable for a benchmark evaluation, but the internal contradictions and the hydrogen-representation confound must be fixed before archival publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a useful, mostly reproducible validity-aware benchmark of five unconditional 3D molecular generators. The protocol is documented in detail, the code and weights are pinned, and the failure decompositions (chemical vs physical, and the component tables) give a good picture of where each model falls down. The observation that standard benchmarks are saturated while physical/chemical validity is not is important for the generative-chemistry subfield, and the post-processing results (GCDM 95.2%, SemlaFlow 92.4%) are the kind of numbers practitioners will want to see.\n\nThe soft spots are concentrated in the abstract and the raw, no-post-processing ranking. The abstract calls SemlaFlow the best method with 87% and 92.4% with post-processing, but Table 3 shows GCDM at 95.2% after post-processing. That is a direct internal contradiction. Second, the claim that 'all five methods generated large sets' is false for GCDM and GeoLDM raw (0.2% and 2.9%). The stress-test note is right: those numbers are dominated by the explicit-hydrogen requirement—the models output heavy atoms only. Table 4 shows GCDM sanitizes at 97.2% but passes explicit-H at 0.2%. So the raw ranking is confounded by output format, not model chemistry. The paper acknowledges this in Section 3, but it is absent from the abstract and from the sentence 'Given that these training data scores are higher than the best model, there is still room for improvement,' which is misleading once you look at GCDM+PP versus GEOM Drugs. Third, the thresholds (25% bond/angle, 30% clash, energy ratio 100) are acknowledged as generous, and rankings could shift under stricter criteria. Fourth, no variance or error bars are reported despite multiple seeds being run for some methods.\n\nNone of this sinks the central argument. Once hydrogens are added uniformly (post-processing), the comparison is fair, and the conclusion that physical validity is the bottleneck stands. The paper should be revised to fix the abstract, qualify the raw numbers as heavy-atom-only representations, and present the post-processing ranking as the headline. As it stands, it earns a serious referee but not an acceptance without those changes.\n\nIf I were editing, I'd send it to peer review with a request for major revision. The core evaluation is worth publishing; the presentation is not yet reliable enough for the abstract to be quoted as is.","headline":"Useful, mostly reproducible validity-aware benchmark of five 3D generators, undercut by an abstract that contradicts its own post-processing table and a raw ranking confounded by hydrogen representation.","tokens_in":11890,"tokens_out":3973,"would_cite":true,"duration_ms":37303,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the near-saturation of standard 3D molecule-generation benchmarks is misleading: once chemical and physical validity are required, the best raw generator reaches 87.0% valid-unique-novel molecules, and its own…","keywords":["3D molecular generation","unconditional generation","molecular validity","physical validity checks","drug discovery","generative model evaluation","post-processing","conformer quality"],"falsifier":"Re-evaluate the same five models' 100,000 samples under stricter physicality cutoffs, for example 10% bond-geometry deviation, 10% van der Waals overlap allowance, and a strain-energy ratio of 20. If the best methods still score above 90% valid-unique-novel, the paper's conclusion that unconditional 3D generation is not saturated would be called into question; if their scores drop sharply, the conclusion stands but is threshold-dependent.","tokens_in":10892,"feed_emoji":"🧪","tokens_out":10295,"duration_ms":100796,"temperature":0.7,"pith_summary":"Unconditional 3D molecular generation is usually scored by whether the generated molecular graph is chemically valid, unique, and novel; recent papers report near-saturation on those benchmarks. This paper argues that such scores are misleading because they ignore the geometry of the 3D conformation. It evaluates five recent models on 100,000 samples each, requiring both a chemically valid graph and a physically plausible conformation, including bond lengths, bond angles, ring planarity, steric clashes, and internal strain. Under this combined test, the best raw model, SemlaFlow, produces 87.0% valid-unique-novel molecules, and after simple post-processing the best model is GCDM at 95.2%. These numbers sit below the 94.2% scored by the benchmark's own training data, so the paper concludes that unconditional 3D generation still has room to improve.","feed_headline":"Best raw 3D molecule generator hits 87 percent validity","feed_subtitle":"Five-model benchmark shows saturation claims are premature; post-processing lifts one model to 95 percent.","key_machinery":"The carrying object is a two-stage validity filter applied to every generated molecule. The first stage is graph-level: the molecule must parse, pass chemical sanitisation, have all explicit hydrogens, and be connected. The second stage is conformation-level: the 3D structure must pass six geometry and energy checks, including bond lengths and angles within 25% of experimental bounds, planar aromatic rings, planar double bonds, no internal steric clash, and a force-field strain-energy ratio below 100. The paper applies this filter, together with uniqueness and novelty checks based on canonical string identifiers, to 100,000 unconditional samples per model, with and without post-processing. Because some thresholds, notably the 30% van der Waals overlap allowance and the strain ratio of 100, are deliberately generous, passing them is a meaningful but not strict test of physical plausibility.","core_discovery":"Put in the authors' terms, the discovery is that standard benchmarks are saturated but the task is not. More than 99% of the valid molecules generated by the five methods are unique and novel, so diversity is not the bottleneck. The bottleneck is validity once it is defined to include the physical conformation: raw success rates are 59.7% for EQGAT-diff, 59.8% for FlowMol, 0.2% for GCDM, 2.9% for GeoLDM, and 87.5% for SemlaFlow. The two low models fail mainly because they do not generate explicit hydrogens; the other failures are split between chemically invalid graphs and physically invalid conformations. Post-processing, meaning keeping the largest fragment, adding hydrogens, and relaxing the structure with a universal force field, changes the ordering, putting GCDM at 95.2%, SemlaFlow at 92.4%, EQGAT-diff and FlowMol at 84.1%, and GeoLDM at 69.6%. Even the best results remain below the 94.2% validity of the training data and the 98.8% of an approved-drug reference set.","pith_inferences":["Using the geometry and strain checks as a rejection filter or scoring term at sampling time could push several methods above 90% valid-unique-novel, an option the paper does not test.","The hydrogen failure pattern implies that changing the generative target from heavy-atom-only structures to full explicit-hydrogen structures could remove the largest chemical-invalidity source at the root.","Since the paper frames unconditional generation as a stepping stone to conditional tasks, the same physical-validity failures would likely appear in conditional generation; auditing conditional outputs with the same conformer checks would test that transfer.","Rankings by a single aggregate rate hide different failure modes: models fail at different stages, so future evaluations should report graph and conformation validity separately, as the paper's own tables do."],"forward_implications":["The reported valid-unique-novel shares are upper bounds on useful output, since several of the physicality thresholds are acknowledged to be generous.","Uniqueness and novelty are effectively solved for these models, with over 99% of valid molecules being novel and unique, so future work should target chemical and physical validity rather than diversity.","Hydrogen handling is a first-order design decision: models that omit explicit hydrogens jump from near-zero raw validity to high validity once hydrogens are added in post-processing.","Because the training set itself scores 94.2% on the same checks, even the best method leaves room for improvement, contradicting the saturated-benchmark narrative.","Distribution metrics such as drug-likeness and synthetic accessibility can match the training distribution while chemical-space coverage remains incomplete, so those metrics alone cannot certify coverage."],"supporting_citations":[{"why":"Supplies the intramolecular geometry and strain checks used to define physical validity.","marker":"Buttenschoen et al., 2024"},{"why":"Provides the experimentally informed distance-geometry bounds that set the bond length and angle tolerances.","marker":"Riniker & Landrum, 2015"},{"why":"Provides the cheminformatics operations for parsing, sanitisation, explicit hydrogen handling, and canonical string generation.","marker":"Landrum et al., 2024"},{"why":"Supplies the universal force field used for energy minimisation and for computing the strain-energy ratio.","marker":"Rappe et al., 1992"},{"why":"Defines the benchmark training set used for training, novelty checking, and as the main comparison distribution.","marker":"Axelrod & G´omez-Bombarelli, 2022"},{"why":"Describes the SemlaFlow model and reports the saturated benchmark scores that motivate re-evaluation.","marker":"Irwin et al., 2024"},{"why":"Earlier call to add empirical geometry-based validity checks to 3D molecular generation benchmarks.","marker":"Baillif et al., 2023"},{"why":"Supplies the approved-drug reference set used as a second, near-perfect baseline.","marker":"Knox et al., 2024"}],"fun_headline_variants":["Raw 3D molecule gen: 87% best, but task not solved","Benchmarks saturated, validity is the real bottleneck","Post-processing flips 3D molecule gen rankings","SemlaFlow tops raw 3D molecule generation at 87%","3D molecule generation: raw validity ranges 0.2% to 87%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the chosen thresholds, a 25% tolerance on bond geometry, a 30% allowance on van der Waals overlap, and a strain-energy ratio of 100, define what counts as physically valid; if a stricter definition is used, the reported success rates and the ranking of methods could change.","fun_headline_variants_meta":{"raw":{"variants":["Raw 3D molecule gen: 87% best, but task not solved","Benchmarks saturated, validity is the real bottleneck","Post-processing flips 3D molecule gen rankings","SemlaFlow tops raw 3D molecule generation at 87%","3D molecule generation: raw validity ranges 0.2% to 87%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000812,"raw_usage":{"total_tokens":3534,"prompt_tokens":892,"completion_tokens":2642,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":2550}},"tokens_in":508,"tokens_out":2642,"duration_ms":19526,"temperature":1.0,"reasoning_tokens":2550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:39:51.946521+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-evaluate the same five models' 100,000 samples under stricter physicality cutoffs, for example 10% bond-geometry deviation, 10% van der Waals overlap allowance, and a strain-energy ratio of 20. If the best methods still score above 90% valid-unique-novel, the paper's conclusion that unconditional 3D generation is not saturated would be called into question; if their scores drop sharply, the conclusion stands but is threshold-dependent.","supporting_citations":[{"cited_title":"Morris, and Charlotte M","cited_arxiv_id":null,"evidence_quote":"Supplies the intramolecular geometry and strain checks used to define physical validity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the experimentally informed distance-geometry bounds that set the bond length and angle tolerances."},{"cited_title":"Scalfani, Rachel Walker, Kazuya Ujihara, Daniel Probst, Juuso Lehtivarjo, Guillaume Godin, Axel Pahl, Fran c ois Fran c ois B \\'e renger, and Hussein Faara","cited_arxiv_id":null,"evidence_quote":"Provides the cheminformatics operations for parsing, sanitisation, explicit hydrogen handling, and canonical string generation."},{"cited_title":"GEOM , energy-annotated molecular conformations for property prediction and molecular generation","cited_arxiv_id":null,"evidence_quote":"Defines the benchmark training set used for training, novelty checking, and as the main comparison distribution."},{"cited_title":"Efficient 3D molecular generation with flow matching and scale optimal transport","cited_arxiv_id":null,"evidence_quote":"Describes the SemlaFlow model and reports the saturated benchmark scores that motivate re-evaluation."},{"cited_title":"Deep generative models for 3D molecular structure","cited_arxiv_id":null,"evidence_quote":"Earlier call to add empirical geometry-based validity checks to 3D molecular generation benchmarks."},{"cited_title":"DrugBank 6.0: The DrugBank knowledgebase for 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the approved-drug reference set used as a second, near-perfect baseline."}],"review_version":1}