{"id":"9152d943-7af9-4c34-8470-57c4592dc00d","arxiv_id":"2506.23008","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"PubChemQCR is a large public dataset of DFT-based molecular relaxation trajectories with energy and force labels, benchmarked with nine machine learning interatomic potentials.","lead":"Researchers release PubChemQCR, a dataset of about 3.5 million molecular relaxation paths and over 300 million quantum-chemistry energy and force labels, giving AI chemistry models a large new training ground. They benchmark nine machine-learned force fields and find Equiformer performs best on the smaller subset. If reliable, the resource could speed up molecular simulation by replacing slow DFT calculations with cheaper surrogates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dataset's central value rests on unvalidated label parsing: no independent check of energies, forces, or stage assignment is reported, and Table 1's 298.75M snapshots do not support the abstract's 'over 300 million' claim.","rationale":"The reader's weakest assumption and my primary concern are the same: label correctness after parsing. I agree with the reader that this is the load-bearing issue. The paper's own text gives the curation pipeline in Section 3.3 but provides no validation step; the benchmark numbers would be meaningless if the labels carry a systematic error, and the dataset's reuse value depends on those labels. I also note the concrete count discrepancy in Table 1 versus the abstract, which is a definite factual issue though secondary to label validity. I do not see an internal inconsistency in the benchmark methodology severe enough to reject; the random split by trajectory is appropriate, and the limitations section honestly flags the near-equilibrium coverage. Therefore the appropriate verdict remains CONDITIONAL: accept the resource as described only after the label-validation check is performed and the 'over 300 million' wording is corrected. This does not change the reader's verdict.","tokens_in":15317,"tokens_out":6821,"duration_ms":73741,"concrete_test":"Take 100 CIDs from PubChemQCR-S, fetch their raw PubChemQC log files, and independently parse them with a separate script written from the GAMESS/Firefly format documentation. Verify that the LMDB entries match in atomic numbers, coordinates, total energy, forces, and stage assignment for every optimization step. Then, for 20 snapshots spanning PM3, HF, DFT_1st, and DFT_2nd stages, rerun single-point calculations at the same level of theory (PM3, HF/STO-6G, B3LYP/6-31G*) in an independent quantum chemistry package and compare total energies to ~1e-3 Hartree and each force component to ~0.01 eV/Å, including the identity F = -∂E/∂x. Any systematic mismatch—especially a force sign flip or unit-conversion offset—would show the dataset labels are not reliable for the claimed MLIP training use.","verdict_should_be":"UNCHANGED","load_bearing_attack":"PubChemQCR's central claim—that it is a large, reusable source of energy and force labels for MLIP training—requires that the parsed labels from the 7 TB of raw logs be correct in units, sign, and stage assignment. Section 3.3 describes removing failed logs and duplicates but reports no validation of the parsed values against either the original calculation or an independent quantum chemistry code. The risk is concrete: force labels in quantum-chemistry outputs are often printed as gradients, so a missing sign flip would produce forces pointing uphill; a Hartree/eV conversion error would shift energies; and a stage misassignment would mix PM3, HF, and DFT labels. Tables 3–5 report small errors (energy MAE ≈ 2–10 meV/atom, force RMSE ≈ 24–61 meV/Å), so any such systematic parser error would be invisible in the aggregate distributions and would propagate into every baseline. The paper's own Section 6 concedes the DFT-stage data are near-equilibrium with small forces, which makes sign errors especially hard to detect in training. A secondary, definite inconsistency: Table 1 sums to 298,751,667 snapshots, so the abstract's and Section 7's 'over 300 million conformations' is not supported by the reported counts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PubChemQCR, a dataset of molecular geometry optimization trajectories curated from roughly 7 TB of raw PubChemQC log files. It reports approximately 3.5 million trajectories and 298,751,667 snapshots across PM3, Hartree–Fock, and two DFT substages, with each snapshot labeled by total energy and atomic forces. The authors benchmark nine machine learning interatomic potentials on a smaller subset (PubChemQCR-S) and two models on the full dataset, and they evaluate geometry optimization performance. The central claim is that PubChemQCR is the largest publicly available dataset of DFT-based relaxation trajectories for small organic molecules and is suitable for training and evaluating MLIPs.","tokens_in":15555,"tokens_out":5137,"duration_ms":50347,"significance":"If the curation is correct, PubChemQCR is a valuable and reusable resource: it is large, publicly released with a custom PyTorch Geometric dataloader, organized into trajectory-level splits, and includes both energy and force labels. The benchmark provides a useful reference point, and the limitations section is candid about the near-equilibrium bias of the DFT-stage data. However, the manuscript's central value depends on the fidelity of the parsed labels, and the absence of any independent validation of energies, forces, and stage assignment is a significant correctness risk. There is also a concrete internal inconsistency in the headline snapshot count. These issues are fixable within the scope of the paper, but they need to be addressed before the resource can be relied upon.","major_comments":[{"comment":"The core value of PubChemQCR rests on parsing 7 TB of raw log files into atomic numbers, coordinates, energies, and forces, but no validation of these parsed labels is reported. A systematic error in units (Hartree vs eV), force sign (gradient vs force), or stage assignment (PM3/HF/DFT_1st/DFT_2nd) would silently propagate into every baseline in Tables 3–5. Please report quantitative checks: for example, recompute a random subset of snapshots with an independent DFT code and compare energies and forces, verify force directions and magnitudes at stationary points, and confirm that atomic numbers and stage labels match the original PubChemQC records. This validation is load-bearing for the central claim that the dataset provides high-quality energy and force labels.","section":"Section 3.3"},{"comment":"Table 1 sums to 298,751,667 snapshots, yet the abstract and Section 7 claim 'over 300 million conformations.' This is an internal numerical inconsistency in a headline quantity. The text should be changed to 'approximately 298.8 million' or the counts should be revised so that the reported total actually exceeds 300 million.","section":"Table 1 / Abstract / Section 7"},{"comment":"The full-dataset benchmarks are said to use 'data from the DFT optimization stage,' which in Section 3.4 includes both DFT_1st (Firefly/SMASH) and DFT_2nd (GAMESS) substages. Section 6 explicitly states that labels from different stages of the same trajectory are 'neither directly comparable nor mutually consistent.' If Table 4 pools DFT_1st and DFT_2nd labels, the training targets are heterogeneous; please either report results separately for each substage or provide a clear justification for pooling them. This affects the interpretation of the full-dataset results.","section":"Section 4.1 / Table 4"},{"comment":"The benchmark tables report single runs without standard deviations over training seeds. Given the small gaps in Table 3 (e.g., test energy MAE of 5.38 meV/atom for Equiformer versus 5.33 for PaiNN), seed variability could change the reported rankings. Please report mean and standard deviation over at least three seeds for the main energy/force and geometry optimization metrics, or otherwise demonstrate that the differences are statistically meaningful.","section":"Tables 3–5"}],"minor_comments":[{"comment":"The text contains a typo: 'NeuqIP' should be 'NequIP' in the Results and Discussions paragraph.","section":"Section 4.1"},{"comment":"The learning-rate scheduler is described as 'REDUCE LRO NPLATEAU'; this should be formatted as 'ReduceLROnPlateau' for readability.","section":"Section 4.1"},{"comment":"The 'largest publicly available dataset' claim would be easier to verify if the paper gave a direct quantitative comparison with the OMol25 trajectory counts under the same scope definition (small organic molecules), rather than listing OMol25 properties only descriptively.","section":"Section 2"},{"comment":"The limitation about near-equilibrium forces is candid and important; it would be helpful to add a short discussion of how future extensions might include off-equilibrium sampling, since this is the main regime where MLIPs are expected to be used.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a dataset/benchmark contribution rather than a new method. The central claim is empirical and plausible, and the curation pipeline is described in enough detail to be reproducible. My main concern is that the absence of label validation is a correctness risk for the entire resource; the count inconsistency and seed-variance issues are fixable. I would recommend major revision rather than rejection, because the authors can address the validation requirement within the scope of the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real product here is PubChemQCR, not the benchmark tables. The authors have parsed the raw PubChemQC geometry optimization logs into per-step energies and forces for roughly 3.5 million trajectories and 299 million snapshots, and they are releasing that. As far as I know this is the first release of the actual relaxation trajectories from PubChemQC, not just the ground-state geometries that Molecule3D and PCQM4Mv2 used. For MLIP training, having off-equilibrium snapshots along realistic optimization paths is valuable, and the scale is genuinely larger than GEOM or ANI-1x. The curation pipeline is described in enough detail to be credible, and the splits are trajectory-level, which is the right call.\n\nThe paper's own limitations section is candid: the DFT-stage data are near-equilibrium, so the abstract's off-equilibrium framing overstates what the DFT portion offers. That is a scope issue, not a fatal one, since the PM3 and HF stages do contain more distorted geometries, even if at lower theory levels. I would have liked a clearer statement that the benchmark uses only the DFT stage and that DFT_1st and DFT_2nd are pooled in the full-dataset runs; mixing Firefly/SMASH and GAMESS labels is a subtle inconsistency.\n\nThe bigger soft spot, and it is real, is that there is no independent validation of the parsed labels. The whole value of the dataset rests on energies, forces, and stage assignments being correct in units and sign. The paper says they removed failed logs and duplicates but does not report recomputing a random subset with DFT or otherwise checking against a known reference. A sign error in forces or a Hartree/eV conversion error would be invisible in the aggregate distributions and would contaminate every benchmark. This is addressable: sample a few hundred trajectories, recompute, and report the comparison. Until then, treat the dataset with caution.\n\nMinor: Table 1 sums to 298,751,667, so 'over 300 million conformations' in the abstract and Section 7 is not supported by the paper's own counts. And the benchmark tables have no seed-to-seed variance, so small differences between models are hard to interpret.\n\nWho benefits: MLIP researchers who need large-scale trajectory data for training and evaluation. I would send this to review, but with the label-validation request as a major revision. It deserves referee time; the resource is likely to be widely used.","headline":"A genuinely useful large-scale relaxation-trajectory dataset, but the missing label validation and a few overclaims need fixing before I'd trust the numbers.","tokens_in":16126,"tokens_out":2225,"would_cite":true,"duration_ms":22361,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper turns 7 TB of raw PubChemQC geometry-optimization logs into PubChemQCR, a benchmark of about 3.5 million relaxation trajectories and 300 million labeled conformations, and benchmarks nine interatomic-potential models on it.","keywords":["machine learning interatomic potentials","quantum chemistry dataset","DFT relaxation trajectories","energy and force labels","PubChemQCR","molecular conformations","geometry optimization benchmark","small organic molecules"],"falsifier":"Recompute total energies and atomic forces with B3LYP/6-31G* for a random sample of a few hundred parsed snapshots (or check that stored forces equal negative gradients of stored energies by finite differences); a systematic mismatch in energy offsets, force magnitudes, or signs would falsify the curation claim.","tokens_in":15122,"feed_emoji":"⚛️","tokens_out":8218,"duration_ms":72855,"temperature":0.7,"pith_summary":"This paper argues that machine-learned interatomic potentials need training and evaluation data that include the off-equilibrium geometries encountered during relaxation, not just final stable structures, and it builds the missing resource at scale. The authors curate the raw geometry-optimization logs of the PubChemQC project into PubChemQCR, which they state is the largest publicly available dataset of DFT-based relaxation trajectories for small organic molecules, with roughly 3.5 million trajectories and over 300 million snapshots, each labeled with total energy and atomic forces. They also provide baseline results for nine machine-learning interatomic potentials on a smaller subset and for two of those models on the full dataset, together with tools for loading the data. A sympathetic reader would take away that trajectory-level supervision is now available in a form that can be used to train and compare MLIPs directly.","feed_headline":"3.5M DFT relaxation paths open to train AI atomic models","feed_subtitle":"PubChemQCR labels 300M conformations with DFT energies and forces, plus nine baseline models to beat.","key_machinery":"The carrying object is PubChemQCR itself, a dataset of relaxation trajectories built by parsing raw PubChemQC log files into memory-mapped database records keyed by PubChem compound ID. Its structure matters: each trajectory is a sequence of snapshots containing atomic numbers, Cartesian coordinates, total energy, and force gradients, organized by the four optimization stages, so a model can be trained on just the DFT-stage portion where labels are most accurate. The paper also computes isolated atomic energies at B3LYP/6-31G* and subtracts them to form energy targets, which centers the energy distribution and makes the regression task easier.","core_discovery":"On its own terms, the paper's central claim is that a careful parsing of the 7 TB of PubChemQC optimization outputs yields a reusable benchmark: approximately 3.5 million relaxation trajectories and 298.75 million molecular snapshots, of which 105.49 million come from the DFT stage, spanning 25 chemical elements. Energy and force labels are stored per snapshot, grouped by optimization stage (PM3, Hartree–Fock, DFT via Firefly/SMASH, and DFT via GAMESS), and formation energies are defined by subtracting isolated atomic energies so the learning target is compactly distributed. On the benchmark, Equiformer achieves the best energy and force errors on the 40,979-trajectory subset, and PaiNN performs best when trained on the full dataset; geometry-optimization experiments show that most models struggle to drive near-equilibrium structures to convergence, with Equiformer the clear outlier. The paper further claims the dataset is the largest of its kind and that its trajectory-level splits avoid information leakage.","pith_inferences":["If the labels survive independent validation, this corpus could serve as a supervised pretraining source for molecular property prediction, since it couples geometry with physical energy and force supervision.","A natural stress test is cross-stage training: train on DFT-stage labels and evaluate on PM3/HF snapshots; the paper's own limitation section predicts large inconsistency, and quantifying it would tell users how much of the multi-level data is usable.","The 25-element coverage and near-equilibrium force distribution mean the dataset is not by itself a substitute for active learning or MD sampling that must explore high-force regions; pairing it with on-the-fly DFT queries is the obvious next step.","Because forces are stored per snapshot, PubChemQCR could be used to test energy-conservation and equivariance properties of MLIP architectures on far more diverse molecules than in earlier MD datasets."],"forward_implications":["Training on PubChemQCR exposes MLIPs to intermediate, non-equilibrium conformations, not only relaxed minima, which is the regime molecular dynamics actually samples.","The benchmark numbers establish a reference point: Equiformer leads on the small subset, while PaiNN leads on the full dataset, so later models can be compared against these numbers.","Trajectory-level data splits mean a model is evaluated on molecules whose relaxation paths were unseen during training, testing generalization rather than memorization.","The formation-energy normalization removes per-atom offsets and is reported to speed convergence of the learned energy model.","Geometry optimization results quantify a remaining gap: from near-equilibrium starting points, only Equiformer reaches chemical accuracy on a substantial fraction of test molecules."],"supporting_citations":[{"why":"Supplies the entire raw corpus: the PubChemQC database of first-principles geometry optimization outputs that this dataset is parsed from.","marker":"[Nakata and Shimazaki, 2017]"},{"why":"Molecule3D, the earlier curation of PubChemQC ground-state geometries, defines the provenance and format the new trajectory curation extends.","marker":"[Xu et al., 2021]"},{"why":"QM9 is the canonical small-molecule quantum chemistry dataset whose limitations (one conformer per molecule, no forces) motivate the need for trajectories.","marker":"[Ramakrishnan et al., 2014]"},{"why":"ANI-1x supplies millions of conformations with forces but only four atom types, one of the coverage gaps PubChemQCR targets.","marker":"[Smith et al., 2020]"},{"why":"GEOM provides large conformational diversity but mostly semi-empirical labels without forces, another comparison point.","marker":"[Axelrod and Gomez-Bombarelli, 2022]"},{"why":"OC20 contributes relaxation trajectories for adsorbate–catalyst systems, showing the trajectory format exists but not for small organic molecules.","marker":"[Chanussot et al., 2021]"},{"why":"MPTrj offers material optimization trajectories and is the closest prior trajectory dataset, but for inorganic materials rather than molecules.","marker":"[Deng et al., 2023]"},{"why":"OMol25 is a recent large molecular dataset with relaxation trajectories, defining the contemporary baseline this work must be distinguished from.","marker":"[Levine et al., 2025]"},{"why":"SchNet is one of the nine benchmarked interatomic potentials and sets the invariant-model baseline.","marker":"[Schütt et al., 2018]"},{"why":"Equiformer is the equivariant-transformer model that wins the benchmark on the small subset and the geometry-optimization task.","marker":"[Liao and Smidt, 2022]"}],"fun_headline_variants":["Largest DFT relaxation dataset: 3.5M trajectories for MLIP training","3.5M DFT trajectories power new quantum chemistry benchmark","Equiformer tops new 3.5M-trajectory quantum chemistry benchmark","New benchmark: 3.5M DFT relaxations with energy and force labels","Quantum chemistry relaxations: 3.5M trajectories, 300M conformations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's validity rests on the assumption that parsing 7 TB of raw log files into atomic numbers, coordinates, total energy, and forces is error-free in units, signs, and stage assignment, and the paper reports no independent check of the parsed labels against recomputed DFT values.","fun_headline_variants_meta":{"raw":{"variants":["Largest DFT relaxation dataset: 3.5M trajectories for MLIP training","3.5M DFT trajectories power new quantum chemistry benchmark","Equiformer tops new 3.5M-trajectory quantum chemistry benchmark","New benchmark: 3.5M DFT relaxations with energy and force labels","Quantum chemistry relaxations: 3.5M trajectories, 300M conformations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000403,"raw_usage":{"total_tokens":2126,"prompt_tokens":999,"completion_tokens":1127,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":615,"completion_tokens_details":{"reasoning_tokens":1025}},"tokens_in":615,"tokens_out":1127,"duration_ms":9178,"temperature":1.0,"reasoning_tokens":1025,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:52:47.370457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute total energies and atomic forces with B3LYP/6-31G* for a random sample of a few hundred parsed snapshots (or check that stored forces equal negative gradients of stored energies by finite differences); a systematic mismatch in energy offsets, force magnitudes, or signs would falsify the curation claim.","supporting_citations":[],"review_version":1}