{"id":"cd31816a-88f5-4e00-b111-ba11629eb13d","arxiv_id":"2502.07335","paper_version":2,"verdict":"UNVERDICTED","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of machine learning interatomic potentials that organizes the field by descriptor type, message-passing architecture, long-range corrections, and universal models, with open challenges.","lead":"This paper is a review of machine learning potentials, computer models that learn how atoms move by fitting to quantum chemistry data. It maps the field's evolution from early fits for small molecules to modern universal models that claim to work across many materials.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Review relies on non-uniform, self-cited comparisons and contains an internal inconsistency (GNoME training-data size in Table II vs text); the central narrative is not independently verified.","rationale":"The reader's verdict (UNVERDICTED) treats the paper as a narrative review with no new data, and identifies as the weakest assumption that the cited literature is accurately represented and the examples are representative. My stress-test agrees that this assumption is load-bearing, and I found a concrete instance of its failure: the GNoME training-data inconsistency between Table II and the text, plus a non-uniform reference basis in Figure 5. These do not invalidate the entire review, but they show that the review's factual scaffolding is not as secure as it should be. I do not move the verdict because the central claim is a descriptive overview rather than a falsifiable scientific result, and the issues, while real, are correctable and do not overturn the overall narrative. The reader's UNVERDICTED stance already captures the epistemic status appropriately. If the authors correct the table and provide evidence that Figure 5's comparisons are controlled, the review would be more reliable, but as is, the verdict should remain unchanged.","tokens_in":45611,"tokens_out":9341,"duration_ms":85810,"concrete_test":"Audit Table II row-by-row against the cited original papers, starting with the GNoME entry: check whether the GNoME paper (Nature 2023) reports a training-set size of 89 million structures; if it does, the '-' in Table II is an error. Then verify every other row's model size, data size, and element count against its source. If any row mismatches, the review's summary data are unreliable and the central overview is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is to provide an accurate overview of the evolution of MLPs and current state-of-the-art. That claim depends on the reliability of its summary tables and on the controlled nature of its comparative figures. Two concrete weaknesses undermine this. First, Figure 5 compares EquiREANN against SchNet, REANN, and sGDML on cumulene energy profiles, but the caption states that the reference data for panels (b) and (c) were generated with different electronic-structure methods (DFT vs MNDO). Because the models are compared on different reference levels across molecules, the qualitative ordering shown is not a controlled comparison, and the figure cannot support the broad conclusion that equivariant MPNNs generally outperform invariant ones for such systems. Second, Table II is meant to summarize training data sizes for universal potentials, yet the GNoME row lists '-' while the main text explicitly states GNoME was trained on '89 million inorganic crystal structures'. This internal inconsistency indicates that the review's factual summaries have not been carefully verified. Since the review's utility lies in guiding readers about which methods to use and what trends matter, these unverified and inconsistent elements weaken the central claim, even though no single computational result in the review is overturned.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review traces the evolution of machine learning potentials (MLPs) from early global descriptor-based fits to modern local descriptor models, message-passing neural networks (including equivariant variants), long-range corrections, and universal potentials. It surveys representative applications in gas-phase reactions, gas-surface dynamics, condensed phases, heterogeneous catalysis, energy materials, and biomolecules, and it compiles software packages and universal-potential models in two tables. The paper's central claim is that MLP development has followed a clear progression from global to local representations, from invariant to equivariant features, and from specialized to general-purpose models.","tokens_in":45788,"tokens_out":5560,"duration_ms":48142,"significance":"If its factual content is reliable, the review is a useful entry point for researchers seeking to navigate the field: it provides a coherent taxonomy of methodological families, a broad table of software packages with links, a second table of universal potentials with model sizes and data sizes, and an informed discussion of open challenges such as many-body completeness, long-range electrostatics, and the data requirements for universal models. The review does not contain new derivations or benchmarks, but it does reproduce learning curves and comparisons from primary sources, which is appropriate for a review when properly credited. The explicit compilation of training-data sizes and architectures for universal potentials is a distinctive feature that makes the paper a practical reference.","major_comments":[{"comment":"The claim that EquiREANN captures subtle torsional energy variations in cumulenes that are 'difficult to be accurately captured by invariant MPNNs such as SchNet, REANN, and even sGDML' rests on Figure 5, whose panels (b) and (c) use reference data generated at different electronic-structure levels (DFT and MNDO). The caption asserts this does not affect the comparison of the energy trend, but this assertion is not self-evident; MNDO is a semi-empirical method and may not reproduce the same torsional barrier shapes as DFT. Because the figure is used to support a general message about the advantage of equivariant MPNNs for nonlocal pi-systems, the authors should either restrict the claim to the specific DFT-referenced cases, provide a consistent reference level across all panels, or explicitly discuss why the mixed references do not alter the qualitative ordering.","section":"Section III.C, Figure 5"},{"comment":"The GNoME row in Table II lists the training data size as '-' and the training set as 'MP, OQMD, WBM', while the main text states that GNoME was trained on '89 million inorganic crystal structures.' This is an internal inconsistency in one of the paper's central summary tables; the table should be corrected to include the 89 million number (or the text amended) so that readers can rely on the table as an accurate comparison of universal potentials.","section":"Section IV, Table II"}],"minor_comments":[{"comment":"The sentence 'This was perhaps the earliest scheme of active learning' should be softened to 'one of the earliest examples' or supported by a specific citation, because without this hedge the historical claim is difficult to verify.","section":"Section II"},{"comment":"The sentence 'These three-body feature-based MPNNs, such as REANN and SpookyNet, significantly outperformed ... on a representative CH4 dataset' would benefit from a specific reference to the figure or table in Ref. 163 that supports the quantitative comparison.","section":"Section III.C"},{"comment":"The claim that the sGDML double-walled nanotube is 'the largest molecule studied to date using global descriptor-based methods' should include a 'to our knowledge' qualifier and a date, since this is a fast-moving area.","section":"Section IV"},{"comment":"The phrase 'end-ot-end manner' appears to be a typo for 'end-to-end manner.'","section":"Section III.A"},{"comment":"The sentence 'Atomistic MLP methods have made significant successes in simulating extended systems' is awkward; consider 'have achieved significant successes.'","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The review is competent and useful, but it carries a noticeable self-citation burden: PIP-NN, EANN, REANN, EquiREANN, and several related applications are the authors' own works and are given prominent placement, including Figure 5. This is not improper, and the works are relevant, but the authors should be encouraged to balance the narrative by explicitly noting the provenance of comparative claims. The manuscript is also long and dense; if the editors allow, a short 'take-home messages' box would help. Overall, the paper is within scope for the journal and likely to be cited widely once the factual inconsistencies are fixed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you my read on the Xia–Zhang–Jiang review. It is exactly what it says: an overview of MLP evolution. No new data, no benchmarks. That is fine for what it is. It is well organized, current through late 2024, and covers the main families (global descriptor, local descriptor, MPNN, equivariant, universal potentials) with good references. As a newcomer map, it serves its purpose.\n\nThe soft spots are real but not disqualifying. The stress-test note about Figure 5 is partially right: panels (b) and (c) come from different electronic-structure references (DFT vs MNDO), so cross-panel rankings do not prove that EquiREANN is generally better. Within each panel, though, the models are compared against the same reference, and the claim in the text is specifically about these cumulene cases, not a universal statement. So I'd call that a caveat, not an error. The bigger issue is Table II: the GNoME row lists '-' for data size while the text says 89 million inorganic crystal structures. That is an internal inconsistency in a table meant to be a quick reference. It is minor, but a review's summary tables need to be accurate. The reader also notes heavy self-citation in the model comparisons (PIP-NN, EANN, REANN, EquiREANN). That is common in reviews where the authors are active contributors, and the cited results are real papers, but it does tilt the narrative toward their own models.\n\nThe non-uniform selection of representative applications is acknowledged as 'selected,' not systematic, so I won't hold that against it. The method descriptions I sampled are consistent with the primary literature. No invented entities or circular derivations.\n\nWho gets value: graduate students and experimentalists wanting a readable map of MLP families and universal potentials. It is not a benchmark paper and should not be cited as one. I would bring it to a reading group maybe, and would cite it as a recent overview if needed. It should go to peer review, despite the minor table slip—reviews like this need referees to catch exactly those errors.","headline":"A solid, current review of MLPs that is useful as an entry point but has a few factual table slips and a comparison figure that is less controlled than it looks.","tokens_in":46305,"tokens_out":2260,"would_cite":true,"duration_ms":20467,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of two decades of machine learning potentials charts the shift from global descriptor fits to local descriptors, equivariant message-passing networks, and universal pretrained models.","keywords":["machine learning potentials","potential energy surfaces","equivariant neural networks","message-passing neural networks","universal potentials","molecular dynamics","reactive scattering","atomistic simulations"],"falsifier":"An independent benchmark that re-trains one representative model from each era on identical datasets and finds that a local-descriptor model matches or beats equivariant models in force accuracy would directly contradict the review's claimed progression.","tokens_in":45378,"feed_emoji":"🧪","tokens_out":9032,"duration_ms":80747,"temperature":0.7,"pith_summary":"This review aims to chart the development of machine learning potentials (MLPs), functions that map nuclear coordinates to potential energy by fitting discrete quantum-chemical data, from their first applications in the 1990s to the present. Its central claim is that the field has followed a recognizable trajectory: global descriptor fits for small molecules gave way to atom-centered local descriptors, then to learnable message-passing features and equivariant tensor operations, and most recently to large universal potentials pretrained across broad chemical space. A sympathetic reader would take away a structured map of the design space and an explicit list of unresolved problems, above all the treatment of long-range interactions and the completeness of atomic descriptors. The authors also select representative applications in spectroscopy, gas-surface dynamics, condensed-phase chemistry, catalysis, and biomolecules to show which methods have become standard in each regime.","feed_headline":"Review maps two decades of machine-learned potentials","feed_subtitle":"From symmetry functions to equivariant networks to universal models, this timeline shows where the field stands.","key_machinery":"The organizing machinery is the choice of atomic representation. The review classifies models by how they build symmetry-invariant structure descriptors: global polynomial or kernel maps of all internuclear distances (PIP, FI-NN, GDML); local atom-centered descriptors such as ACSFs, SOAP, and moment or atomic-cluster expansions; and learned features from message-passing neural networks, including equivariant tensors built from spherical harmonics and Clebsch-Gordan couplings (NequIP, MACE, EquiREANN). A second mechanism is range separation, where total energy is split into a short-range learned term plus explicit electrostatics, dispersion, or charge-equilibration terms to capture long-range physics. These two mechanisms, representation and range separation, carry the narrative and structure the authors' comparison tables and timeline.","core_discovery":"The authors' core assertion is that the past two decades of MLP research can be organized as a series of architectural discoveries about how to encode atomic environments. The breakthrough of local decomposition, writing total energy as a sum of atomic energies each depending on a symmetry-preserving descriptor of the surrounding atoms, made high-dimensional and periodic systems tractable. Message-passing networks then replaced fixed descriptors with features learned by repeatedly exchanging information between neighbors, and equivariant networks added explicit rotational tensors to gain data efficiency. The review closes with universal potentials trained on datasets with tens of millions of structures, arguing that these are the emerging frontier even though they are not yet validated for reactive chemistry. Throughout, the paper claims that no single architecture dominates: global-descriptor models remain best for small, high-accuracy spectroscopy and reaction dynamics, while local and equivariant models dominate extended systems.","pith_inferences":["A consequence the authors leave implicit is that if data efficiency is the limiting resource, the practical question is not global versus local versus equivariant architecture, but how each architecture behaves under active learning on a fixed ab initio budget; the review's comparisons do not settle this.","The completeness failures of low-body-order descriptors suggest that test suites built from deliberately 'pathological' geometries would provide a sharper test of equivariant networks than the standard benchmarks.","The universal-potential trend points toward a future of foundation models fine-tuned per system; the review's own open challenges imply that hybrid designs with a pretrained backbone plus a physical long-range correction are a plausible next step."],"forward_implications":["Small-molecule and reaction-dynamics studies should continue to favor global-descriptor fits such as PIP and FI-NN, which deliver spectroscopic accuracy with far less data than local models.","For extended materials, biomolecules, and heterogeneous interfaces, equivariant message-passing models are positioned as the default because they combine accuracy with data efficiency.","Long-range interactions remain the main structural weakness of local and message-passing potentials, so range-separated schemes and charge-equilibration networks are the practical remedies until a cheaper complete representation appears.","Universal potentials are a real trend, but the review expects their current coverage of crystals and equilibrium structures to be insufficient for reactive and non-equilibrium chemistry, so fine-tuning and broader sampling will be needed."],"supporting_citations":[{"why":"It introduces the atom-centered symmetry function scheme that anchors the local-descriptor lineage.","marker":"82"},{"why":"It establishes the Gaussian approximation potential, the kernel-based branch of local descriptors.","marker":"84"},{"why":"It defines the smooth overlap of atomic positions descriptor used across local, kernel, and universal models.","marker":"144"},{"why":"It provides the symmetric gradient-domain model that keeps global descriptors competitive for flexible molecules.","marker":"86"},{"why":"It presents an early message-passing potential whose learned features replaced hand-coded descriptors.","marker":"156"},{"why":"It sets out the atomic cluster expansion, a systematic many-body descriptor later combined with equivariant message passing.","marker":"148"},{"why":"It introduces an E(3)-equivariant message-passing model whose data efficiency supports the recent accuracy claims.","marker":"166"},{"why":"It combines high body order with equivariant message passing to reach state-of-the-art accuracy in the review's benchmarks.","marker":"174"},{"why":"It supplies a universal-potential example trained on millions of relaxation geometries, illustrating the final trend.","marker":"265"}],"fun_headline_variants":["MLP evolution: from symmetry functions to universal models","No single architecture dominates: MLP landscape mapped","20 years of MLPs: local, global, and equivariant approaches","Universal potentials emerge, but reactive chemistry awaits","How machine-learned potentials went from niche to universal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusions depend on the cited studies being representative and accurately summarized; none of the model comparisons it highlights are independently reproduced in the review itself.","fun_headline_variants_meta":{"raw":{"variants":["MLP evolution: from symmetry functions to universal models","No single architecture dominates: MLP landscape mapped","20 years of MLPs: local, global, and equivariant approaches","Universal potentials emerge, but reactive chemistry awaits","How machine-learned potentials went from niche to universal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1495,"prompt_tokens":825,"completion_tokens":670,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":593}},"tokens_in":441,"tokens_out":670,"duration_ms":5843,"temperature":1.0,"reasoning_tokens":593,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T13:04:02.430478+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent benchmark that re-trains one representative model from each era on identical datasets and finds that a local-descriptor model matches or beats equivariant models in force accuracy would directly contradict the review's claimed progression.","supporting_citations":[],"review_version":1}