{"id":"f1bb5d62-c3fc-4724-b903-6535f69e595a","arxiv_id":"2508.15614","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Benchmarking 11 universal machine learning interatomic potentials on a new 40,000-structure, 0D-3D dataset shows energy and geometry errors grow as dimensionality falls, with eSEN the most transferable.","lead":"A benchmark of 11 universal machine learning interatomic potentials on 40,000 materials, spanning molecules, wires, sheets, and bulk crystals, shows accuracy drops as dimensionality decreases. One model, eSEN, holds errors near DFT quality across all dimensionalities, pointing toward cheap simulation of realistic surfaces and interfaces.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline error statistics are conditional on successful relaxations; excluding 620 eqV2 0D failures and other fragmented systems inflates the 'across all dimensionalities' accuracy claim.","rationale":"The reader's weakest assumption correctly identifies that excluding failed and fragmented relaxations undermines the unconditional accuracy claim and the model ranking. This is the most load-bearing concern because it directly bears on the paper's central promise: that these uMLIPs can replace DFT for geometry relaxations across dimensionalities. If a substantial fraction of systems fail to relax, reporting only the error of the successes does not establish replacement-level reliability. The proposed test is cheap and decisive: recompute all headline metrics and rankings with failures included. I do not see a reason to move the verdict beyond CONDITIONAL; the resource (0123D dataset) and the qualitative trend are still valuable, and the issue is addressable. The ORB-2-based test-set selection is a related but secondary concern; it is more expensive to test and does not invalidate the main conditional comparison as directly as censored failures do.","tokens_in":13098,"tokens_out":4729,"duration_ms":55602,"concrete_test":"Re-run the benchmark including all failures: for each model and each of the 40,000 systems, assign unconverged/fragmented relaxations an error cap at the 99th percentile of converged position/energy errors (or treat them as infinite errors), then recompute the per-dimensionality MAE, the fraction of systems with energy error <10 meV/atom and position MAE <0.02 Å, and the Section II.D ranking. If eqV2 drops below eSEN/ORB-2 on 0D, or if no model achieves '<10 meV/atom across all dimensionalities' when failures are included, the abstract and conclusion must be revised to say 'conditional on successful convergence.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—'errors in the atomic positions in the range of 0.01–0.02 Å and errors in the energy below 10 meV/atom across all dimensionalities'—is computed only for relaxations that converged and did not fragment. Section III.B explicitly discards 2D slabs thicker than 7.5 Å, 1D systems wider than 12.5 Å, and 0D systems with diameter larger than 20 Å; relaxations are also labeled unconverged if they exceed 15,000 steps or forces exceed 10,000 eV/Å. Table II shows these deletions are not negligible: eqV2 fails 620/10,000 0D systems, ORB-3d fails 89, and even eSEN has 32 2D failures. Figures 5–6 and the statement in Section II.D that ORB-2, eqV2, and eSEN are the best geometry models are therefore conditional on the survivors. If failures were counted as errors, eqV2's 0D position error would be effectively unbounded and the ranking—and the 'sufficient accuracy to serve as direct replacements for DFT' conclusion—could change. The eSEN 'more than 75% of energy predictions below 10 meV/atom' statistic is likewise conditional on successfully converged structures. The headline as written overstates unconditional accuracy, because a meaningful fraction of low-dimensional systems cannot be relaxed by some of the recommended models.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a benchmark dataset (0123D) of 10,000 PBE-relaxed compounds per dimensionality (0D, 1D, 2D, 3D) and evaluates 11 universal machine learning interatomic potentials (uMLIPs) by re-relaxing these structures and comparing final energies and Kabsch-aligned atomic positions. The central claims are that the best models—particularly eSEN, with ORB-2 and eqV2 for geometries—reach errors of 0.01–0.02 Å in positions and below 10 meV/atom in energies across all dimensionalities, and are therefore accurate enough to replace DFT for near-equilibrium simulations of mixed-dimensional systems.","tokens_in":13391,"tokens_out":4344,"duration_ms":51551,"significance":"If the claims hold, this is a useful contribution: it provides a consistent PBE-level benchmark spanning dimensionalities, uses a sensible geometry metric, and reports code and data availability. The comparison of 11 uMLIPs, including non-conservative and conservative architectures, is timely and valuable. However, the central quantitative claims rest on error statistics computed only for successfully converged relaxations, which weakens the 'across all dimensionalities' headline. The dataset construction also depends on ORB-2 for pre-relaxation and hull filtering, which may bias rankings. These issues are fixable and do not invalidate the dataset itself, but they require revision before the paper's conclusions can be accepted.","major_comments":[{"comment":"The headline error metrics are conditional on relaxations that converged without fragmentation. Table II shows that eqV2 fails on 620 of 10,000 0D systems and ORB-3d on 89, yet these models are ranked among the 'best performing models for geometry optimization' in §II.D using error distributions that exclude these failures. A failed relaxation represents an unbounded geometry error, so the ranking and the abstract's '0.01–0.02 Å / below 10 meV/atom across all dimensionalities' overstate unconditional performance. Please report failure counts as a primary metric, or include failures as unrelaxed/infinite error, and revise the conclusions accordingly.","section":"§III.B, Table II, Figs. 5–6, §II.D"},{"comment":"For the 0D–2D subsets, candidate structures were pre-relaxed with ORB-2 and selected by distance to the convex hull computed using ORB-2 energies. ORB-2 is then one of the models evaluated and ranked among the best. This selection may favor structures on which ORB-2 has an advantage, potentially inflating its rank relative to models not used in dataset construction. Please test sensitivity, e.g., by constructing an independent subset with a different selector or by reporting rankings on the unfiltered random structures, or explicitly state this as a limitation.","section":"§III.A"},{"comment":"The text states that the dataset was constructed to 'minimize potential contamination' with uMLIP training sets, but no deduplication protocol is described. The 0D subset includes molecular structures from the Materials Project and generated clusters, which may overlap with training sets such as SPICE, ANI, or MPtrj. Please provide the exact similarity/removal criteria used and report how many structures were removed at each step. This is load-bearing for the benchmark's validity as an unbiased evaluation.","section":"§III.A and §IV"}],"minor_comments":[{"comment":"'Clockwise from the top' is ambiguous; please label the subpanels explicitly with 0D/1D/2D/3D or use a legend.","section":"Fig. 1"},{"comment":"The header 'N w Targets' is unclear. Define 'w' (weights?) and expand 'Targets' (E, F, S, D/G) directly in the caption.","section":"Table I"},{"comment":"For 0D systems, 'random space groups' is confusing because molecules do not have periodic space groups; clarify whether this refers to the simulation cell or to molecular point groups.","section":"§III.A"},{"comment":"The statement 'more than 75% of the energy predictions on 0123D dataset have an error lower than 10 meV/atom' should specify whether this is over all 40,000 systems or per dimensionality, and should note that it applies only to converged relaxations.","section":"§II.D"},{"comment":"The data availability section gives a general Alexandria URL rather than a direct link to the 0123D dataset; please provide a specific identifier/path.","section":"§IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid benchmark study with a useful dataset, but the two load-bearing issues—exclusion of failed relaxations from accuracy claims and the use of ORB-2 in dataset construction—need to be addressed before publication. The failure-exclusion problem is easy to fix in reporting, but the selection-bias concern may require additional experiments. I would not reject the manuscript; a major revision is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read this one. It's a benchmark, not a new method, but the dataset it ships—0123D, 40,000 PBE-relaxed structures across 0D/1D/2D/3D—is a genuinely useful resource. And the headline accuracy numbers are conditional: they're computed only for relaxations that converged without fragmenting, and the abstract overstates the result.\n\nWhat the paper does well: the benchmark design is mostly sound. Consistent PBE references, explicit overlap checks against training sets, and—credit where due—they report failure counts openly in Table II. The main finding (accuracy degrades as you move away from 3D training data, with eSEN the most transferable) is plausible, and the geometry comparison via Kabsch alignment is reasonable.\n\nThe soft spots are real but addressable. First, the error distributions in Figs. 5 and 6 and the 'best geometry models' ranking exclude unconverged or fragmented structures. eqV2 fails 620 of 10,000 0D systems, ORB-3d fails 89; if those count as infinite-error, eqV2's 0D geometry ranking changes. The paper reports these failures, so it's not hiding anything, but the abstract's 'errors below 10 meV/atom across all dimensionalities' is too strong—the body itself qualifies it as more than 75% for eSEN.\n\nSecond, the low-dimensional test set was pre-filtered with ORB-2 itself. Section III.A says structures were pre-relaxed with ORB-2 and hull distances were computed with ORB-2 energies. That's a mild but real selection bias in ORB-2's favor, and it's not flagged as a limitation. It's not fatal—ORB-2 also does well in 3D where selection used ALIGNN—but it should be disclosed.\n\nThe fixes are straightforward: report error statistics with and without failures, add numeric tables, and either select structures on an independent set or label the bias.\n\nWho benefits: anyone choosing a uMLIP for catalysts, interfaces, or 2D heterostructures gets a useful comparison; the dataset is a reusable benchmark for future models. It's not a deep theoretical advance, but it's a solid, useful study. It deserves a serious referee—the dataset alone justifies that. With tightening of the abstract and clearing up the failure accounting, this would be a good contribution.","headline":"Useful benchmark dataset and mostly honest comparison, but the headline accuracy claim is conditional on successful relaxations and the ORB-2 pre-filtering is a quiet bias.","tokens_in":13980,"tokens_out":3769,"would_cite":true,"duration_ms":36636,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"eSEN reaches DFT-level accuracy from isolated molecules to bulk crystals","keywords":["universal machine learning interatomic potentials","dimensionality transferability","0D-3D benchmark","0123D dataset","geometry relaxation","eSEN","equivariant neural network potentials","DFT-level accuracy"],"falsifier":"Rerun the benchmark with failures included as errors—assigning a failed relaxation an energy error at least as large as the full structure's energy difference, or a position error equal to the cell size—and check whether eSEN still stays below 10 meV/atom and within 0.02 Å across all dimensionalities. A second check: repeat the relaxations starting from slightly perturbed geometries and see if the reported sub-10 meV/atom errors persist.","tokens_in":12946,"feed_emoji":"⚛️","tokens_out":6364,"duration_ms":58966,"temperature":0.7,"pith_summary":"This paper benchmarks eleven universal machine-learning interatomic potentials (uMLIPs) on a new dataset of 40,000 relaxed compounds spanning four dimensionalities: isolated molecules and clusters (0D), nanowires and nanoribbons (1D), atomic layers and slabs (2D), and bulk crystals (3D). It finds that most models, which are trained predominantly on 3D bulk data, lose accuracy as dimensionality decreases, but a few maintain near-DFT performance everywhere. The best overall model, the equivariant Smooth Energy Network (eSEN), keeps energy errors below 10 meV/atom on more than 75% of all systems and position errors around 0.01–0.02 Å across all dimensionalities. The authors conclude that these top potentials are accurate enough to replace DFT for near-equilibrium geometry relaxation and energy evaluation across the full range from isolated atoms to solids, including mixed-dimensionality systems such as surfaces and interfaces.","feed_headline":"eSEN reaches DFT accuracy from isolated molecules to bulk crystals","feed_subtitle":"A benchmark of 11 universal potentials on 40,000 structures shows 0.01–0.02 Å position and under 10 meV/atom energy errors.","key_machinery":"The load-bearing components are the 0123D dataset and the relaxation workflow that turns the benchmark into an end-to-end test. The dataset holds 10,000 PBE-relaxed compounds per dimensionality, generated with dimension-specific strategies—a generative model for 3D candidates and PyXtal for lower-dimensional structures—then pre-relaxed with ORB-2 and refined with DFT. The evaluation protocol starts each uMLIP from the DFT equilibrium geometry, relaxes with the FIRE optimizer, and measures the mean absolute error in atomic positions (after Kabsch alignment) and the energy difference relative to the PBE reference, with failures and fragmentation separately recorded. This design tests whether u","core_discovery":"The central claim is that state-of-the-art universal machine-learning interatomic potentials have reached DFT-level accuracy in geometry relaxation and energy prediction across the entire range of system dimensionality, from isolated molecules to bulk crystals. The paper supports this with a purpose-built benchmark: 10,000 relaxed PBE-quality structures for each dimensionality (the 0123D dataset), constructed to avoid overlap with existing training sets. Forcing each uMLIP to relax these structures from the DFT equilibrium geometry, the authors measure failure rates, optimization steps, energy errors, and position errors. Across all four dimensionalities, the best models—ORB-v2, eqV2, and es","pith_inferences":["A testable extension would be to fine-tune existing uMLIPs on a fraction of the 0123D dataset and measure how much the dimensionality gap closes; the paper's 'do not train on this' request makes this a controlled experiment for future work.","The concentration of failures in force-direct (non-conservative) models suggests that an energy-derived force architecture, or a hybrid that projects forces to a conservative form, may be the most robust path for low-dimensional applications—a hypothesis the paper raises but does not itself pursue.","Because the benchmark only samples near-equilibrium geometries, the DFT-replacement claim likely holds for relaxation and static energetics, but extrapolating to reactive pathways, finite-temperature dynamics, or charged defects would require separate validation.","For practical users, starting relaxation from the DFT geometry (as done here) may inflate apparent performance relative to starting from a rough or random structure; applying both protocols would quantify this gap."],"forward_implications":["If the benchmark numbers hold, eSEN and the top rivals can replace DFT for routine geometry relaxation and single-point energies across molecules, wires, layers, and crystals, at a fraction of the cost.","Simulations that couple subsystems of different dimensionality—a molecule on a slab, a nanowire on a support—are no longer forced to mix incompatible levels of theory.","The measured degradation from 3D to 0D exposes a training-data bias; models trained with more dimension-balanced data should close the gap, a direct incentive for new dataset construction.","The high failure rates of the non-conservative models (eqV2, ORB-3d) in low-dimensional relaxation imply that direct force prediction, unless carefully controlled, undermines practical reliability in the very regimes where DFT replacement is most wanted."],"supporting_citations":[{"why":"Defines the eSEN model, the paper's top performer on energy accuracy and the main subject of the headline claim.","marker":"[25]"},{"why":"Supplies the eqV2 model and the OMat24 training dataset that many competing models rely on, making it the primary comparative baseline.","marker":"[8]"},{"why":"Describes ORB-2, used both as a leading geometry-optimization model and as the pre-relaxation tool for building the 0123D dataset.","marker":"[9]"},{"why":"Introduces the ORB-3 models, whose high failure rates in low-dimensional relaxations anchor the paper's assessment of non-conservative force prediction.","marker":"[28]"},{"why":"Defines the Alexandria dataset and its computational parameters, which the 0123D dataset mirrors for consistency and which many uMLIPs use for training.","marker":"[16]"},{"why":"Provides the generative model that produced the 3D candidate structures entering the benchmark dataset.","marker":"[39]"},{"why":"Supplies the PyXtal structure generator used to create the 0D, 1D, and 2D candidates, underpinning the dataset's dimensional diversity.","marker":"[40]"},{"why":"Gives the Kabsch alignment algorithm that computes the position-error metric underlying the 0.01–0.02 Å claim.","marker":"[43]"},{"why":"Specifies the FIRE optimizer used in every relaxation run, forming the protocol that all benchmark results depend on.","marker":"[42]"}],"fun_headline_variants":["esEN, ORB-v2, eqV2 match DFT from molecules to solids","Universal ML potentials now DFT-accurate for all dimensions","From 0D to 3D: ML potentials hit DFT accuracy","40,000 structures show: ML potentials rival DFT everywhere","Atom to bulk: ML potentials achieve DFT-level accuracy"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The headline error ranges (0.01–0.02 Å, below 10 meV/atom) are computed only over relaxations that converged without fragmentation; for models like eqV2, 620 of 10,000 0D systems failed and are excluded, so counting failures as errors would shift the rankings and the headline numbers.","fun_headline_variants_meta":{"raw":{"variants":["esEN, ORB-v2, eqV2 match DFT from molecules to solids","Universal ML potentials now DFT-accurate for all dimensions","From 0D to 3D: ML potentials hit DFT accuracy","40,000 structures show: ML potentials rival DFT everywhere","Atom to bulk: ML potentials achieve DFT-level accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1536,"prompt_tokens":770,"completion_tokens":766,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":677}},"tokens_in":514,"tokens_out":766,"duration_ms":7967,"temperature":1.0,"reasoning_tokens":677,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:47:56.878749+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the benchmark with failures included as errors—assigning a failed relaxation an energy error at least as large as the full structure's energy difference, or a position error equal to the cell size—and check whether eSEN still stays below 10 meV/atom and within 0.02 Å across all dimensionalities. A second check: repeat the relaxations starting from slightly perturbed geometries and see if the reported sub-10 meV/atom errors persist.","supporting_citations":[{"cited_title":"Kabsch, A solution for the best rotation to relate two sets of vectors, Foundations of Crystallography32, 922 (1976)","cited_arxiv_id":null,"evidence_quote":"Gives the Kabsch alignment algorithm that computes the position-error metric underlying the 0.01–0.02 Å claim."},{"cited_title":"Bitzek, P","cited_arxiv_id":null,"evidence_quote":"Specifies the FIRE optimizer used in every relaxation run, forming the protocol that all benchmark results depend on."}],"review_version":1}