Pith. sign in

REVIEW 4 major objections 5 minor 100 references

This paper claims that nine pretrained universal machine learning potentials span from near-ab initio to essentially non-predictive in crystal structure global optimization, with eSEN matching DFT to 5.28 meV/atom ranking error.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 20:18 UTC pith:A3DO3LYK

load-bearing objection A large, credible benchmark of nine uMLPs on global structure searches with a real performance spread, but the pool-merge protocol inflates the quantitative rankings and the abstract lists models that are not in the results. the 4 major comments →

arxiv 2602.23515 v2 pith:A3DO3LYK submitted 2026-02-26 cond-mat.mtrl-sci

Performance of universal machine learning potentials in global optimization of inorganic crystal structures

classification cond-mat.mtrl-sci
keywords universal machine learning potentialscrystal structure predictionglobal optimizationevolutionary algorithmsground statesdensity functional theorybenchmarkinginorganic compounds
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper benchmarks nine pretrained universal machine learning potentials by running evolutionary searches for ground-state crystal structures across twelve multicomponent inorganic compounds, then checking low-energy candidates against density functional theory. It aims to show that these models span a wide performance range, from near-ab initio to essentially non-predictive, and that at least one model, eSEN, reproduces the reference DFT method more accurately than the systematic differences among DFT approximations themselves. The searches also surfaced two new candidate ground states, one of which (oI28-MgB3C3) is favored by every tested functional. A sympathetic reader should care because, if the claim holds, off-the-shelf potentials can accelerate discovery of new stable inorganic phases without bespoke fitting.

Core claim

The central claim is that the latest generation of universal machine learning potentials can serve as reliable surrogate energy surfaces in global structure optimization, but with a wide performance spread: eSEN's ranking RMSE is 5.28 meV/atom, while M3GNet's is 23.97 meV/atom with an energy proximity above 70 meV/atom. eSEN's accuracy is better than the systematic spread among PBE, PBEsol, and r2SCAN themselves. The evolutionary searches found an oI28-MgB3C3 phase uniformly favored by all tested functionals, and a tI10-Na2CN2 phase favored only by PBE, likely a functional-specific artifact. The paper argues that the benchmark demonstrates current pretrained uMLPs can already boost ab initio

What carries the argument

The central mechanism is a benchmarking protocol built around merged pools of low-energy minima. Each uMLP runs evolutionary searches and contributes candidate structures; all pools are merged so every test set contains the DFT ground state, then re-relaxed with each uMLP and re-ranked. The ranking RMSE metric, defined as the average squared deviation between uMLP and DFT energies after removing the average pool shift, together with structure and energy proximity metrics, isolates a model's ability to order competing phases within low-energy basins.

Load-bearing premise

The claim that uMLPs can boost unconstrained global optimization rests on the merging step: every model's test pool is seeded with the ground state found by other models, and a model counts as successful even when it only fails to destroy an injected foreign minimum rather than discovering it independently.

What would settle it

Run each uMLP in a fully unconstrained evolutionary search with no merged-pool injection and no ground-state seeding, using many independent random seeds, and count how often each model's own shortlist contains the DFT ground state; if eSEN's independent discovery rate is far below its near-perfect merged-pool success rate, the conclusion that uMLPs are ready for unconstrained searches would be refuted.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the claim holds, off-the-shelf universal potentials, especially eSEN, can replace system-specific potentials as the default starting point for crystal structure prediction.
  • The oI28-MgB3C3 phase is a concrete, testable prediction: deintercalating the MgB2C2 precursor may yield a 3D-connected BC framework rather than the previously proposed honeycomb layers.
  • Relatively larger models with more expressive descriptors and broader training data transfer more reliably across chemistries and competing structures.
  • Several uMLPs resolve near-equilibrium distortions driven by electronic structure, enabling large-scale exploration of superconducting or superhard boride phase space without exhaustive DFT analysis of every distortion pattern.
  • The tI10-Na2CN2 finding illustrates that uMLP-driven searches can surface DFT-functional-specific artifacts, so multi-functional validation remains necessary.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The reported per-model success rates likely overstate unconstrained discovery capability: each model's pool was merged with minima found by other models, and success was credited if the ground state merely appeared in the shortlist, not if the model discovered it independently, so single-model deployments could show a wider performance gap.
  • The ranking RMSE metric, which subtracts the average energy shift and focuses only on near-ground-state basins, may be a more transferable measure of a potential's usefulness for structure prediction than absolute formation-energy errors reported in standard benchmarks.
  • The weakness of most uMLPs on zinc's c/a anomaly and on layer stacking in vdW materials suggests that some ground-state questions will still require hybrid workflows that call DFT for final ranking or retraining on target motifs.
  • The two newly proposed phases are falsifiable by synthesis attempts or by higher-level electronic structure methods beyond the tested DFT functionals, providing a direct route to test whether the uMLP-accelerated search found real physics or only a DFT-level artifact.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper benchmarks nine universal machine-learning potentials (M3GNet, MACE, SevenNet, EquiformerV2, MatterSim, GRACE, eSEN, Orb-v3, PET-MAD) as surrogate models in evolutionary crystal-structure searches for 12 inorganic compounds. It combines qualitative success rates (Table II), ranking-RMSE and proximity metrics on merged pools (Fig. 2, Table III), and targeted near-equilibrium tests (Zn c/a, Cr/Mn/FeB4, LiBy). The central claim is that uMLPs span a wide performance range from near-DFT to essentially non-predictive, with eSEN the most consistent, and that current uMLPs can already accelerate ab initio global optimization. The paper also reports two new candidate phases, oI28-MgB3C3 and tI10-Na2CN2.

Significance. If the claims hold, this is a valuable large-scale benchmark: over two million uMLP relaxations, more than seven thousand DFT re-optimizations, explicit and well-motivated metrics, multi-functional DFT controls, and testable structural predictions. The perturbation tests (Zn, MB4, LiBy) add useful information beyond simple formation-energy benchmarks, and the comparison with system-specific NN potentials gives practical perspective. However, the load-bearing quantitative ordering is entangled with the merged-pool protocol and with in-house/unverified target structures, so the significance is currently conditional on a focused re-analysis.

major comments (4)
  1. [Section II C, Fig. 2, Table III] The pooled-merge step merges all uMLP minima into a single pool 'to ensure that each uMLP test pool included the ground state' before per-model re-relaxation and DFT ranking. This means Fig. 2's ranking RMSE and proximity scores do not measure a model's own unconstrained discovery ability: a model that never reaches the ground-state basin can still receive a low RMSE if it does not destroy an injected foreign minimum. The summary statement that uMLPs 'can already boost ab initio global optimization' and the EN-vs-MG ordering (5.28 vs 23.97 meV/atom) therefore conflate ranking fidelity on a common candidate set with search performance. Please recompute the pooled metrics on each model's original pre-merge minima, or otherwise show sensitivity to pool composition (e.g., restrict to structures originating from that model's own search).
  2. [Section III, Table II, Section V] Three of the twelve target ground states are the authors' own prior predictions (Li3Sn, Pd5Sn3, MgB3C3), and oI28-MgB3C3 was actually discovered in this same benchmark before being used as the reference ground state. Scoring models against a structure generated by the model ensemble is circular for the discovery claim; it also makes the success rates vulnerable to errors in these unpublished target structures. Please report results with and without these in-house targets, or at least separate 'known targets' from 'self-discovered targets' and discuss the effect on Table II and Table III scores.
  3. [Section III, Table III] The success rates are based on 12 compounds, with at most three random seeds per failed run. Differences such as 11/12 vs 10/12 vs 9/12 are within sampling noise, despite the paper's own admission that the dataset is not statistically sufficient (Section III) and the caveat in Section V. The table and Summary nevertheless use these rates to rank models. Please provide confidence intervals, bootstrap estimates, or a statistical test, and soften model-level claims derived from success counts.
  4. [Section II C, Fig. 2 (MG row)] The text states that 'the representative EN set' was used to evaluate MG because MG sets were excluded from the merge. If the MG row in Fig. 2 is computed on EN's pooled structures and not on M3GNet's own relaxations, then the 'essentially non-predictive' label for MG is not a fair test of that model. Please clarify what was actually done and recompute MG metrics on MG's own pool (or explain why the EN set is a valid proxy).
minor comments (5)
  1. [Abstract vs full text] The abstract lists twelve models (including MACE-MATPES, MACE-mh-1, EquiformerV3, PET-OAM), while the full text, Table I, and all analyses treat nine models. Please correct this inconsistency.
  2. [Table II] The AgClO4 entry for MC is '-1e40', which appears to be a placeholder or numerical artifact; please replace with a meaningful value or state explicitly that the model produced no valid minimum. Also clarify in the caption that negative entries correspond to uMLP minima lying below the reference DFT ground state.
  3. [Table II caption] The superscripts indicating the rank of the reference structure are not readable in the text-only version and their convention (including a 'zero superscript') is not defined in the caption. Please define all symbols.
  4. [Table III] The composite score on a 12-point scale is not defined with enough detail to be reproducible. Please provide the exact rubric or a worked example for at least one model.
  5. [Fig. 4] The caption identifies 'three DFT approximations' but does not state which curve corresponds to PBE, PBEsol, or r2SCAN. Please label the curves directly or add a legend.

Circularity Check

0 steps flagged

No load-bearing circularity found; benchmark targets external DFT energies and no uMLP parameter is fitted in this work.

full rationale

The paper's central claims—that the nine uMLPs span a wide performance range and that eSEN approaches the reference DFT accuracy—are supported by evolutionary searches and by ranking RMSEs computed on pools of candidate structures whose energies are re-evaluated with DFT (PBE or PBEsol, according to each model's training functional). No uMLP parameter is fitted in the paper, and the reference energies are external DFT results for the same structures, so the quantitative comparisons are not derived from the uMLPs' own outputs. Self-citations are present (own MAISE search engine, own prior NN potentials, own previously predicted Li3Sn, Pd5Sn3, and MgB3C3 ground states used as benchmark targets), but these are contextual: the benchmark targets are re-optimized and re-scored with DFT in the present paper, and the search engine is a tool rather than an input that determines the outcome. The main caveat is methodological rather than circular: Section II C merges minima found by all uMLPs 'to ensure that each uMLP test pool included the ground state,' and Section III defines success as the DFT ground state appearing 'within the low-energy shortlist, even if it was not ranked first.' This means the success metric tests whether a model preserves an injected foreign minimum, not whether it can discover that minimum on its own, and the Fig. 2 ranking RMSEs inherit the merged-pool composition. That is a benchmark-design limitation that could affect the practical conclusions about unconstrained deployment, but it does not reduce any predicted quantity to an input by construction: the RMSEs still compare genuine model energies with genuine DFT energies of the pooled structures. No uniqueness theorem, no ansatz-by-citation, and no renaming of known results occurs. Overall, the derivation chain is largely self-contained and any circularity is minor and non-load-bearing.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central claim rests on four hand-chosen protocol parameters (pool energy windows, fingerprint cutoffs, convergence tolerances, and the per-pool mean-shift subtraction in the ranking metric), on the validity of PBE as the arbiter across all models, and on the correctness of the reference ground states — several of which are the authors' own prior predictions. No new physical entities are postulated; the two new crystal phases are outputs of the method, not inputs, and each carries its own evidentiary status.

free parameters (4)
  • per-pool mean energy shift ΔE = per-pool average E_uMLP − E_DFT (values not tabulated; aggregated in Fig. 2)
    Subtracted in the ranking RMSE definition (Section II C) to remove average shifts; a calibration offset fitted separately for each compound pool, and therefore a free parameter in the reported accuracy metric.
  • energy window for pool filtering = 20 meV/atom initially, stepped by 10 up to 100 meV/atom
    Hand-chosen post-hoc window to ensure at least 10 minima per pool (Section II C); changes which minima enter the pool and hence the ranking and proximity metrics.
  • SCUT structure-fingerprint thresholds = 0.92 default; 0.82 for low-symmetry satellite minima
    Hand-chosen similarity cutoffs controlling duplicate elimination and pool size (Section II C); the 0.82 value is introduced only for cases where models improperly model structures as dynamically unstable.
  • force-convergence tolerances = 0.05 eV/Å (coarse), 0.001 eV/Å (fine)
    Chosen convergence criteria defining which minima are visited and retained during the evolutionary search (Section II C); affects the set of candidate structures in each pool.
axioms (4)
  • domain assumption DFT-PBE (with PAW, 500 eV cutoff, Δk ≤ 0.025 Å⁻¹) provides the reference ground states used to score all models; PBEsol and r2SCAN are used only as cross-checks.
    The benchmark's ground truth is a single semi-local DFT flavor. PBE's own errors propagate into the model rankings: e.g., PBE places the Zn c/a minimum at 1.894 vs the experimental 1.826 (Section IV A), and the tI10-Na2CN2 phase is favored only by PBE (Section III). Section II A states the DFT setup.
  • domain assumption The 12 selected compounds' reported ground states, including the authors' own recently proposed Li3Sn, Pd5Sn3, and MgB3C3 prototypes and the newly found oI28-MgB3C3 phase, are the true ambient-pressure ground states.
    If any reference structure were wrong, the corresponding benchmark column would be mislabeled. For MgB3C3 the benchmark's own product (oI28) is the claimed ground state, with no experimental verification. Sections III and V.
  • domain assumption MAISE evolutionary search (100 structures, 100 generations, 20/60/20 mutation/crossover/injection) with force tolerance 0.05 eV/Å adequately samples low-energy basins for every model.
    Search failure vs model failure cannot be separated; the authors state the dataset "is not statistically sufficient to disentangle whether missed ground states reflect surrogate deficiencies or seed/trajectory dependence" (Section III).
  • domain assumption Pretrained uMLP checkpoints are used as-is, and their energies are comparable without absolute-reference calibration.
    Stated in Section II B; the ranking metric removes absolute shifts precisely because raw energies are not directly comparable across models and functionals.
invented entities (2)
  • oI28-MgB3C3 polymorph independent evidence
    purpose: Proposed new ambient-pressure ground state of MgB3C3 with a 3D sp3-connected BC framework, reported below the hP14/hP7 layered superconducting phases by all tested DFT functionals.
    Falsifiable prediction: synthesis via MgB2C2 deintercalation or higher-level electronic-structure calculations can confirm or refute it. Currently DFT-only (Section III, Fig. 1).
  • tI10-Na2CN2 phase no independent evidence
    purpose: Proposed less-dense Na2CN2 packing, 1.9 meV/atom below the known mS10 phase at the PBE level.
    The paper itself flags it as a "likely PBE-specific artifact" (Section III): all other tested functionals disfavor it by 16-56 meV/atom. No experimental handle is offered.

pith-pipeline@v1.3.0-alltime-deepseek · 20128 in / 22756 out tokens · 192861 ms · 2026-08-02T20:18:12.014604+00:00 · methodology

0 comments
read the original abstract

Rapid development of universal machine learning potentials (uMLPs) and expansion of training data sets are reshaping the state of the art in atomistic simulation, highlighting the need for concurrent systematic benchmarking of their capabilities. Global optimization is among the most demanding uMLP applications because unconstrained exploration includes probing motifs not present in reference sets. We examined twelve pretrained uMLPs in unconstrained evolutionary searches to assess whether these models can consistently predict complex nonmagnetic crystal structure ground states under ambient pressure across diverse inorganic systems. Our findings demonstrate that the considered M3GNet, MACE-MATPES, MACE-mh-1, SevenNet, EquiformerV2, EquiformerV3, MatterSim, GRACE, eSEN, Orb-v3, PET-MAD, and PET-OAM models span a wide performance range, from near {\it ab initio} to essentially non-predictive, in their ability to resolve competing phases within low-energy basins. Additional tests on hcp-Zn, MB$_4$ (M = Cr, Mn, and Fe), and LiB$_{y}$ ($y\approx 0.9$) ground states reveal that several uMLPs capture fine energy differences arising from subtle electronic structure features.

Figures

Figures reproduced from arXiv: 2602.23515 by Aleksey N. Kolmogorov, Edan T. Marcial, Laxman Chaudhary, Olesya Gorbunova.

Figure 1
Figure 1. Figure 1: FIG. 1. Stability of the tI10-Na [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Band structure and density of states calculated with [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Stability of distorted oP10 and mP20 derivatives [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Stability ranges for Li [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

100 extracted references · 1 canonical work pages

  1. [1]

    Behler and M

    J. Behler and M. Parrinello, Generalized neural-network representation of high-dimensional potential-energy sur- faces, Phys. Rev. Lett. 98, 146401 (2007)

  2. [2]

    A. P. Bart´ ok, M. C. Payne, R. Kondor, and G. Cs´ anyi, Gaussian approximation potentials: The accuracy of quantum mechanics, without the electrons, Phys. Rev. Lett. 104, 136403 (2010)

  3. [3]

    V. L. Deringer, M. A. Caro, and G. Cs´ anyi, Machine learning interatomic potentials as emerging tools for materials science, Adv. Mater. 31, 1902765 (2019)

  4. [4]

    G. L. W. Hart, T. Mueller, C. Toher, and S. Curtarolo, Machine learning for alloys, Nat. Rev. Mater. 6, 730 (2021). 11

  5. [5]

    Mishin, Machine-learning interatomic potentials for materials science, Acta Materialia 214, 116980 (2021)

    Y. Mishin, Machine-learning interatomic potentials for materials science, Acta Materialia 214, 116980 (2021)

  6. [6]

    Huang, C

    S.-D. Huang, C. Shang, P.-L. Kang, and Z.-P. Liu, Atomic structure of boron resolved using machine learn- ing and global sampling, Chem. Sci. 9, 8644 (2018)

  7. [7]

    E. V. Podryabinkin, E. V. Tikhonov, A. V. Shapeev, and A. R. Oganov, Accelerating crystal structure pre- diction by machine-learning interatomic potentials with active learning, Phys. Rev. B 99, 064114 (2019)

  8. [8]

    Behler, R

    J. Behler, R. Martoˇ n´ ak, D. Donadio, and M. Parrinello, Metadynamics simulations of the high-pressure phases of silicon employing a high-dimensional neural network potential, Phys. Rev. Lett. 100, 185501 (2008)

  9. [9]

    V. L. Deringer, D. M. Proserpio, G. Cs´ anyi, and C. J. Pickard, Data-driven learning and prediction of in- organic crystal structures, Faraday Discuss. 211, 45 (2018)

  10. [10]

    Ibarra-Hern´ andez, S

    W. Ibarra-Hern´ andez, S. Hajinazar, G. Avenda˜ no- Franco, A. Bautista-Hern´ andez, A. N. Kolmogorov, and A. H. Romero, Structural search for stable Mg–Ca alloys accelerated with a neural network interatomic model, Phys. Chem. Chem. Phys. 20, 27545 (2018)

  11. [11]

    Kharabadze, A

    S. Kharabadze, A. Thorn, E. A. Koulakova, and A. N. Kolmogorov, Prediction of stable Li-Sn compounds: boosting ab initio searches with neural network poten- tials, npj Comput. Mater. 8, 136 (2022)

  12. [12]

    Thorn, D

    A. Thorn, D. Gochitashvili, S. Kharabadze, and A. N. Kolmogorov, Machine learning search for stable binary Sn alloys with Na, Ca, Cu, Pd, and Ag, Phys. Chem. Chem. Phys. 25, 22415 (2023)

  13. [13]

    Ouyang, Y

    R. Ouyang, Y. Xie, and D.-e. Jiang, Global minimiza- tion of gold clusters by combining neural network po- tentials and the basin-hopping method, Nanoscale 7, 14817–14821 (2015)

  14. [14]

    Gubaev, E

    K. Gubaev, E. V. Podryabinkin, G. L. W. Hart, and A. V. Shapeev, Accelerating high-throughput searches for new alloys with active learning of interatomic poten- tials, Comput. Mater. Sci. 156, 148 (2019)

  15. [15]

    Hajinazar, J

    S. Hajinazar, J. Shao, and A. N. Kolmogorov, Strati- fied construction of neural network based interatomic models for multicomponent materials, Phys. Rev. B 95, 014114 (2017)

  16. [16]

    Hajinazar, A

    S. Hajinazar, A. Thorn, E. D. Sandoval, S. Kharabadze, and A. N. Kolmogorov, MAISE: Construction of neural network interatomic models and evolutionary structure optimization, Comput. Phys. Commun. 259, 107679 (2021)

  17. [17]

    P. E. Dolgirev, I. A. Kruglov, and A. R. Oganov, Machine learning scheme for fast extraction of chemi- cally interpretable interatomic potentials, AIP Adv. 6, 085318 (2016)

  18. [18]

    Artrith, T

    N. Artrith, T. Morawietz, and J. Behler, High- dimensional neural-network potentials for multicompo- nent systems: Applications to zinc oxide, Phys. Rev. B 83, 153101 (2011)

  19. [19]

    C. J. Pickard, Ephemeral data derived potentials for random structure search, Phys. Rev. B 106, 014102 (2022)

  20. [20]

    Chen and S

    C. Chen and S. P. Ong, A universal graph deep learning interatomic potential for the periodic table, Nat. Com- put. Sci. 2, 718 (2022)

  21. [21]

    P´ ota, P

    B. P´ ota, P. Ahlawat, G. Cs´ anyi, and M. Simon- celli, Thermal conductivity predictions with foun- dation atomistic models, arXiv:2408.00755 (2024), arXiv:2408.00755 [cond-mat.mtrl-sci]

  22. [22]

    Sharma, A

    K. Sharma, A. Loew, H. Wang, F. A. Nilsson, M. Jain, M. A. L. Marques, and K. S. Thygesen, Accelerating point defect photo-emission calculations with machine learning interatomic potentials, npj Comput. Mater.11, 18201 (2025)

  23. [23]

    H. Yu, M. Giantomassi, G. Materzanini, J. Wang, and G.-M. Rignanese, Systematic assessment of various uni- versal machine-learning interatomic potentials, Mater. Genome Eng. Adv. 2, e58 (2024)

  24. [24]

    A. Loew, D. Sun, H.-C. Wang, S. Botti, and M. A. L. Marques, Universal machine learning interatomic poten- tials are ready for phonons, npj Comput. Mater. 11, 10 (2025)

  25. [25]

    Riebesell, R

    J. Riebesell, R. E. A. Goodall, P. Benner, Y. Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, A framework to evaluate machine learn- ing crystal stability predictions, Nat. Mach. Intell. 7, 836 (2025)

  26. [26]

    Chiang, T

    Y. Chiang, T. Kreiman, C. Zhang, M. C. Kuner, E. Weaver, I. Amin, H. Park, Y. Lim, J. Kim, D. Chrzan, A. Walsh, S. M. Blau, M. Asta, and A. S. Krishnapriyan, MLIP Arena: Advanc- ing Fairness and Transparency in Machine Learn- ing Interatomic Potentials via an Open, Accessible Benchmark Platform, arXiv preprint arXiv:2509.20630 10.48550/arXiv.2509.20630 (2025)

  27. [27]

    Tahmasbi, A

    H. Tahmasbi, A. Kn¨ upfer, T. D. K¨ uhne, and H. Mirhos- seini, Benchmarking Universal Machine Learning Inter- atomic Potentials on Elemental Systems, arXiv preprint arXiv:2512.20230 10.48550/arXiv.2512.20230 (2025)

  28. [28]

    A. Peng, C. Cai, M. Guo, D. Zhang, C. Zhang, W. Jiang, Y. Wang, A. Loew, C. Wu, W. E, L. Zhang, and H. Wang, Lambench: a benchmark for large atomistic models, npj Computat. Mater. 12, 62 (2026)

  29. [29]

    A. Loew, J. Schmidt, S. Botti, and M. A. L. Marques, Universal machine learning potentials under pressure, Journal of Physics: Materials 9, 015010 (2025)

  30. [30]

    Lyngby, C

    P. Lyngby, C. Larsen, and K. W. Jacobsen, Bayesian optimization of atomic structures with prior probabil- ities from universal interatomic potentials, Phys. Rev. M 8 (2024)

  31. [31]

    Mair, H.-G

    G. Mair, H.-G. von Schnering, M. W¨ orle, and R. Nesper, Dilithium Hexaboride, Li 2B6, Z. Anorg. Allg. Chem. 625, 1207 (1999)

  32. [32]

    Wang, H.-Y

    D.-H. Wang, H.-Y. Zhou, C.-H. Hu, Y. Zhong, A. R. Oganov, and G.-H. Rao, Prediction of thermodynami- cally stable Li–B compounds at ambient pressure, Phys. Chem. Chem. Phys. 19, 8471 (2017)

  33. [33]

    S. K. Choi, R. Coldea, A. N. Kolmogorov, T. Lancaster, I. I. Mazin, S. J. Blundell, P. G. Radaelli, Y. Singh, P. Gegenwart, K. R. Choi, S.-W. Cheong, P. J. Baker, C. Stock, and J. Taylor, Spin Waves and Revised Crystal Structure of Honeycomb Iridate Na 2IrO3, Phys. Rev. Lett. 108, 127204 (2012)

  34. [34]

    H. J. Becher and A. Sch¨ afer, Darstellung und Struktur des Berylliumborids Be 4B, Z. Anorg. Allg. Chem. 318, 304 (1962)

  35. [35]

    Hermann, N

    A. Hermann, N. W. Ashcroft, and R. Hoffmann, Binary compounds of boron and beryllium: A rich structural arena with space for predictions, Chemistry – A Euro- pean Journal 19, 4184 (2013)

  36. [36]

    C. J. Howard, T. M. Sabine, and F. Dickson, Structural and thermal parameters for rutile and anatase, Acta 12 Crystallogr. Sect. B: Struct. Sci. 47, 462 (1991)

  37. [37]

    Eguchi, D

    G. Eguchi, D. C. Peets, M. Kriener, Y. Maeno, E. Nishi- bori, Y. Kumazawa, K. Banno, S. Maki, and H. Sawa, Crystallographic and superconducting properties of the fully gapped noncentrosymmetric 5d-electron supercon- ductors CaM Si3 (M = Ir, Pt), Phys. Rev. B 83, 024512 (2011)

  38. [38]

    Ahn, H.-J

    K.-H. Ahn, H.-J. Lee, D. Johrendt, and K. A. M¨ uller, In- fluence of pressure on the structure and electronic prop- erties of the layered superconductor Y 2C2I2, J. Phys. Condens. Matter 17, L637 (2005)

  39. [39]

    Errandonea, L

    D. Errandonea, L. Gracia, A. Beltr´ an, A. Vegas, and Y. Meng, Pressure-induced phase transitions in AgClO4, Phys. Rev. B 84, 064103 (2011)

  40. [40]

    Zhang, J

    Y. Zhang, J. W. Furness, B. Xiao, and J. Sun, Sub- tlety of TiO 2 phase stability: Reliability of the density functional theory predictions and persistence of the self- interaction error, J. Chem. Phys. 150, 014105 (2019)

  41. [41]

    C. R. Tomassetti, D. Gochitashvili, C. Renskers, E. R. Margine, and A. N. Kolmogorov, First-principles design of ambient-pressure Mg xB2C2 and Na xBC supercon- ductors, Phys. Rev. Mater. 8, 114801 (2024)

  42. [42]

    Batatia, D

    I. Batatia, D. P. Kov´ acs, G. N. C. Simm, C. Ortner, and G. Cs´ anyi, MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields, in Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 35 (2022) pp. 11423–11436

  43. [43]

    Batatia, S

    I. Batatia, S. Batzner, D. P. Kov´ acs, A. Musaelian, G. N. C. Simm, R. Drautz, C. Ortner, B. Kozinsky, and G. Cs´ anyi, The design space of E(3)-equivariant atom- centred interatomic potentials, Nature Machine Intelli- gence 7, 56–67 (2025)

  44. [44]

    Y. Park, J. Kim, S. Hwang, and S. Han, Scalable paral- lel algorithm for graph neural network interatomic po- tentials in molecular dynamics simulations, J. Chem. Theory Comput. 20, 4857 (2024)

  45. [45]

    J. Kim, J. Kim, J. Kim, J. Lee, Y. Park, Y. Kang, and S. Han, Data-Efficient Multifidelity Training for High-Fidelity Machine Learning Interatomic Potentials, J. Am. Chem. Soc. 147, 1042 (2024)

  46. [46]

    Y.-L. Liao, B. Wood, A. Das, and T. Smidt, EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations, in Interna- tional Conference on Learning Representations (ICLR) (2024)

  47. [47]

    Barroso-Luque, M

    L. Barroso-Luque, M. Shuaibi, X. Fu, B. M. Wood, M. Dzamba, M. Gao, A. Rizvi, C. L. Zitnick, and Z. W. Ulissi, Open Materials 2024 (OMat24) Inor- ganic Materials Dataset and Models, arXiv preprint arXiv:2410.12771 10.48550/arXiv.2410.12771 (2024)

  48. [48]

    Liao and T

    Y.-L. Liao and T. Smidt, Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic Graphs, in International Conference on Learning Representa- tions (ICLR) (2023)

  49. [49]

    Chanussot, A

    L. Chanussot, A. Das, S. Goyal, T. Lavril, M. Shuaibi, M. Riviere, K. Tran, J. Heras-Domingo, C. Ho, W. Hu, A. Palizhati, A. Sriram, B. Wood, J. Yoon, D. Parikh, C. L. Zitnick, and Z. Ulissi, Open Catalyst 2020 (OC20) Dataset and Community Challenges, ACS Catal. 11, 6059 (2021)

  50. [50]

    H. Yang, C. Hu, Y. Zhou, X. Liu, Y. Shi, J. Li, G. Li, Z. Chen, S. Chen, C. Zeni, M. Horton, R. Pinsler, A. Fowler, D. Z¨ ugner, T. Xie, J. Smith, L. Sun, Q. Wang, L. Kong, C. Liu, H. Hao, and Z. Lu, Mat- terSim: A Deep Learning Atomistic Model Across El- ements, Temperatures and Pressures, arXiv preprint arXiv:2405.04967 10.48550/arXiv.2405.04967 (2024)

  51. [51]

    Bochkarev, Y

    A. Bochkarev, Y. Lysogorskiy, and R. Drautz, Graph atomic cluster expansion for semilocal interactions be- yond equivariant message passing, Phys. Rev. X 14, 021036 (2024)

  52. [52]

    Lysogorskiy, A

    Y. Lysogorskiy, A. Bochkarev, and R. Drautz, Graph atomic cluster expansion for foundational ma- chine learning interatomic potentials, arXiv preprint arXiv:2508.17936 10.48550/arXiv.2508.17936 (2025)

  53. [53]

    X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zitnick, Learning Smooth and Expressive Interatomic Poten- tials for Physical Property Prediction, arXiv preprint arXiv:2502.12147 10.48550/arXiv.2502.12147 (2025)

  54. [54]

    Rhodes, S

    B. Rhodes, S. Vandenhaute, V. ˇSimkus, J. Gin, J. God- win, T. Duignan, and M. Neumann, Orb-v3: atomistic simulation at scale, arXiv preprint arXiv:2504.06231 10.48550/arXiv.2504.06231 (2025)

  55. [55]

    Neumann, J

    M. Neumann, J. Gin, B. Rhodes, S. Bennett, Z. Li, H. Choubisa, A. Hussey, and J. Godwin, Orb: A Fast, Scalable Neural Network Potential, arXiv preprint arXiv:2410.22570 10.48550/arXiv.2410.22570 (2024)

  56. [56]

    Mazitov, F

    A. Mazitov, F. Bigi, M. Kellner, P. Pegolo, D. Tisi, G. Fraux, S. Pozdnyakov, P. Loche, and M. Ceriotti, PET-MAD as a lightweight universal interatomic po- tential for advanced materials modeling, Nat. Commun. 16, 65662 (2025)

  57. [57]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Generalized Gradient Approximation Made Simple, Phys. Rev. Lett. 77, 3865 (1996)

  58. [58]

    J. P. Perdew, A. Ruzsinszky, G. I. Csonka, O. A. Vy- drov, G. E. Scuseria, L. A. Constantin, X. Zhou, and K. Burke, Restoring the density-gradient expansion for exchange in solids and surfaces, Phys. Rev. Lett. 100, 136406 (2008)

  59. [59]

    J. W. Furness, A. D. Kaplan, J. Ning, J. P. Perdew, and J. Sun, Accurate and numerically efficient r 2SCAN meta-generalized gradient approximation, J. Phys. Chem. Lett. 11, 8208 (2020)

  60. [60]

    Kresse and J

    G. Kresse and J. Hafner, Ab initio molecular dynamics for liquid metals, Phys. Rev. B 47, 558 (1993)

  61. [61]

    Kresse and J

    G. Kresse and J. Hafner, Ab initio molecular- dynamics simulation of the liquid-metal–amorphous- semiconductor transition in germanium, Phys. Rev. B 49, 14251 (1994)

  62. [62]

    Kresse and J

    G. Kresse and J. Furthm¨ uller, Efficient iterative schemes for ab initio total-energy calculations using a plane- wave basis set, Phys. Rev. B 54, 11169 (1996)

  63. [63]

    Kresse and J

    G. Kresse and J. Furthm¨ uller, Efficiency of ab-initio to- tal energy calculations for metals and semiconductors using a plane-wave basis set, Comput. Mater. Sci. 6, 15 (1996)

  64. [64]

    P. E. Bl¨ ochl, Projector augmented-wave method, Phys. Rev. B 50, 17953 (1994)

  65. [65]

    H. J. Monkhorst and J. D. Pack, Special points for Brillouin-zone integrations, Phys. Rev. B 13, 5188 (1976)

  66. [66]

    P. E. Bl¨ ochl, O. Jepsen, and O. K. Andersen, Im- proved tetrahedron method for Brillouin-zone integra- tions, Phys. Rev. B 49, 16223 (1994)

  67. [67]

    D. C. Langreth and M. J. Mehl, Beyond the local- density approximation in calculations of ground-state 13 electronic properties, Phys. Rev. B 28, 1809 (1983)

  68. [68]

    A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. A. Persson, Commentary: The Ma- terials Project: A materials genome approach to ac- celerating materials innovation, APL Mater. 1, 011002 (2013)

  69. [69]

    Klimeˇ s, D

    J. Klimeˇ s, D. R. Bowler, and A. Michaelides, Chemical accuracy for the van der Waals density functional, J. Phys.: Condens. Matter 22, 022201 (2009)

  70. [70]

    Klimeˇ s, D

    J. Klimeˇ s, D. R. Bowler, and A. Michaelides, Van der Waals density functionals applied to solids, Phys. Rev. B 83, 195131 (2011)

  71. [71]

    J. Ning, M. Kothakonda, J. W. Furness, A. D. Ka- plan, S. Ehlert, J. G. Brandenburg, J. P. Perdew, and J. Sun, Workhorse minimally empirical dispersion- corrected density functional with tests for weakly bound systems: r 2SCAN + rVV10, Phys. Rev. B 106, 075422 (2022)

  72. [72]

    Chakraborty, K

    D. Chakraborty, K. Berland, and T. Thonhauser, Next- Generation Nonlocal van der Waals Density Functional, J. Chem. Theory Comput. 16, 5893 (2020)

  73. [73]

    Batzner, A

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun. 13, 2453 (2022)

  74. [74]

    Passaro and C

    S. Passaro and C. L. Zitnick, Reducing SO(3) Convolu- tions to SO(2) for Efficient Equivariant GNNs, in Proc. Mach. Learn. Res. (PMLR), Vol. 202 (PMLR, 2023) pp. 27424–27438

  75. [75]

    T. W. Ko, B. Deng, M. Nassar, L. Barroso-Luque, R. Liu, J. Qi, A. C. Thakur, A. R. Mishra, E. Liu, G. Ceder, S. Miret, and S. P. Ong, Materials Graph Library (MatGL), an open-source graph deep learning library for materials science and chemistry, npj Comput. Mater. 11, 253 (2025)

  76. [76]

    A. H. Larsen, J. J. Mortensen, J. Blomqvist, I. E. Castelli, R. Christensen, M. Du lak, J. Friis, M. Groves, B. Hammer, C. Hargus, E. Hermes, P. C. Jennings, P. B. Jensen, J. Kermode, J. R. Kitchin, E. Kolsbjerg, J. Kubal, K. Kaasbjerg, S. Lysgaard, J. B. Marons- son, T. Maxson, T. Olsen, L. Pastewka, A. T. Peterson, C. Rostgaard, J. Schiøtz, O. Sch¨ utt,...

  77. [77]

    Thorn, J

    A. Thorn, J. Rojas-Nunez, S. Hajinazar, S. E. Baltazar, and A. N. Kolmogorov, Toward ab Initio Ground States of Gold Clusters via Neural Network Modeling, J. Phys. Chem. C 123, 30088 (2019)

  78. [78]

    A. N. Kolmogorov, S. Shah, E. R. Margine, et al., New Superconducting and Semiconducting Fe-B Compounds Predicted with an Ab Initio Evolutionary Search, Phys. Rev. Lett. 105, 217003 (2010)

  79. [79]

    Gochitashvili, M

    D. Gochitashvili, M. Meyers, C. Wang, and A. N. Kol- mogorov, Improving structure search with hyperspatial optimization and TETRIS seeding, Phys. Chem. Chem. Phys. 27, 25636 (2025)

  80. [80]

    Van Der Geest and A

    A. Van Der Geest and A. Kolmogorov, Stability of 41 metal–boron systems at 0GPa and 30GPa from first principles, Calphad 46, 184 (2014)

Showing first 80 references.