Pith. sign in

REVIEW 2 major objections 6 minor 79 references

Equivariant foundation potentials can be both fast enough for large-scale MD and accurate enough to match leading community benchmarks, while materials-discovery errors are limited more by data diversity and inconsistent transition-metal en

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 06:39 UTC pith:AXETT33A

load-bearing objection Solid engineering paper: fast equivariant foundation checkpoints on the accuracy–speed front, with a useful but still correlational diagnosis of the discovery bottleneck. the 2 major comments →

arxiv 2607.28461 v1 pith:AXETT33A submitted 2026-07-30 physics.comp-ph cond-mat.mtrl-sciphysics.app-phphysics.chem-ph

Fast and Accurate Foundation Models for Equivariant Machine-Learned Interatomic Potentials

classification physics.comp-ph cond-mat.mtrl-sciphysics.app-phphysics.chem-ph
keywords machine-learned interatomic potentialsequivariant graph neural networksfoundation modelsNequIPAllegromaterials discoverythermal conductivitymolecular dynamics scaling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Universal machine-learned interatomic potentials are now used as foundation models that get fine-tuned for specific chemistry, but scientific work such as molecular dynamics needs both high accuracy and high speed. This paper shows that equivariant graph models—architectures that bake rotation, translation, and inversion symmetry into the network—can hit both targets even on extremely large training sets where data efficiency matters less. Using software accelerations (graph compilation, faster tensor-product kernels, mixed precision, and multi-GPU training), the authors train a family of NequIP and Allegro foundation models on progressively larger inorganic datasets up to roughly 113 million structures. The largest models match top published accuracies on materials discovery, thermal conductivity, and near-equilibrium mechanical and thermodynamic tests, while defining the accuracy–speed Pareto front and scaling molecular dynamics to tens of millions of atoms on modest GPU counts. Separately, the authors argue that further gains on materials discovery will come less from bigger models or more of the same rattled geometries, and more from broader chemical diversity and consistent treatment of transition-metal energy surfaces.

Core claim

A family of NequIP and Allegro equivariant foundation potentials, trained on MPtrj, Alexandria, and OMat24 (OAM), achieves leading molecular-dynamics inference speeds and strong multi-GPU scalability while matching the best community accuracies across materials discovery, thermal conductivity, and MatCalc near-equilibrium benchmarks. Training accelerations cut the cost of high-accuracy foundation training on ultra-large data to a few hundred GPU-hours. Materials-discovery accuracy saturates when adding OMat24 because that set adds off-equilibrium geometries without new compositions and inherits inconsistent Hubbard-U labeling on multi-valent transition metals.

What carries the argument

Equivariant NequIP (atom-centred message passing) and Allegro (strictly local pair-centred) architectures, combined with train-time graph compilation, accelerated tensor-product kernels, mixed-precision training, distributed data-parallel training, and a two-stage force-then-energy loss schedule, which together make large equivariant foundation models both trainable at scale and fast at inference.

Load-bearing premise

That the high per-element discovery errors on first-row transition metals and the weak accuracy gain from MPA to OAM mainly prove that inconsistent transition-metal energy labels and missing chemical diversity—not residual model limits or the discovery benchmark’s construction—are the dominant bottleneck.

What would settle it

Retrain the same large NequIP/Allegro models on a version of the OAM data with fully consistent Hubbard-U (or multi-fidelity heads for distinct +U schemes) plus substantially new transition-metal compositions beyond ionic substitution; if materials-discovery convex-hull MAE does not drop well below ~20 meV/atom while thermal-conductivity gains stay similar, the bottleneck claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • High-accuracy equivariant foundation potentials become practical starting points for fine-tuning because training cost falls to hundreds of GPU-hours and inference leads the accuracy–speed frontier.
  • Message-passing NequIP models can run multi-GPU MD at ~100M-atom scale on modest node counts via ghost-atom exchange interfaces.
  • Dataset builders targeting discovery should prioritise new compositions and consistent d/f-electron treatments rather than only more rattled copies of existing structures.
  • Dataset builders targeting transport and finite-temperature properties should keep prioritising off-equilibrium geometries, which continue to improve force-constant-derived metrics under power-law scaling.
  • Architectures that adapt radial/angular resolution to local chemistry could avoid uniformly oversized capacity for simple versus correlated bonding.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If multi-fidelity heads cleanly separate +U and non-+U energy surfaces, many existing DFT trajectories could be reused without full recalculation, changing the economics of foundation-data curation.
  • The stronger thermal-conductivity scaling versus discovery scaling suggests community leaderboards may increasingly split into ‘equilibrium ranking’ versus ‘dynamics/smoothness’ tracks with different optimal data recipes.
  • Allegro’s weaker κSRME learning rates relative to NequIP, despite comparable force errors, point to a testable difference between pair-local and message-passing receptive fields for higher-order force constants.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The manuscript presents a family of equivariant foundation MLIPs in the NequIP and Allegro architectures (S/M/L/XL), trained on DIRECT, MPtrj, MPA, and OAM (OMat24 pretrain then MPA fine-tune) datasets. Leveraging NequIP infrastructure accelerations (graph compilation, OpenEquivariance/cuEquivariance kernels, mixed precision, DDP), the authors report 5–10× train-time speedups and training costs of ~100–750 H100/H200 GPU-hours on >100M structures. Models are evaluated on matbench-discovery (EWBM MAE, RMSD), κSRME thermal conductivity, and MatCalc near-equilibrium/off-equilibrium metrics, with NequIP-OAM-XL among the strongest reported (e.g. EWBM MAE 19.7 meV/atom, κSRME 0.129). Model- and data-size learning curves, ASE/LAMMPS accuracy–speed Pareto comparisons, and multi-GPU strong/weak scaling to ~100M atoms are provided. A secondary analysis attributes limited materials-discovery gains from MPA→OAM primarily to insufficient chemical diversity and inconsistent selective Hubbard-U labeling of transition-metal compounds rather than further scale alone.

Significance. If the reported speed/accuracy trade-off and scaling hold under community reuse, this is a substantial practical contribution: open, fine-tunable equivariant foundation weights that sit on the MD inference Pareto front while remaining competitive on standard discovery and phonon-related benchmarks, plus infrastructure that lowers the cost of training on ultra-large datasets to individual-group scale. The multi-GPU ML-IAP path for message-passing NequIP and the packaged hyperparameters/models are concrete, reusable deliverables. The element-wise error and MPA→OAM differential analysis is a useful, falsifiable steer for dataset design even if not fully causal. Strengths include external community benchmarks, public comparison setups, learning-curve fits with reported α/R², and explicit scaling measurements on Perlmutter and Frontier.

major comments (2)
  1. [Abstract; §IV.A; Figs. 2–4; S3] Abstract and §IV.A present as a main result that materials-discovery accuracy “should focus on dataset diversity and improved, consistent descriptions of transition metal compound energy surfaces,” based on Fig. 4 element-wise EWBM MAE, weak MPA→OAM discovery improvement versus strong κSRME gains (Figs. 2–3), OMat24’s lack of new compositions under the same selective +U scheme, and similar patterns for eSEN/MACE (S3). This is well-motivated but correlational: confounders include WBM ionic-substitution prototype construction, residual radial/angular capacity on directional d/f bonding, and other DFT label noise. Excluding V/Cr/Mn/Fe/Np/Pu (19.7→13.9 meV/atom) does not isolate fixability by consistency or diversity. Please qualify the Abstract to match the body’s “suggests” language, and state explicitly that no controlled retrain on consistent-+U or multi-fidelity TM labels was performed—
  2. [§IV.B; Table 2; Figs. 2–3, 5; S4] §IV.B notes Allegro’s systematically worse κSRME and shallower learning-curve exponents (α≈0.33/0.11 vs NequIP 0.48/0.24) despite comparable force validation errors, and leaves open strictly-local receptive field versus PES smoothness. Given that CPS weights thermal conductivity heavily (5:4:1) and Fig. 5 places Allegro on or near the front partly via that score, the manuscript should either (i) report a short controlled check (e.g. larger effective cutoff / more layers at matched cost, or displacement-sensitivity table beyond the brief remark) or (ii) clearly separate NequIP vs Allegro claims in the Abstract/Conclusion so “excellent accuracies” on phonon-derived properties is not read as equal for both families.
minor comments (6)
  1. [Table 1; Figs. 2–3] Table 1 lists Allegro Extra-Large as absent while NequIP has XL; naming in figures sometimes mixes {dataset}-L vs XL. A single schematic of which sizes exist per architecture would reduce confusion.
  2. [§IV.D; Fig. 5] Fig. 5 CPS definition (5:4:1 discovery:κSRME:RMSD) is version-dependent; state the matbench-discovery CPS version/date used so comparisons remain reproducible as the leaderboard evolves.
  3. [§VII] Methods: FIRE force threshold 0.015 eV/Å and 0.035 Å phonon displacement are appropriate but should be flagged when comparing to other entries that may use different defaults.
  4. [§VII] Parity restriction (parity: false) is justified empirically (~1.5–2× speed, <3% accuracy); a one-sentence note that this remains SO(3)-equivariant but is not full O(3) with independent parity channels would help non-e3nn readers.
  5. [§VI; S6] Supplementary S6 type-embedding PCA/UMAP/t-SNE is valuable; consider moving one labelled panel into the main text or citing it in the Conclusion where chemical organisation is claimed.
  6. [Throughout; §VII] Minor copy-editing: “ad-ditional”, “train-time speed-ups” hyphenation consistency; ensure arXiv/DOI placeholders (zenodo-to-publish) are final before journal production.

Circularity Check

0 steps flagged

No significant circularity: external DFT datasets and community benchmarks; descriptive scaling fits; self-citations are infrastructure lineage only.

full rationale

This is an empirical ML-engineering paper. Models are trained on public first-principles datasets (MPtrj, Alexandria/sAlex, OMat24) and scored on independent community suites (matbench-discovery WBM hull energies/RMSD, Póta et al. κSRME, MatCalc near-equilibrium and force-magnitude tests). Reported accuracies, MD timesteps/s, and multi-GPU scaling are measured outputs, not quantities defined from the models’ own fitted parameters. Power-law learning curves (L = a · x^{-α} on model size and dataset size) are post-hoc descriptive fits to those measured errors, not recycled as theory or as predictions of held-out targets. Self-citations to NequIP, Allegro, and the NequIP infrastructure (graph compilation, OpenEquivariance/cuEquivariance, ML-IAP) document implementation lineage and prior architecture, not a uniqueness theorem or load-bearing proof of accuracy. The secondary claim that materials-discovery error is limited by chemical diversity and inconsistent Hubbard-U TM surfaces is a correlational inference from element-wise MAE and MPA→OAM differentials; that may be causally under-supported, but it is not circular—the claim does not reduce by construction to a fitted input or a self-defined quantity. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggling chain is present.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 2 invented entities

The work is empirical ML engineering on top of established equivariant GNN potentials and public DFT corpora. Load-bearing background includes SO(3) equivariance as an architectural prior, DFT (with Materials Project-style +U) as the training/test oracle, and community benchmark definitions as proxies for discovery and transport utility. Many architectural and optimization hyperparameters are chosen by hand; no new physical entities are postulated.

free parameters (6)
  • Model-size hyperparameter sets (cutoff, l_max, feature multiplicities, layers, MLP depth) = e.g. NequIP-XL: 6 Å, lmax=4, 320/96/64/32/32, 6 layers, 32.1M params
    Table 1 defines S/M/L/XL configurations that determine capacity–speed trade-offs; chosen to span the Pareto front rather than derived.
  • Energy:force:stress loss weights and two-stage reweighting = 1:5:0.1 then 1:1:0.1 (plus dataset-specific LR/weight-decay/grad-clip)
    Initial 1:5:0.1 then continuation 1:1:0.1 with LR drop; directly shapes what the optimizer prioritizes on large data.
  • Mixed-precision casting policy (float64 embeddings/readout, float32 tensor products, AMP/TF32 in training only) = train mixed; inference full float32
    Engineering choice reported to preserve accuracy in training but disabled at inference due to force noise; affects claimed train-time speedups.
  • Phonon finite-difference displacement for κSRME = 0.035 Å
    Displacement 0.035 Å (Methods) affects third-order force constants and reported thermal-conductivity errors, especially for Allegro sensitivity notes.
  • Geometry-relaxation convergence settings for matbench-discovery = fmax=0.015 eV/Å, max 200 steps
    FIRE, fmax 0.015 eV/Å, max 200 steps; changes EWBM/RMSD metrics used in CPS and discovery claims.
  • Power-law learning-curve exponents α = e.g. NequIP model-size α≈0.21 (EWBM), 0.48 (κSRME)
    Descriptive least-squares fits L=a·x^{-α} on model/data size; not physical constants but used to narrate scaling and saturation.
axioms (5)
  • domain assumption Rotation/inversion/translation equivariance encoded in the network is a valid and beneficial inductive bias for interatomic potentials.
    Stated in Introduction as motivation for NequIP/Allegro; accuracy claims are interpreted within this model class.
  • domain assumption Semi-local DFT energies/forces/stresses (Materials Project-compatible settings, including selective Hubbard U) are the appropriate supervised targets and benchmark references.
    All training sets and matbench/MatCalc/κ comparisons are defined against this oracle; +U inconsistency is discussed as label pathology, not abandonment of DFT.
  • domain assumption Community metrics (WBM hull MAE/RMSD, κSRME, MatCalc moduli/Cv, f/fDFT, CPS weighting 5:4:1) are meaningful proxies for materials-discovery and dynamical fidelity.
    Section IV frames excellence via these suites; Pareto and ‘leading’ claims are with respect to them.
  • standard math Standard ML optimization mathematics (AdamW, Huber losses, DDP data parallelism, SWA) behaves as usual; no new learning theory is required.
    Methods rely on these without proof obligations beyond empirical validation error.
  • ad hoc to paper Restricting irrep parity to spherical-harmonic parity preserves sufficient equivariant expressivity for the reported tasks.
    Methods: parity:false for ~1.5–2× speed with <3% accuracy effect; architectural choice specific to these runs.
invented entities (2)
  • NequIP/Allegro-OAM-{S,M,L,XL} foundation model family (and MP/DIRECT/MPA variants) independent evidence
    purpose: Concrete pre-trained weights offered as fine-tuning starting points and speed/accuracy baselines.
    Not a new physical object; a named artifact class. Independent use is possible once checkpoints are public.
  • OAM training protocol (pretrain OMat24 then fine-tune MPA for MP settings) no independent evidence
    purpose: Reconcile large off-equilibrium OMat24 data with Materials Project-compatible discovery benchmarks.
    Procedural construct; success judged only via the same benchmarks it is tuned to serve.

pith-pipeline@v1.2.0-daily-grok45 · 34670 in / 4038 out tokens · 85399 ms · 2026-07-31T06:39:10.450551+00:00 · methodology

0 comments
read the original abstract

Machine-learned interatomic potentials (MLIPs) have emerged as a transformative tool for computational materials science and chemistry, with universal potentials trained on large and diverse datasets now routinely deployed as 'foundation models' for downstream fine-tuning in targeted chemical spaces. Many scientific applications of the resulting models, such as molecular dynamics (MD), require high inference and training speeds as well as accuracy. In this work, we examine the limits of equivariant MLIPs, which directly encode physical symmetries in model architectures, to achieve these competing targets -- particularly in the regime of extremely large datasets where data efficiency is less critical. We show how this trade-off can be addressed, and present a family of foundation potentials in the NequIP and Allegro equivariant MLIP architectures which achieve leading inference speeds and strong scalability as well as excellent accuracies across a range of community benchmarks -- spanning materials discovery, thermal conductivity prediction, and near-equilibrium mechanical and thermodynamic properties. Accelerations implemented within the NequIP infrastructure now permit training of high-accuracy foundation potentials on ultra-large datasets with dramatically reduced computational cost. Alongside, we show that efforts to improve model accuracy for materials discovery should focus on dataset diversity and improved, consistent descriptions of transition metal compound energy surfaces.

Figures

Figures reproduced from arXiv: 2607.28461 by Albert Musaelian, Anders Johansson, Austin Glover, Boris Kozinsky, Chuin Wei Tan, Francesco Libbi, Gabriel de Miranda Nascimento, Laura Zichi, Marc L. Descoteaux, Menghang Wang, Norma Rivano, Se\'an R. Kavanagh, Ulrik Unneberg, Vivek Bharadwaj, William C. Witt.

Figure 1
Figure 1. Figure 1: Accelerated Model Training. Training speeds for large NequIP and Allegro models with the MPtrj dataset,4 plotted as the number of data frames (atomic structures with energy and forces) trained on per GPU-hour. Training speeds are measured without and with graph compilation (torch.compile), using standard tensor product kernels or the recently-implemented OpenEquivariance24 (for NequIP) and cuEquivariance25… view at source ↗
Figure 2
Figure 2. Figure 2: Model Size Learning Curve. Test errors for NequIP and Allegro models trained on the OAM dataset (NequIP/Allegro-OAM- {model size}) and evaluated on the Matbench Discovery26 benchmark suite, as a function of model size (Tables 1 and 2). This includes the mean absolute error (MAE) in convex hull energy predictions for unseen compounds (EWBM, MAE, left) and the dimensionless Symmetric Relative Mean Error in t… view at source ↗
Figure 3
Figure 3. Figure 3: Training Data Learning Curve. Test errors for the largest NequIP and Allegro models trained in this work (NequIP-{dataset}-XL and Allegro-{dataset}-L) and evaluated on the Matbench Discovery26 benchmark suite, as a function of training dataset size. This includes the mean absolute error (MAE) in convex hull energy predictions for unseen compounds (EWBM, MAE, left) and the dimensionless Symmetric Relative M… view at source ↗
Figure 4
Figure 4. Figure 4: Per-Element Errors for Materials Discovery. Distribution of element-wise convex hull energy mean absolute error (MAE) across the periodic table for the matbench-discovery26 materials discovery benchmark, using the NequIP-OAM-XL model. Errors are shown in meV/atom, using a logarithmic colourbar scale. The largest element-wise errors are seen for compounds containing multi-valent first-row transition metals … view at source ↗
Figure 5
Figure 5. Figure 5: Accuracy-Speed Trade-off. matbench-discovery26 ‘Combined Performance Score’ (CPS) plotted against Molecular Dynamics (MD) inference speed for Si diamond systems ranging from 64 to 13,824 atoms, using a single NVIDIA A100-SXM4 (80 GB) GPU. MD simulations were performed with both the Atomic Simulation Environment (ASE)48 (hollow markers) and LAMMPS49 (filled markers) – for models with LAMMPS interfaces – usi… view at source ↗
Figure 6
Figure 6. Figure 6: Multi-GPU Scaling with NequIP. Strong (top) and weak (bottom) scaling performance of the large and small NequIP models (NequIP-OAM-L and NequIP-OAM-S) in LAMMPS49 using the ML-IAP interface52 with OpenEquivariance24 kernels and torch.compile. MD simulations were performed using diamond-structured silicon systems ranging from 32.8k to 102.5M atoms, with NVIDIA A100 (80 GB) and AMD MI250X GPUs on the Perlmut… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 13 linked inside Pith

  1. [1]

    Behler and M

    J. Behler and M. Parrinello, Physical Review Letters98, 146401 (2007)

  2. [2]

    A. M. Mroz, A. R. Basford, F. Hastedt, I. S. Jayasekera, I. Mosquera-Lois, R. Sedgwick, P. J. Ballester, J. D. Bocarsly, E. A. d. R. Chanona, M. L. Evans, J. M. Frost, A. M. Ganose, R. L. Greenaway, K. K. M. Hii, Y . Li, R. Misener, A. Walsh, D. Zhang, and K. E. Jelfs, Chemical Society Reviews 54, 5433 (2025)

  3. [3]

    Mannodi-Kanakkithodi, M

    A. Mannodi-Kanakkithodi, M. Huang, P. Gorai, and S. R. Ka- vanagh, MRS Bulletin 51, 600 (2026)

  4. [4]

    B. Deng, P. Zhong, K. Jun, J. Riebesell, K. Han, C. J. Bartel, and G. Ceder, Nature Machine Intelligence 5, 1031 (2023)

  5. [5]

    Schmidt, T

    J. Schmidt, T. F. T. Cerqueira, A. H. Romero, A. Loew, F. J¨ager, H.-C. Wang, S. Botti, and M. A. L. Marques, Materials Today Physics 48, 101560 (2024)

  6. [6]

    Barroso-Luque, M

    L. Barroso-Luque, M. Shuaibi, X. Fu, B. M. Wood, 14 M. Dzamba, M. Gao, A. Rizvi, C. L. Zitnick, and Z. W. Ulissi, Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models (2024), arXiv:2410.12771 [cond-mat]

  7. [7]

    A. D. Kaplan, R. Liu, J. Qi, T. W. Ko, B. Deng, J. Riebe- sell, G. Ceder, K. A. Persson, and S. P. Ong, A Founda- tional Potential Energy Surface Dataset for Materials (2025), arXiv:2503.04070 [cond-mat]

  8. [8]

    Mazitov, F

    A. Mazitov, F. Bigi, M. Kellner, P. Pegolo, D. Tisi, G. Fraux, S. Pozdnyakov, P. Loche, and M. Ceriotti, Nature Communica- tions 16, 10653 (2025)

  9. [9]

    D. S. Levine, M. Shuaibi, E. W. C. Spotte-Smith, M. G. Taylor, M. R. Hasyim, K. Michel, I. Batatia, G. Cs ´anyi, M. Dzamba, P. Eastman, N. C. Frey, X. Fu, V . Gharakhanyan, A. S. Kr- ishnapriyan, J. A. Rackers, S. Raja, A. Rizvi, A. S. Rosen, Z. Ulissi, S. Vargas, C. L. Zitnick, S. M. Blau, and B. M. Wood, The Open Molecules 2025 (OMol25) Dataset, Evaluat...

  10. [10]

    H. Kaur, F. D. Pia, I. Batatia, X. R. Advincula, B. X. Shi, J. Lan, G. Cs´anyi, A. Michaelides, and V . Kapil, Faraday Discussions 256, 120 (2025)

  11. [11]

    Mosquera-Lois, S

    I. Mosquera-Lois, S. R. Kavanagh, A. M. Ganose, and A. Walsh, npj Computational Materials 10, 1 (2024)

  12. [12]

    J. L. A. Gardner, H. Schulz, J. Helie, L. Sun, and G. N. C. Simm, Understanding multi-fidelity training of machine- learned force-fields (2026), arXiv:2506.14963 [physics.chem- ph]

  13. [13]

    J. Kim, J. Kim, J. Kim, J. Lee, Y . Park, Y . Kang, and S. Han, Journal of the American Chemical Society 147, 1042 (2025)

  14. [14]

    Batzner, A

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, Na- ture Communications 13, 2453 (2022)

  15. [15]

    Musaelian, S

    A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Nature Communications 14, 579 (2023)

  16. [16]

    Batzner, A

    S. Batzner, A. Musaelian, and B. Kozinsky, Nature Reviews Physics 5, 437 (2023)

  17. [17]

    Batatia, D

    I. Batatia, D. P. Kov ´acs, G. N. C. Simm, C. Ortner, and G. Cs´anyi, MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields (2023), arXiv:2206.07697 [cond-mat, physics:physics, stat]

  18. [18]

    Bochkarev, Y

    A. Bochkarev, Y . Lysogorskiy, and R. Drautz, Physical Review X 14, 021036 (2024)

  19. [19]

    Merchant, S

    A. Merchant, S. Batzner, S. S. Schoenholz, M. Aykol, G. Cheon, and E. D. Cubuk, Nature 624, 80 (2023)

  20. [20]

    Y . Park, J. Kim, S. Hwang, and S. Han, Journal of Chemical Theory and Computation 20, 4857 (2024)

  21. [21]

    Batatia, S

    I. Batatia, S. Batzner, D. P. Kov ´acs, A. Musaelian, G. N. C. Simm, R. Drautz, C. Ortner, B. Kozinsky, and G. Cs ´anyi, Na- ture Machine Intelligence 7, 56 (2025)

  22. [22]

    Kozinsky, A

    B. Kozinsky, A. Musaelian, A. Johansson, and S. Batzner, in Proceedings of the International Conference for High Perfor- mance Computing, Networking, Storage and Analysis , SC ’23 (Association for Computing Machinery, New York, NY , USA,

  23. [23]

    C. W. Tan, M. L. Descoteaux, M. Kotak, G. d. M. Nasci- mento, S. R. Kavanagh, L. Zichi, M. Wang, A. Saluja, Y . R. Hu, T. Smidt, A. Johansson, W. C. Witt, B. Kozinsky, and A. Musaelian, Digital Discovery 10.1039/D5DD00423C (2026)

  24. [25]

    Accelerate Drug and Material Discovery with New Math Li- brary NVIDIA cuEquivariance (2024)

  25. [26]

    Riebesell, R

    J. Riebesell, R. E. A. Goodall, P. Benner, Y . Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, Na- ture Machine Intelligence 7, 836 (2025)

  26. [27]

    P ´ota, P

    B. P ´ota, P. Ahlawat, G. Cs ´anyi, and M. Simoncelli, Thermal Conductivity Predictions with Foundation Atomistic Models (2024), arXiv:2408.00755 [cond-mat] version: 4

  27. [29]

    Kharya, NVIDIA Blogs: TensorFloat-32 Accelerates AI Training HPC upto 20x (2020)

    P. Kharya, NVIDIA Blogs: TensorFloat-32 Accelerates AI Training HPC upto 20x (2020)

  28. [31]

    Cheng, npj Computational Materials 11, 80 (2025)

    B. Cheng, npj Computational Materials 11, 80 (2025)

  29. [32]

    T. W. Ko, R. Liu, A. R. Mishra, Z. Yu, J. Qi, and S. P. Ong, A Fast, Accurate, and Reactive Equivariant Foundation Potential (2025), arXiv:2511.07249 [cond-mat.mtrl-sci]

  30. [33]

    Izmailov, D

    P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson, 34th Conference on Uncertainty in Artificial Intelli- gence 2018, UAI 2018 34th Conference on Uncertainty in Ar- tificial Intelligence 2018, UAI 2018, 876 (2018)

  31. [34]

    J. Qi, T. W. Ko, B. C. Wood, T. A. Pham, and S. P. Ong, npj Computational Materials 10, 43 (2024)

  32. [35]

    A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. a. Persson, APL Materials 1, 011002 (2013)

  33. [36]

    B. Deng, Y . Choi, P. Zhong, J. Riebesell, S. Anand, Z. Li, K. Jun, K. A. Persson, and G. Ceder, npj Computational Ma- terials 11, 9 (2025)

  34. [37]

    H.-C. Wang, S. Botti, and M. A. L. Marques, npj Computational Materials 7, 12 (2021)

  35. [38]

    C. J. Owen, S. B. Torrisi, Y . Xie, S. Batzner, K. Bystrom, J. Coulter, A. Musaelian, L. Sun, and B. Kozinsky, npj Com- putational Materials 10, 92 (2024)

  36. [39]

    Warford, F

    T. Warford, F. L. Thiemann, and G. Cs ´anyi, Better without U: Impact of Selective Hubbard U Correction on Foundational MLIPs (2026), arXiv:2601.21056 [physics] version: 1

  37. [40]

    Mazitov, S

    A. Mazitov, S. Chorna, G. Fraux, M. Bercx, G. Pizzi, S. De, and M. Ceriotti, Scientific Data 12, 1857 (2025)

  38. [41]

    Rhodes, S

    B. Rhodes, S. Vandenhaute, V . ˇSimkus, J. Gin, J. Godwin, T. Duignan, and M. Neumann, Orb-v3: atomistic simulation at scale (2025), arXiv:2504.06231 [cond-mat]

  39. [42]

    Simeon and G

    G. Simeon and G. De Fabritiis, in Advances in Neural Infor- mation Processing Systems , V ol. 36 (Curran Associates, Inc.,

  40. [43]

    H. Yang, C. Hu, Y . Zhou, X. Liu, Y . Shi, J. Li, G. Li, Z. Chen, S. Chen, C. Zeni, M. Horton, R. Pinsler, A. Fowler, D. Z¨ugner, T. Xie, J. Smith, L. Sun, Q. Wang, L. Kong, C. Liu, H. Hao, and Z. Lu, MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures (2024), arXiv:2405.04967 [cond-mat.mtrl-sci]

  41. [44]

    B. M. Wood, M. Dzamba, X. Fu, M. Gao, M. Shuaibi, L. Barroso-Luque, K. Abdelmaqsoud, V . Gharakhanyan, J. R. Kitchin, D. S. Levine, K. Michel, A. Sriram, T. Cohen, A. Das, A. Rizvi, S. J. Sahoo, Z. W. Ulissi, and C. L. Zit- nick, UMA: A Family of Universal Models for Atoms (2026), arXiv:2506.23971 [cs.LG]

  42. [45]

    R. Tran, J. Lan, M. Shuaibi, B. M. Wood, S. Goyal, A. Das, J. Heras-Domingo, A. Kolluru, A. Rizvi, N. Shoghi, A. Sriram, F. Therrien, J. Abed, O. V oznyy, E. H. Sargent, Z. Ulissi, and C. L. Zitnick, ACS Catalysis 13, 3066 (2023)

  43. [46]

    Sriram, L

    A. Sriram, L. M. Brabson, X. Yu, S. Choi, K. Abdelmaqsoud, E. Moubarak, P. d. Haan, S. L ¨owe, J. Brehmer, J. R. Kitchin, M. Welling, C. L. Zitnick, Z. Ulissi, A. J. Medford, and D. S. 15 Sholl, The Open DAC 2025 Dataset for Sorbent Discovery in Direct Air Capture (2025), arXiv:2508.03162 [cond-mat]

  44. [47]

    Gharakhanyan, L

    V . Gharakhanyan, L. Barroso-Luque, Y . Yang, M. Shuaibi, K. Michel, D. S. Levine, M. Dzamba, X. Fu, M. Gao, X. Liu, H. Ni, K. Noori, B. M. Wood, M. Uyttendaele, A. Boromand, C. L. Zitnick, N. Marom, Z. W. Ulissi, and A. Sriram, Open Molecular Crystals 2025 (OMC25) Dataset and Models (2025), arXiv:2508.02651 [physics]

  45. [48]

    A. H. Larsen, J. J. Mortensen, J. Blomqvist, I. E. Castelli, R. Christensen, M. Du \lak, J. Friis, M. N. Groves, B. Ham- mer, C. Hargus, E. D. Hermes, P. C. Jennings, P. B. Jensen, J. Kermode, J. R. Kitchin, E. L. Kolsbjerg, J. Kubal, K. Kaas- bjerg, S. Lysgaard, J. B. Maronsson, T. Maxson, T. Olsen, L. Pastewka, A. Peterson, C. Rostgaard, J. Schiøtz, O. ...

  46. [49]

    A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in ’t Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton, Computer Physics Communica- tions 271, 108171 (2022)

  47. [50]

    J. Moon, U. Jeon, S. Choung, and J. W. Han, Cell Reports Phys- ical Science 6, 10.1016/j.xcrp.2025.102968 (2025)

  48. [51]

    Chiang, T

    Y . Chiang, T. Kreiman, C. Zhang, M. C. Kuner, E. Weaver, I. Amin, H. Park, Y . Lim, J. Kim, D. Chrzan, A. Walsh, S. M. Blau, M. Asta, and A. S. Krishnapriyan, MLIP Arena: Ad- vancing Fairness and Transparency in Machine Learning Inter- atomic Potentials via an Open, Accessible Benchmark Platform (2025), arXiv:2509.20630 [physics]

  49. [53]

    Nguyen-Cong, J

    K. Nguyen-Cong, J. T. Willman, S. G. Moore, A. B. Be- lonoshko, R. Gayatri, E. Weinberg, M. A. Wood, A. P. Thomp- son, and I. I. Oleynik, in Proceedings of the International Con- ference for High Performance Computing, Networking, Storage and Analysis , SC ’21 (Association for Computing Machinery, New York, NY , USA, 2021) pp. 1–12

  50. [54]

    A. P. Thompson, L. P. Swiler, C. R. Trott, S. M. Foiles, and G. J. Tucker, Journal of Computational Physics 285, 316 (2015)

  51. [55]

    D. Kim, X. Wang, S. Vargas, P. Zhong, D. S. King, T. J. Inizan, and B. Cheng, Journal of Chemical Theory and Computation 21, 12709 (2025)

  52. [56]

    M. U. Maruf, S. Kim, and Z. Ahmad, The Journal of Physical Chemistry Letters 16, 9078 (2025)

  53. [57]

    Falletta, A

    S. Falletta, A. Cepellotti, A. Johansson, C. W. Tan, M. L. De- scoteaux, A. Musaelian, C. J. Owen, and B. Kozinsky, Nature Communications 16, 4031 (2025)

  54. [58]

    G. d. M. Nascimento, M. L. Descoteaux, L. Zichi, C. W. Tan, W. C. Witt, N. Molinari, S. Mantha, D. Kitchaev, M. Ko- rnbluth, K. Gadelrab, C. Tuffile, and B. Kozinsky, Mixture of Experts Framework in Machine Learning Interatomic Po- tentials for Atomistic Simulations (2026), arXiv:2604.26143 [physics.comp-ph]

  55. [59]

    L. A. Gomes, S. Larmore, M. Wang, C. W. Tan, B. Kozin- sky, and S. A. Lopez, ChemRxiv 2026, 10.26434/chem- rxiv.15001631/v1

  56. [60]

    S. R. Kavanagh, Journal of Physics: Energy 7, 045002 (2025)

  57. [61]

    LMDB: Lightning Memory-Mapped Database Manager (LMDB)

  58. [62]

    Ocampo, D

    D. Ocampo, D. Posso, R. Namakian, and W. Gao, Computa- tional Materials Science 244, 113155 (2024)

  59. [63]

    Kuryla, F

    D. Kuryla, F. Berger, G. Cs ´anyi, and A. Michaelides, How Ac- curate Are DFT Forces? Unexpectedly Large Uncertainties in Molecular Datasets (2025), arXiv:2510.19774 [physics.chem- ph]

  60. [64]

    S. J. Sahoo, M. Maraschin, J. B. Varley, D. S. Levine, Z. Ulissi, C. L. Zitnick, W. Takemura, J. A. Gauthier, N. Govindarajan, and M. Shuaibi, Insights into CO dimerization at electrified Cu interfaces from large-scale machine learning simulations (2025), version Number: 2

  61. [66]

    Bitzek, P

    E. Bitzek, P. Koskinen, F. G¨ahler, M. Moseler, and P. Gumbsch, Physical Review Letters 97, 170201 (2006)

  62. [69]

    Koker, A

    T. Koker, A. Gangan, M. Kotak, J. Marian, and T. Smidt, PFT: Phonon Fine-tuning for Machine Learned Interatomic Poten- tials (2026), arXiv:2601.07742 [cond-mat.mtrl-sci]

  63. [70]

    Y .-L. Liao, A. J. Hoffman, S. C. Shen, A. Duval, S. W. Nor- wood, and T. Smidt, EquiformerV3: Scaling Efficient, Expres- sive, and General SE(3)-Equivariant Graph Attention Trans- formers (2026), arXiv:2604.09130 [cs.LG]

  64. [72]

    J. Qi, T. W. Ko, B. C. Wood, T. A. Pham, and S. P. Ong, npj Computational Materials10, 43 (2024)

  65. [73]

    Bharadwaj, A

    V . Bharadwaj, A. Glover, A. Buluc, and J. Demmel, An Efficient Sparse Kernel Generator for O(3)-Equivariant Deep Networks (2025), arXiv:2501.13986 [cs]

  66. [74]

    Accelerate Drug and Material Discovery with New Math Library NVIDIA cuEquivariance (2024)

  67. [75]

    BFloat16: The secret to high performance on Cloud TPUs

  68. [76]

    Riebesell, R

    J. Riebesell, R. E. A. Goodall, P. Benner, Y . Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, Nature Machine Intelligence 7, 836 (2025)

  69. [77]

    X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zitnick, Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction (2025), arXiv:2502.12147 [physics]

  70. [78]

    Batatia, P

    I. Batatia, P. Benner, Y . Chiang, A. M. Elena, D. P. Kov´acs, J. Riebesell, X. R. Advincula, M. Asta, M. Avaylon, W. J. Baldwin, F. Berger, N. Bernstein, A. Bhowmik, S. M. Blau, V . C˘arare, J. P. Darby, S. De, F. D. Pia, V . L. Deringer, R. Elijoˇsius, Z. El-Machachi, F. Falcioni, E. Fako, A. C. Ferrari, A. Genreith-Schriever, J. George, R. E. A. Goodal...

  71. [79]

    A. P. Thompson, H. M. Aktulga, R. Berger, D. S. Bolintineanu, W. M. Brown, P. S. Crozier, P. J. in ’t Veld, A. Kohlmeyer, S. G. Moore, T. D. Nguyen, R. Shan, M. J. Stevens, J. Tranchida, C. Trott, and S. J. Plimpton, Computer Physics Communications271, 108171 (2022)

  72. [80]

    Bochkarev, Y

    A. Bochkarev, Y . Lysogorskiy, and R. Drautz, Physical Review X14, 021036 (2024)

  73. [81]

    Lysogorskiy, A

    Y . Lysogorskiy, A. Bochkarev, and R. Drautz, Graph atomic cluster expansion for foundational machine learning interatomic potentials (2026), arXiv:2508.17936 [cond-mat.mtrl-sci]

  74. [82]

    J. Kim, J. You, Y . Park, Y . Lim, Y . Kang, J. Kim, H. Jeon, S. Ju, D. Hong, S. Y . Lee, S. Choi, Y . Kim, J. W. Lee, and S. Han, Optimizing Cross-Domain Transfer for Universal Machine Learning Interatomic Potentials (2025), arXiv:2510.11241 [cond-mat.mtrl-sci]

  75. [83]

    Mazitov, F

    A. Mazitov, F. Bigi, M. Kellner, P. Pegolo, D. Tisi, G. Fraux, S. Pozdnyakov, P. Loche, and M. Ceriotti, Nature Communications 16, 10653 (2025)

  76. [84]

    F. Bigi, P. Pegolo, A. Mazitov, J. Schmidt, and M. Ceriotti, Pushing the limits of unconstrained machine-learned interatomic potentials (2026), arXiv:2601.16195 [physics.chem-ph]

  77. [85]

    Johansson, E

    A. Johansson, E. Weinberg, C. R. Trott, M. J. McCarthy, and S. G. Moore, in Proceedings of the SC ’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (2025) pp. 1217–1232, arXiv:2508.13523 [cs.DC]

  78. [86]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V . Dubourg, J. Van- derplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, Journal of Machine Learning Research 12, 2825 (2011)

  79. [87]

    McInnes, J

    L. McInnes, J. Healy, N. Saul, and L. Großberger, Journal of Open Source Software 3, 861 (2018)