Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?

T0 review · 2 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Benchmarking 80 AI interatomic potentials finds no universal winner for heterogeneous catalysis, with training data often mattering more than architecture and many models failing catastrophically on magnetic surfaces.

desk verdict Large, careful MLIP benchmark for catalysis with one real flaw: the gas-phase offset correction fits the test set, so corrected formation/adsorption rankings are not zero-shot; the main qualitative conclusions still hold. read the letter →

arxiv 2512.16702 v1 pith:2IPH4NE3 submitted 2025-12-18 cond-mat.mtrl-sci cs.LGphysics.chem-ph

classification cond-mat.mtrl-scics.LGphysics.chem-ph
keywords machinelearninginteratomicpotentialsheterogeneouscatalysiszero-shotbenchmarkfoundationmodelsadsorptionenergiestransitionstatesmagneticmaterialsstructurerelaxation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Foundational machine learning interatomic potentials (MLIPs) promise near-DFT accuracy at low cost, but their published benchmarks usually cover ordered bulk crystals, not the surfaces, adsorbates, and transition states that matter for catalysis. This paper systematically tests 80 pretrained MLIPs zero-shot, without fine-tuning, on 13 catalysis-relevant properties across six datasets, including oxide surfaces, alloyed metal surfaces, and metal–oxide interfaces. The central finding is that no single MLIP wins everywhere: performance depends strongly on which dataset the model was trained on, possibly more than on its architecture, so a model that excels at oxide vacancy energies can fail on magnetic alloy surfaces. The paper shows that current-generation models are accurate enough for screening, with best-case errors around 0.1–0.3 eV, but not yet negligible relative to DFT's intrinsic errors, and that relaxing structures within the MLIP typically increases prediction error.

What carries the argument

The central object is a benchmark database: over 700,000 single-point potential-energy evaluations from 80 pretrained MLIPs spanning 16 architecture families, scored against DFT reference values on 6 catalysis-relevant datasets and 13 target properties. The argument is carried by comparing root-mean-square and maximum absolute errors across models, grouped by architecture, training data, and model size, and by isolating dataset-driven trends such as the magnetic-element failure.

What would settle it

Re-run the same evaluation but let each MLIP relax every structure to its own local minimum (rather than only a 10% subset of one dataset) and recalculate RMSEs and rankings; if rankings change substantially or the best models no longer remain on top, the paper's recommendation of which models to start with would need to be revised.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a systematic mapping of zero-shot MLIP accuracy across heterogeneous catalysis tasks. Evaluated on DFT-relaxed geometries, the best models (the eSEN, Orb, and UMA families, trained on datasets containing oxides or mixed inorganic materials) reach RMSEs around 0.1–0.3 eV for activation and formation energies, and as low as 0.01 eV for zero-point vibrational energies of supported nanoclusters. Task-specific models trained directly on the evaluation data can match or beat these MLIPs on accuracy, though they are far less general. A significant fraction of MLIPs exhibit catastrophic errors, tens of eV, on surfaces containing cobalt or nickel, and the pa

Load-bearing premise

The headline accuracy numbers assume the user starts from DFT-relaxed geometries, not from structures relaxed by the MLIP itself, and the paper itself notes this is not a fully realistic use case.

Editorial extensions

If this is right

  • No single MLIP can be chosen a priori for a catalysis problem; users must screen several models for their specific application, and the paper recommends starting with the eSEN, Orb, and UMA families.
  • The strong training-data dependence implies that foundation-model accuracy is set largely by dataset coverage: models trained on data including oxides and spin-polarized systems generalize better to oxide and magnetic surfaces, respectively.
  • Current MLIPs are adequate for initial screening (narrowing candidate materials and mechanisms) but not yet accurate enough to replace DFT in microkinetic modeling, since best-case errors are comparable to intrinsic DFT errors.
  • Relaxing structures with the MLIP rather than evaluating DFT-relaxed geometries raises errors for all models, so reported RMSEs on fixed geometries are optimistic for realistic workflows.
  • Cheap task-specific models (scaling relations, graph-based Gaussian processes, and similar) can compete with the best foundational MLIPs on accuracy for their narrow target, at lower cost and with interpretability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If training data dominates over architecture, then future gains may come as much from broadening and re-weighting training datasets, especially including spin-polarized, magnetic, and disordered surface configurations, as from architectural innovations.
  • The relaxation penalty observed here suggests the MLIP error surface is shifted relative to DFT at off-equilibrium geometries; this points to fine-tuning or delta-learning on the target system as a promising low-cost route to make zero-shot MLIPs reliable for genuine structure searches.
  • The catastrophic magnetic-element failures imply that any MLIP claiming to be universal should be stress-tested on magnetic surfaces before being used for transition-metal catalysis; a simple test on cobalt- or nickel-containing surfaces would be a useful addition to standard benchmarks.
  • Since the benchmark uses DFT-relaxed geometries from different functionals and numerical settings than the training data, small shifts in ranking could occur if the reference calculations were redone with a consistent functional, suggesting a follow-up test with a unified reference.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper benchmarks 80 pretrained machine-learning interatomic potentials (MLIPs) across 6 catalysis-relevant datasets and 13 target properties, using single-point energy evaluations on DFT-relaxed geometries plus a smaller relaxation experiment. The authors report systematic RMSE/MaxAE comparisons, identify catastrophic failures on Co- and Ni-containing surfaces (attributed to magnetic/spin-polarization treatment in training data), and compare zero-shot MLIP accuracy with low-cost task-specific models trained on the same datasets. They conclude that no single MLIP is universally best, that the training dataset strongly influences performance possibly more than architecture, and recommend eSEN, Orb, and UMA as starting points. The paper is transparent about the single-point limitation and provides extensive supporting tables and figures.

Significance. If the benchmark is accepted, it provides a valuable large-scale, systematic assessment of foundational MLIPs for heterogeneous catalysis, filling a gap left by benchmarks limited to ordered bulk crystals. The strength of the paper is its breadth: 80 models, 16 architectures, multiple training datasets, and 13 properties, with results tabulated for all models (Tables S2–S15) so that readers can inspect model- and dataset-level behavior. The paper also ships a high-throughput workflow and commits to releasing code and data, which supports reproducibility. The qualitative finding that training-data composition can dominate architecture is useful guidance for practitioners. However, the validity of the headline accuracy numbers and model rankings depends on the gas-phase error correction in Eq. (9), which is not validated as a pure gas-phase correction and is fitted to the evaluation set; this is a load-bearing issue that must be resolved before the quantitative claims can be taken at face value.

major comments (2)
  1. [§2.4, Eq. (9)] The gas-phase mean-error correction is computed on the same data set that is then used to report RMSE/MaxAE. Subtracting the per-model mean error over the entire evaluation set is an affine fit to the test labels; it always reduces RMSE and can arbitrarily favor models with large systematic biases. The correction is averaged over a mixed set of gas-phase references (H2, H2O, CO2, CO) and is acknowledged in the text to "might not accurately reflect gas-phase prediction error." Because the reported RMSEs in Tables S5–S10/S12 and Figures 2b/2c are corrected values, they are not true zero-shot errors. Since the recommendation of eSEN, Orb, and UMA is based in part on these corrected metrics, removing or properly validating the correction could change the rankings. Please report uncorrected errors alongside corrected ones and validate the gas-phase offset on independent gas-phase molecules.
  2. [§3.4, Figure 5] The main benchmark is restricted to single-point energies on DFT-relaxed geometries, which the paper explicitly states is "not a fully realistic use case." The relaxation experiment on a 10% subset shows that MLIP relaxation increases errors substantially (e.g., eSEN-30M OAM RMSE rises by 1.6×) and that the best single-point models do not reach below ~0.18 eV RMSE after relaxation. This means the abstract's practical-accuracy claim — how accurate MLIPs are for heterogeneous catalysis — is only supported for single-point use on pre-relaxed geometries. The paper's model rankings and the recommendation of starting points should be qualified as applying to single-point evaluation, not to MLIP-relaxed workflows, unless further relaxation data are provided.
minor comments (5)
  1. [§3.1] Typo: "possibly moreso than" should be "possibly more so than."
  2. [Figure 2b caption] The white circle representing scaling relations is a training error, not a test error; this is stated in §3.5 but should appear in the caption or legend for clarity.
  3. [§2.4] The text says errors are corrected "except when indicated otherwise," but it is not always clear which tables and figures use the correction. Please state explicitly for each property whether Eq. (9) was applied.
  4. [Table S1] The footnotes (a–h) are not defined in the table caption. Add a footnote legend explaining the markings for model variants and training data splits.
  5. [§4] The statement "over 700,000 potential energy evaluations" would benefit from a brief derivation (number of models × number of structures × number of properties) to make the scale transparent.

Circularity Check

1 steps flagged · score 6.0 of 10

Test-set gas-phase offset (Eq. 9) is a per-model fit; headline RMSEs for formation/adsorption energies are post-fit residuals.

  1. fitted input called prediction [Section 2.4, Eq. (9); used in Section 3.1 and Section S2 (Tables S5–S10, S12; Figs. 2b, 2c)]
    "For a more fair comparison of MLIP accuracy, we present errors for the prediction of these properties after correcting the MLIP property predictions by the mean error across the entire set of predicted properties ... Ecorr_MLIP,i = EMLIP,i − 1/N Σ_j (EMLIP,j − EDFT,j) ... Note that the correction value is calculated over the entire data set, not split by molecule ... the error correction might not accurately reflect gas-phase prediction error."

    The correction subtracts, for each MLIP, the mean error over the same evaluation set used to report RMSE/MaxAE. This is a per-model constant fitted to the target labels, so the reported RMSE is the centered residual sqrt(RMSE^2 − μ^2), which is always ≤ the raw zero-shot RMSE and can change model rankings. The headline zero-shot claims and the recommendation of eSEN/Orb/UMA as best-performing rest substantially on these corrected formation and adsorption energies (Figs. 2b–c; Tables S5–S10/S12). The paper's own caveat admits the mean is taken over a mixed dataset and may not isolate gas-phase error, so the gas-phase justification is unvalidated; the metric is nevertheless presented as out-of-the-box performance in Section 3.5.

full rationale

This is primarily an external benchmark, not a derivation: 80 pretrained MLIPs are evaluated zero-shot on catalysis-relevant datasets, and the main qualitative results (training-data dependence, magnetic-element failures, relaxation-induced error growth, and the absence of a universally best model) are genuine empirical findings that do not reduce to the inputs. The one load-bearing circular element is the gas-phase error correction of Eq. (9): for every formation/adsorption energy, a per-model mean error computed on the same test set is subtracted before RMSE/MaxAE are reported, making those headline errors post-fit residuals. The paper is transparent about this procedure and notes that the correction is not split by molecule, but no validation is provided to show the offset is actually a gas-phase error, and the model rankings and the recommended starting points (eSEN, Orb, UMA) depend on the corrected metrics. Self-citations to the authors' prior datasets and task-specific models are used as benchmark/test references rather than to justify the MLIP predictions, so they are not circular. Overall, the central claims have substantial independent content, but one class of headline 'zero-shot' predictions is partly fit to the evaluation labels, giving a partial circularity score of 6.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

This is an empirical benchmark paper; it introduces no new physical entities or fitted force-field parameters beyond the data-driven gas-phase correction. The main load-bearing assumptions are the fidelity of DFT references, the harmonic approximation for ZPE, and the post-hoc interpretation of magnetic-element failures.

free parameters (1)
  • Gas-phase mean-error correction offset (Eq. 9) = varies per model and dataset; not tabulated
    For formation and adsorption energies, the mean error over the entire evaluation set is subtracted from each MLIP prediction before computing RMSE. This effectively fits one parameter per model to the same data being evaluated, reducing reported errors.
assumptions (4)
  • domain assumption DFT reference energies (PBE, PBE+U, BEEF-vdW, PBED4) are treated as ground truth for all error calculations
    All RMSE/MaxAE values measure deviation from DFT labels, not experimental truth. The paper itself notes in Section 4 that DFT functionals have intrinsic errors of 0.2–0.3 eV relative to experiments.
  • domain assumption Harmonic approximation for zero-point energies (Eq. 7)
    Zero-point energies are computed from finite-difference Hessians in the harmonic limit; anharmonic effects are neglected. This is a standard approximation for the nanocluster systems studied.
  • ad hoc to paper The catastrophic failure on Co/Ni surfaces is caused by magnetic/spin-polarization treatment in training data
    The paper states 'We hypothesize that this is caused by a combination of these materials being magnetic materials and the utilization of spin-polarization in the preparation of the training data sets.' This is a correlation-based hypothesis, not a direct causal test.
  • domain assumption The 10% random sample used for MLIP relaxation analysis is representative of the full formate adsorption dataset
    Section 3.4 states a random 10% was sampled; the paper does not provide error bars or repeated sampling to confirm stability of the relaxation comparison.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?." pith.science (2026). https://pith.science/paper/2IPH4NE3

@misc{pith2026251216702,
  author       = {Pith},
  title        = {Pith review of: How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2IPH4NE3}},
  note         = {Machine review of arXiv:2512.16702}
}
read the original abstract

Foundational machine learning interatomic potentials (MLIPs) are being developed at a rapid pace, promising closer and closer approximation to ab initio accuracy. This unlocks the possibility to simulate much larger length and time scales. However, benchmarks for these MLIPs are usually limited to ordered, crystalline and bulk materials. Hence, reported performance does not necessarily accurately reflect MLIP performance in real applications such as heterogeneous catalysis. Here, we systematically analyze zero-shot performance of 80 different MLIPs, evaluating tasks typical for heterogeneous catalysis across a range of different data sets, including adsorption and reaction on surfaces of alloyed metals, oxides, and metal-oxide interfacial systems. We demonstrate that current-generation foundational MLIPs can already perform at high accuracy for applications such as predicting vacancy formation energies of perovskite oxides or zero-point energies of supported nanoclusters. However, limitations also exist. We find that many MLIPs catastrophically fail when applied to magnetic materials, and structure relaxation in the MLIP generally increases the energy prediction error compared to single-point evaluation of a previously optimized structure. Comparing low-cost task-specific models to foundational MLIPs, we highlight some core differences between these model approaches and show that -- if considering only accuracy -- these models can compete with the current generation of best-performing MLIPs. Furthermore, we show that no single MLIP universally performs best, requiring users to investigate MLIP suitability for their desired application.

Figures

Figures reproduced from arXiv: 2512.16702 by the authors.

Figure 1
Figure 1. An overview of the data sets and foundational MLIPs assessed in this work. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Prediction performance of the MLIPs for three different data sets and target properties: (a) slab vacancy [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Comparison between RMSEs of formation energy predictions for TSs on metals and SAAs for the entire [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Prediction performance of the MLIPs for the zero-point energy of metal oxide nanoclusters on metal [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Comparison between RMSEs of adsorption energy predictions for formate on metal-supported oxide [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data-driven atomistic modelling of hybrid halide perovskite passivation

    cond-mat.mtrl-sci 2026-07 accept novelty 5.0 of 10

    A continual fine-tuning protocol for machine-learned interatomic potentials enables large-scale simulation of amino-silane passivation at hybrid perovskite surfaces, revealing coverage-dependent lattice disruption.

Reference graph

Works this paper leans on

5 extracted references · 4 canonical work pages · cited by 1 Pith paper

  1. [1]

    Cheula, R., Tran, T. A. M. Q. & Andersen, M. Unraveling the Effect of Dopants in Zirconia-Based Catalysts for CO2 Hydrogenation to Methanol.ACS Catalysis 14, 13126–13135. issn: 2155-5435. http://dx.doi. org/10.1021/acscatal.4c03206 (2024)

  2. [2]

    & Andersen, M

    Cheula, R. & Andersen, M. Transition States Energies from Machine Learning: An Application to Reverse Water–Gas Shift on Single-Atom Alloys.ACS Catalysis 15, 11377–11388. issn: 2155-5435. http://dx. doi.org/10.1021/acscatal.5c02818 (2025)

  3. [3]

    & Bruix, A

    Reichenbach, T., Walter, M., Moseler, M., Hammer, B. & Bruix, A. Effects of Gas-Phase Conditions and Particle Size on the Properties of Cu(111)-Supported ZnyOx Particles Revealed by Global Optimization and Ab Initio Thermodynamics.The Journal of Physical Chemistry C123, 30903–30916. issn: 1932-7455. http://dx.doi.org/10.1021/acs.jpcc.9b07715 (2019)

  4. [4]

    Kempen, L. H. E. & Andersen, M. Inverse catalysts: tuning the composition and structure of oxide clusters through the metal support.npj Computational Materials11, 8. issn: 2057-3960. http://dx.doi.org/10. 1038/s41524-024-01507-z (2025)

  5. [5]

    J., Kempen, L

    Nielsen, M. J., Kempen, L. H. E., de Neergaard Ravn, J., Cheula, R. & Andersen, M. Interpretable machine learned predictions of adsorption energies at the metal–oxide interface.The Journal of Chemical Physics 163, 044708. issn: 1089-7690. http://dx.doi.org/10.1063/5.0282674 (2025). S37

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.