REVIEW 2 major objections 5 minor 1 cited by
How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?
T0 review · 2 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Benchmarking 80 AI interatomic potentials finds no universal winner for heterogeneous catalysis, with training data often mattering more than architecture and many models failing catastrophically on magnetic surfaces.
desk verdict Large, careful MLIP benchmark for catalysis with one real flaw: the gas-phase offset correction fits the test set, so corrected formation/adsorption rankings are not zero-shot; the main qualitative conclusions still hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a benchmark database: over 700,000 single-point potential-energy evaluations from 80 pretrained MLIPs spanning 16 architecture families, scored against DFT reference values on 6 catalysis-relevant datasets and 13 target properties. The argument is carried by comparing root-mean-square and maximum absolute errors across models, grouped by architecture, training data, and model size, and by isolating dataset-driven trends such as the magnetic-element failure.
What would settle it
Re-run the same evaluation but let each MLIP relax every structure to its own local minimum (rather than only a 10% subset of one dataset) and recalculate RMSEs and rankings; if rankings change substantially or the best models no longer remain on top, the paper's recommendation of which models to start with would need to be revised.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a systematic mapping of zero-shot MLIP accuracy across heterogeneous catalysis tasks. Evaluated on DFT-relaxed geometries, the best models (the eSEN, Orb, and UMA families, trained on datasets containing oxides or mixed inorganic materials) reach RMSEs around 0.1–0.3 eV for activation and formation energies, and as low as 0.01 eV for zero-point vibrational energies of supported nanoclusters. Task-specific models trained directly on the evaluation data can match or beat these MLIPs on accuracy, though they are far less general. A significant fraction of MLIPs exhibit catastrophic errors, tens of eV, on surfaces containing cobalt or nickel, and the pa
Load-bearing premise
The headline accuracy numbers assume the user starts from DFT-relaxed geometries, not from structures relaxed by the MLIP itself, and the paper itself notes this is not a fully realistic use case.
Editorial extensions
If this is right
- No single MLIP can be chosen a priori for a catalysis problem; users must screen several models for their specific application, and the paper recommends starting with the eSEN, Orb, and UMA families.
- The strong training-data dependence implies that foundation-model accuracy is set largely by dataset coverage: models trained on data including oxides and spin-polarized systems generalize better to oxide and magnetic surfaces, respectively.
- Current MLIPs are adequate for initial screening (narrowing candidate materials and mechanisms) but not yet accurate enough to replace DFT in microkinetic modeling, since best-case errors are comparable to intrinsic DFT errors.
- Relaxing structures with the MLIP rather than evaluating DFT-relaxed geometries raises errors for all models, so reported RMSEs on fixed geometries are optimistic for realistic workflows.
- Cheap task-specific models (scaling relations, graph-based Gaussian processes, and similar) can compete with the best foundational MLIPs on accuracy for their narrow target, at lower cost and with interpretability.
Reading between the lines
- If training data dominates over architecture, then future gains may come as much from broadening and re-weighting training datasets, especially including spin-polarized, magnetic, and disordered surface configurations, as from architectural innovations.
- The relaxation penalty observed here suggests the MLIP error surface is shifted relative to DFT at off-equilibrium geometries; this points to fine-tuning or delta-learning on the target system as a promising low-cost route to make zero-shot MLIPs reliable for genuine structure searches.
- The catastrophic magnetic-element failures imply that any MLIP claiming to be universal should be stress-tested on magnetic surfaces before being used for transition-metal catalysis; a simple test on cobalt- or nickel-containing surfaces would be a useful addition to standard benchmarks.
- Since the benchmark uses DFT-relaxed geometries from different functionals and numerical settings than the training data, small shifts in ranking could occur if the reference calculations were redone with a consistent functional, suggesting a follow-up test with a unified reference.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper benchmarks 80 pretrained machine-learning interatomic potentials (MLIPs) across 6 catalysis-relevant datasets and 13 target properties, using single-point energy evaluations on DFT-relaxed geometries plus a smaller relaxation experiment. The authors report systematic RMSE/MaxAE comparisons, identify catastrophic failures on Co- and Ni-containing surfaces (attributed to magnetic/spin-polarization treatment in training data), and compare zero-shot MLIP accuracy with low-cost task-specific models trained on the same datasets. They conclude that no single MLIP is universally best, that the training dataset strongly influences performance possibly more than architecture, and recommend eSEN, Orb, and UMA as starting points. The paper is transparent about the single-point limitation and provides extensive supporting tables and figures.
Significance. If the benchmark is accepted, it provides a valuable large-scale, systematic assessment of foundational MLIPs for heterogeneous catalysis, filling a gap left by benchmarks limited to ordered bulk crystals. The strength of the paper is its breadth: 80 models, 16 architectures, multiple training datasets, and 13 properties, with results tabulated for all models (Tables S2–S15) so that readers can inspect model- and dataset-level behavior. The paper also ships a high-throughput workflow and commits to releasing code and data, which supports reproducibility. The qualitative finding that training-data composition can dominate architecture is useful guidance for practitioners. However, the validity of the headline accuracy numbers and model rankings depends on the gas-phase error correction in Eq. (9), which is not validated as a pure gas-phase correction and is fitted to the evaluation set; this is a load-bearing issue that must be resolved before the quantitative claims can be taken at face value.
major comments (2)
- [§2.4, Eq. (9)] The gas-phase mean-error correction is computed on the same data set that is then used to report RMSE/MaxAE. Subtracting the per-model mean error over the entire evaluation set is an affine fit to the test labels; it always reduces RMSE and can arbitrarily favor models with large systematic biases. The correction is averaged over a mixed set of gas-phase references (H2, H2O, CO2, CO) and is acknowledged in the text to "might not accurately reflect gas-phase prediction error." Because the reported RMSEs in Tables S5–S10/S12 and Figures 2b/2c are corrected values, they are not true zero-shot errors. Since the recommendation of eSEN, Orb, and UMA is based in part on these corrected metrics, removing or properly validating the correction could change the rankings. Please report uncorrected errors alongside corrected ones and validate the gas-phase offset on independent gas-phase molecules.
- [§3.4, Figure 5] The main benchmark is restricted to single-point energies on DFT-relaxed geometries, which the paper explicitly states is "not a fully realistic use case." The relaxation experiment on a 10% subset shows that MLIP relaxation increases errors substantially (e.g., eSEN-30M OAM RMSE rises by 1.6×) and that the best single-point models do not reach below ~0.18 eV RMSE after relaxation. This means the abstract's practical-accuracy claim — how accurate MLIPs are for heterogeneous catalysis — is only supported for single-point use on pre-relaxed geometries. The paper's model rankings and the recommendation of starting points should be qualified as applying to single-point evaluation, not to MLIP-relaxed workflows, unless further relaxation data are provided.
minor comments (5)
- [§3.1] Typo: "possibly moreso than" should be "possibly more so than."
- [Figure 2b caption] The white circle representing scaling relations is a training error, not a test error; this is stated in §3.5 but should appear in the caption or legend for clarity.
- [§2.4] The text says errors are corrected "except when indicated otherwise," but it is not always clear which tables and figures use the correction. Please state explicitly for each property whether Eq. (9) was applied.
- [Table S1] The footnotes (a–h) are not defined in the table caption. Add a footnote legend explaining the markings for model variants and training data splits.
- [§4] The statement "over 700,000 potential energy evaluations" would benefit from a brief derivation (number of models × number of structures × number of properties) to make the scale transparent.
Circularity Check
Test-set gas-phase offset (Eq. 9) is a per-model fit; headline RMSEs for formation/adsorption energies are post-fit residuals.
-
fitted input called prediction
[Section 2.4, Eq. (9); used in Section 3.1 and Section S2 (Tables S5–S10, S12; Figs. 2b, 2c)]
"For a more fair comparison of MLIP accuracy, we present errors for the prediction of these properties after correcting the MLIP property predictions by the mean error across the entire set of predicted properties ... Ecorr_MLIP,i = EMLIP,i − 1/N Σ_j (EMLIP,j − EDFT,j) ... Note that the correction value is calculated over the entire data set, not split by molecule ... the error correction might not accurately reflect gas-phase prediction error."
The correction subtracts, for each MLIP, the mean error over the same evaluation set used to report RMSE/MaxAE. This is a per-model constant fitted to the target labels, so the reported RMSE is the centered residual sqrt(RMSE^2 − μ^2), which is always ≤ the raw zero-shot RMSE and can change model rankings. The headline zero-shot claims and the recommendation of eSEN/Orb/UMA as best-performing rest substantially on these corrected formation and adsorption energies (Figs. 2b–c; Tables S5–S10/S12). The paper's own caveat admits the mean is taken over a mixed dataset and may not isolate gas-phase error, so the gas-phase justification is unvalidated; the metric is nevertheless presented as out-of-the-box performance in Section 3.5.
full rationale
This is primarily an external benchmark, not a derivation: 80 pretrained MLIPs are evaluated zero-shot on catalysis-relevant datasets, and the main qualitative results (training-data dependence, magnetic-element failures, relaxation-induced error growth, and the absence of a universally best model) are genuine empirical findings that do not reduce to the inputs. The one load-bearing circular element is the gas-phase error correction of Eq. (9): for every formation/adsorption energy, a per-model mean error computed on the same test set is subtracted before RMSE/MaxAE are reported, making those headline errors post-fit residuals. The paper is transparent about this procedure and notes that the correction is not split by molecule, but no validation is provided to show the offset is actually a gas-phase error, and the model rankings and the recommended starting points (eSEN, Orb, UMA) depend on the corrected metrics. Self-citations to the authors' prior datasets and task-specific models are used as benchmark/test references rather than to justify the MLIP predictions, so they are not circular. Overall, the central claims have substantial independent content, but one class of headline 'zero-shot' predictions is partly fit to the evaluation labels, giving a partial circularity score of 6.
Assumptions & free parameters
free parameters (1)
- Gas-phase mean-error correction offset (Eq. 9) =
varies per model and dataset; not tabulated
assumptions (4)
- domain assumption DFT reference energies (PBE, PBE+U, BEEF-vdW, PBED4) are treated as ground truth for all error calculations
- domain assumption Harmonic approximation for zero-point energies (Eq. 7)
- ad hoc to paper The catastrophic failure on Co/Ni surfaces is caused by magnetic/spin-polarization treatment in training data
- domain assumption The 10% random sample used for MLIP relaxation analysis is representative of the full formate adsorption dataset
Cite this review
Pith. "Pith review of How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?." pith.science (2026). https://pith.science/paper/2IPH4NE3
@misc{pith2026251216702,
author = {Pith},
title = {Pith review of: How accurate are foundational machine learning interatomic potentials for heterogeneous catalysis?},
year = {2026},
howpublished = {\url{https://pith.science/paper/2IPH4NE3}},
note = {Machine review of arXiv:2512.16702}
}
read the original abstract
Foundational machine learning interatomic potentials (MLIPs) are being developed at a rapid pace, promising closer and closer approximation to ab initio accuracy. This unlocks the possibility to simulate much larger length and time scales. However, benchmarks for these MLIPs are usually limited to ordered, crystalline and bulk materials. Hence, reported performance does not necessarily accurately reflect MLIP performance in real applications such as heterogeneous catalysis. Here, we systematically analyze zero-shot performance of 80 different MLIPs, evaluating tasks typical for heterogeneous catalysis across a range of different data sets, including adsorption and reaction on surfaces of alloyed metals, oxides, and metal-oxide interfacial systems. We demonstrate that current-generation foundational MLIPs can already perform at high accuracy for applications such as predicting vacancy formation energies of perovskite oxides or zero-point energies of supported nanoclusters. However, limitations also exist. We find that many MLIPs catastrophically fail when applied to magnetic materials, and structure relaxation in the MLIP generally increases the energy prediction error compared to single-point evaluation of a previously optimized structure. Comparing low-cost task-specific models to foundational MLIPs, we highlight some core differences between these model approaches and show that -- if considering only accuracy -- these models can compete with the current generation of best-performing MLIPs. Furthermore, we show that no single MLIP universally performs best, requiring users to investigate MLIP suitability for their desired application.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Data-driven atomistic modelling of hybrid halide perovskite passivation
A continual fine-tuning protocol for machine-learned interatomic potentials enables large-scale simulation of amino-silane passivation at hybrid perovskite surfaces, revealing coverage-dependent lattice disruption.
Reference graph
Works this paper leans on
-
[1]
Cheula, R., Tran, T. A. M. Q. & Andersen, M. Unraveling the Effect of Dopants in Zirconia-Based Catalysts for CO2 Hydrogenation to Methanol.ACS Catalysis 14, 13126–13135. issn: 2155-5435. http://dx.doi. org/10.1021/acscatal.4c03206 (2024)
-
[2]
Cheula, R. & Andersen, M. Transition States Energies from Machine Learning: An Application to Reverse Water–Gas Shift on Single-Atom Alloys.ACS Catalysis 15, 11377–11388. issn: 2155-5435. http://dx. doi.org/10.1021/acscatal.5c02818 (2025)
-
[3]
Reichenbach, T., Walter, M., Moseler, M., Hammer, B. & Bruix, A. Effects of Gas-Phase Conditions and Particle Size on the Properties of Cu(111)-Supported ZnyOx Particles Revealed by Global Optimization and Ab Initio Thermodynamics.The Journal of Physical Chemistry C123, 30903–30916. issn: 1932-7455. http://dx.doi.org/10.1021/acs.jpcc.9b07715 (2019)
-
[4]
Kempen, L. H. E. & Andersen, M. Inverse catalysts: tuning the composition and structure of oxide clusters through the metal support.npj Computational Materials11, 8. issn: 2057-3960. http://dx.doi.org/10. 1038/s41524-024-01507-z (2025)
-
[5]
Nielsen, M. J., Kempen, L. H. E., de Neergaard Ravn, J., Cheula, R. & Andersen, M. Interpretable machine learned predictions of adsorption energies at the metal–oxide interface.The Journal of Chemical Physics 163, 044708. issn: 1089-7690. http://dx.doi.org/10.1063/5.0282674 (2025). S37
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.