{"id":"59c1e224-f3ab-4718-9143-2fb599ea4a7c","arxiv_id":"2512.16702","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A zero-shot benchmark of 80 MLIPs across six catalysis datasets shows no universally best model, strong training-data dependence, and catastrophic errors on magnetic Co/Ni surfaces.","lead":"This paper tested 80 off-the-shelf machine learning interatomic potentials on catalysis tasks—adsorption, reaction barriers, vacancy formation, and vibrational energies—against DFT reference data. It finds that no single model wins everywhere, training data choice matters as much as architecture, and many models fail badly on cobalt- and nickel-containing surfaces.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gas-phase mean-error correction (Eq. 9) is fitted to the test set and may dominate reported rankings; removal could change the recommended models.","rationale":"The reader's weakest assumption was the single-point DFT-relaxed geometry limitation. That is a legitimate concern, and the paper's own relaxation experiment (Section 3.4) shows errors increase and rankings may shift. However, that limitation is explicitly acknowledged and partially tested. The gas-phase error correction is a more subtle but potentially more damaging issue: it is fitted to the same test data and affects the reported numbers for all formation/adsorption energies, which are central to the 'best-performing models' claim. The paper itself flags the assumption in Section 2.4, but does not provide uncorrected values or a sensitivity analysis. The correction could artificially boost any model with a large systematic offset, and no independent validation shows it isolates gas-phase error. The reader mentioned this as a secondary reason for CONDITIONAL, but did not treat it as the weakest assumption. I agree with the CONDITIONAL verdict because the concern is addressable and may not be fatal to the qualitative 'no universally best' claim, but it could affect the specific model recommendations. A concrete recomputation would settle whether the rankings and recommendations are robust.","tokens_in":59493,"tokens_out":5426,"duration_ms":61253,"concrete_test":"Recompute all RMSE/MaxAE values in Tables S5–S10 and S12 without the Eq. (9) correction, or with a correction based only on MLIP errors for the isolated gas-phase reference molecules (H2, H2O, CO2, CO) as defined in Section 2.4. Then compare the per-property top-5 model ranking and the stated recommendation of eSEN/Orb/UMA against the published values. If the top-5 changes by more than two models for any property, or if eSEN/Orb/UMA no longer appear in the top tier across most properties, the central recommendation is not robust to the correction.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline zero-shot accuracy numbers and model rankings rest on the gas-phase error correction defined in Eq. (9), applied to all formation and adsorption energies. For each MLIP, this subtracts the mean error over the entire target dataset from every prediction. Because the correction is computed on the same test set used to report RMSE/MaxAE, it is a per-model affine fit to the test data. It always reduces RMSE (RMSE_centered = sqrt(RMSE^2 - mean_error^2)) and can arbitrarily improve models with large systematic biases. The paper assumes this mean error 'mostly comprises' gas-phase errors, but the mean is taken over a mixed set of adsorbates and TSs with different gas-phase references (H2, H2O, CO2, CO), and it also absorbs any solid-state offset. The authors explicitly note the correction is not split by molecule and 'might not accurately reflect gas-phase prediction error' (Section 2.4). No validation isolates the gas-phase contribution. Consequently, the RMSE/MaxAE values in Tables S5–S10/S12 and Figures 2b/2c are not true zero-shot errors. The central claim that eSEN, Orb, and UMA are among the best-performing models is based on these corrected metrics, so a hidden test-set fit could inflate their apparent accuracy. If the correction is removed or computed solely from gas-phase references, the rankings and practical recommendations may shift. This is load-bearing because the paper's quantitative evidence for its main qualitative conclusions depends on this unvalidated correction.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks 80 pretrained machine-learning interatomic potentials (MLIPs) across 6 catalysis-relevant datasets and 13 target properties, using single-point energy evaluations on DFT-relaxed geometries plus a smaller relaxation experiment. The authors report systematic RMSE/MaxAE comparisons, identify catastrophic failures on Co- and Ni-containing surfaces (attributed to magnetic/spin-polarization treatment in training data), and compare zero-shot MLIP accuracy with low-cost task-specific models trained on the same datasets. They conclude that no single MLIP is universally best, that the training dataset strongly influences performance possibly more than architecture, and recommend eSEN, Orb, and UMA as starting points. The paper is transparent about the single-point limitation and provides extensive supporting tables and figures.","tokens_in":59867,"tokens_out":4153,"duration_ms":41232,"significance":"If the benchmark is accepted, it provides a valuable large-scale, systematic assessment of foundational MLIPs for heterogeneous catalysis, filling a gap left by benchmarks limited to ordered bulk crystals. The strength of the paper is its breadth: 80 models, 16 architectures, multiple training datasets, and 13 properties, with results tabulated for all models (Tables S2–S15) so that readers can inspect model- and dataset-level behavior. The paper also ships a high-throughput workflow and commits to releasing code and data, which supports reproducibility. The qualitative finding that training-data composition can dominate architecture is useful guidance for practitioners. However, the validity of the headline accuracy numbers and model rankings depends on the gas-phase error correction in Eq. (9), which is not validated as a pure gas-phase correction and is fitted to the evaluation set; this is a load-bearing issue that must be resolved before the quantitative claims can be taken at face value.","major_comments":[{"comment":"The gas-phase mean-error correction is computed on the same data set that is then used to report RMSE/MaxAE. Subtracting the per-model mean error over the entire evaluation set is an affine fit to the test labels; it always reduces RMSE and can arbitrarily favor models with large systematic biases. The correction is averaged over a mixed set of gas-phase references (H2, H2O, CO2, CO) and is acknowledged in the text to \"might not accurately reflect gas-phase prediction error.\" Because the reported RMSEs in Tables S5–S10/S12 and Figures 2b/2c are corrected values, they are not true zero-shot errors. Since the recommendation of eSEN, Orb, and UMA is based in part on these corrected metrics, removing or properly validating the correction could change the rankings. Please report uncorrected errors alongside corrected ones and validate the gas-phase offset on independent gas-phase molecules.","section":"§2.4, Eq. (9)"},{"comment":"The main benchmark is restricted to single-point energies on DFT-relaxed geometries, which the paper explicitly states is \"not a fully realistic use case.\" The relaxation experiment on a 10% subset shows that MLIP relaxation increases errors substantially (e.g., eSEN-30M OAM RMSE rises by 1.6×) and that the best single-point models do not reach below ~0.18 eV RMSE after relaxation. This means the abstract's practical-accuracy claim — how accurate MLIPs are for heterogeneous catalysis — is only supported for single-point use on pre-relaxed geometries. The paper's model rankings and the recommendation of starting points should be qualified as applying to single-point evaluation, not to MLIP-relaxed workflows, unless further relaxation data are provided.","section":"§3.4, Figure 5"}],"minor_comments":[{"comment":"Typo: \"possibly moreso than\" should be \"possibly more so than.\"","section":"§3.1"},{"comment":"The white circle representing scaling relations is a training error, not a test error; this is stated in §3.5 but should appear in the caption or legend for clarity.","section":"Figure 2b caption"},{"comment":"The text says errors are corrected \"except when indicated otherwise,\" but it is not always clear which tables and figures use the correction. Please state explicitly for each property whether Eq. (9) was applied.","section":"§2.4"},{"comment":"The footnotes (a–h) are not defined in the table caption. Add a footnote legend explaining the markings for model variants and training data splits.","section":"Table S1"},{"comment":"The statement \"over 700,000 potential energy evaluations\" would benefit from a brief derivation (number of models × number of structures × number of properties) to make the scale transparent.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful and timely benchmark, but the Eq. (9) correction is a test-set fit that could materially affect the reported rankings and the central recommendation. The authors should be asked to provide uncorrected results and a proper gas-phase validation. The relaxation caveat is acknowledged, but the abstract and discussion could overstate practical accuracy; this should be tempered. Overall, the paper is within scope and likely to be a valuable contribution once the correction issue is resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a genuinely useful, large-scale zero-shot benchmark of 80 MLIPs on catalysis-specific datasets and properties — the kind of systematic comparison the field has been missing. It does not introduce new methods, but it gives practical guidance on model selection. The headline claims — no universal best MLIP, training data often matters more than architecture, many models fail badly on magnetic surfaces, and relaxation degrades accuracy — are supported by the evidence.\n\nWhat's new is the breadth. Six datasets, 13 properties, from perovskite vacancies to zero-point energies to transition-state formation energies, all evaluated consistently. The paper is careful about workflows: DFT-relaxed geometries for the main screening, a smaller relaxation experiment, and a fair comparison to task-specific models. The tables are extensive and internally consistent. The magnetic failure analysis is especially convincing: errors exceeding 15 eV on Co/Ni surfaces, with clear training-data dependence, and the CHGNet/AQCat spin-aware models as a nice control.\n\nWhere I'd push back: the gas-phase mean-error correction (Eq. 9). It subtracts the mean error over the entire test set from each property before computing RMSE. That's a per-model fit to the data being evaluated. The authors justify it as compensating for gas-phase errors, but the mean is taken over a mixed set of molecules and also absorbs solid-state offsets. They acknowledge it isn't split by molecule. The practical effect is that corrected RMSE and MaxAE values for formation and adsorption energies are not true zero-shot errors. Tables S5–S10, S12, and Figures 2b/2c rely on this correction, and if it were removed those specific rankings could shift. The stress-test note is right to flag this as load-bearing for those numbers.\n\nThat said, I don't think it sinks the paper. Perovskite vacancy formation energies, activation/reaction energies, and zero-point energies are uncorrected and show the same trends: eSEN, Orb, and UMA are consistently strong, and training-data effects dominate. The magnetic failure analysis is shown uncorrected in the SI. The authors are transparent about the correction's limitations. The fix is straightforward: report uncorrected values alongside, or validate the gas-phase error on a separate set.\n\nOne more thing: data and code are promised only \"upon publication,\" which limits reproducibility now. For a benchmark paper, earlier release would help.\n\nOverall, this deserves serious peer review. A referee should ask for the uncorrected numbers and a clearer separation of gas-phase and solid-state errors. It belongs in the conversation — I'd cite it as the current state of the art on MLIP zero-shot performance for catalysis.\n\nRecommendation: send it out, with requests for major revision on the error-correction transparency and data availability.","headline":"Large, careful MLIP benchmark for catalysis with one real flaw: the gas-phase offset correction fits the test set, so corrected formation/adsorption rankings are not zero-shot; the main qualitative conclusions still hold.","tokens_in":60287,"tokens_out":3138,"would_cite":true,"duration_ms":32995,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Benchmarking 80 AI interatomic potentials finds no universal winner for heterogeneous catalysis, with training data often mattering more than architecture and many models failing catastrophically on magnetic surfaces.","keywords":["machine learning interatomic potentials","heterogeneous catalysis","zero-shot benchmark","foundation models","adsorption energies","transition states","magnetic materials","structure relaxation"],"falsifier":"Re-run the same evaluation but let each MLIP relax every structure to its own local minimum (rather than only a 10% subset of one dataset) and recalculate RMSEs and rankings; if rankings change substantially or the best models no longer remain on top, the paper's recommendation of which models to start with would need to be revised.","tokens_in":59398,"feed_emoji":"🧪","tokens_out":4162,"duration_ms":41301,"temperature":0.7,"pith_summary":"Foundational machine learning interatomic potentials (MLIPs) promise near-DFT accuracy at low cost, but their published benchmarks usually cover ordered bulk crystals, not the surfaces, adsorbates, and transition states that matter for catalysis. This paper systematically tests 80 pretrained MLIPs zero-shot, without fine-tuning, on 13 catalysis-relevant properties across six datasets, including oxide surfaces, alloyed metal surfaces, and metal–oxide interfaces. The central finding is that no single MLIP wins everywhere: performance depends strongly on which dataset the model was trained on, possibly more than on its architecture, so a model that excels at oxide vacancy energies can fail on magnetic alloy surfaces. The paper shows that current-generation models are accurate enough for screening, with best-case errors around 0.1–0.3 eV, but not yet negligible relative to DFT's intrinsic errors, and that relaxing structures within the MLIP typically increases prediction error.","feed_headline":"Benchmarking 80 AI potentials finds no universal winner for catalysis","feed_subtitle":"Zero-shot tests show training data often beats architecture; many models fail on magnetic surfaces.","key_machinery":"The central object is a benchmark database: over 700,000 single-point potential-energy evaluations from 80 pretrained MLIPs spanning 16 architecture families, scored against DFT reference values on 6 catalysis-relevant datasets and 13 target properties. The argument is carried by comparing root-mean-square and maximum absolute errors across models, grouped by architecture, training data, and model size, and by isolating dataset-driven trends such as the magnetic-element failure.","core_discovery":"On the paper's own terms, the discovery is a systematic mapping of zero-shot MLIP accuracy across heterogeneous catalysis tasks. Evaluated on DFT-relaxed geometries, the best models (the eSEN, Orb, and UMA families, trained on datasets containing oxides or mixed inorganic materials) reach RMSEs around 0.1–0.3 eV for activation and formation energies, and as low as 0.01 eV for zero-point vibrational energies of supported nanoclusters. Task-specific models trained directly on the evaluation data can match or beat these MLIPs on accuracy, though they are far less general. A significant fraction of MLIPs exhibit catastrophic errors, tens of eV, on surfaces containing cobalt or nickel, and the pa","pith_inferences":["If training data dominates over architecture, then future gains may come as much from broadening and re-weighting training datasets, especially including spin-polarized, magnetic, and disordered surface configurations, as from architectural innovations.","The relaxation penalty observed here suggests the MLIP error surface is shifted relative to DFT at off-equilibrium geometries; this points to fine-tuning or delta-learning on the target system as a promising low-cost route to make zero-shot MLIPs reliable for genuine structure searches.","The catastrophic magnetic-element failures imply that any MLIP claiming to be universal should be stress-tested on magnetic surfaces before being used for transition-metal catalysis; a simple test on cobalt- or nickel-containing surfaces would be a useful addition to standard benchmarks.","Since the benchmark uses DFT-relaxed geometries from different functionals and numerical settings than the training data, small shifts in ranking could occur if the reference calculations were redone with a consistent functional, suggesting a follow-up test with a unified reference."],"forward_implications":["No single MLIP can be chosen a priori for a catalysis problem; users must screen several models for their specific application, and the paper recommends starting with the eSEN, Orb, and UMA families.","The strong training-data dependence implies that foundation-model accuracy is set largely by dataset coverage: models trained on data including oxides and spin-polarized systems generalize better to oxide and magnetic surfaces, respectively.","Current MLIPs are adequate for initial screening (narrowing candidate materials and mechanisms) but not yet accurate enough to replace DFT in microkinetic modeling, since best-case errors are comparable to intrinsic DFT errors.","Relaxing structures with the MLIP rather than evaluating DFT-relaxed geometries raises errors for all models, so reported RMSEs on fixed geometries are optimistic for realistic workflows.","Cheap task-specific models (scaling relations, graph-based Gaussian processes, and similar) can compete with the best foundational MLIPs on accuracy for their narrow target, at lower cost and with interpretability."],"fun_headline_variants":["80 AI potentials: no universal winner for catalysis","Zero-shot AI potentials fail on magnetic catalytic surfaces","Catalysis benchmarks show task-specific models can rival MLIPs","AI interatomic potentials: high accuracy on some, catastrophic on others","Systematic test finds no best AI potential for catalysis tasks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline accuracy numbers assume the user starts from DFT-relaxed geometries, not from structures relaxed by the MLIP itself, and the paper itself notes this is not a fully realistic use case.","fun_headline_variants_meta":{"raw":{"variants":["80 AI potentials: no universal winner for catalysis","Zero-shot AI potentials fail on magnetic catalytic surfaces","Catalysis benchmarks show task-specific models can rival MLIPs","AI interatomic potentials: high accuracy on some, catastrophic on others","Systematic test finds no best AI potential for catalysis tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1099,"prompt_tokens":790,"completion_tokens":309,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":229}},"tokens_in":534,"tokens_out":309,"duration_ms":3671,"temperature":1.0,"reasoning_tokens":229,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T15:25:11.543491+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same evaluation but let each MLIP relax every structure to its own local minimum (rather than only a 10% subset of one dataset) and recalculate RMSEs and rankings; if rankings change substantially or the best models no longer remain on top, the paper's recommendation of which models to start with would need to be revised.","supporting_citations":[],"review_version":1}