{"id":"8816c013-2e29-461b-9068-6902a6c2d07d","arxiv_id":"2507.21297","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Weights from OMat24 force errors turn an eleven-model heterogeneous uMLIP ensemble into an uncertainty metric U that correlates with true force errors across material families and drives low-DFT distillation.","lead":"An uncertainty score for AI atomistic force fields is built by combining eleven pre-trained models, each weighted by its measured accuracy on a large reference set. The score flags unreliable predictions without new DFT calculations and allows training accurate system-specific potentials with a fraction of the usual DFT data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The OMat24 test set is used both to set the inverse-RMSE weights and the ensemble size and to report the headline Spearman ρ=0.87, so the central 'universal' claim rests on an in-sample calibration that is not directly tested for portability of the weighting scheme.","rationale":"The reader's weakest_assumption identifies exactly the concern that OMat24 test labels set the weights and ensemble selection before external validation, so the OMat24 ρ is in-sample and portability of the weighting is asserted, not demonstrated. My reading agrees: the central claim is that U is a universal uncertainty metric, and the most load-bearing part of that claim is that a single weighting derived from OMat24 remains valid for arbitrary target systems. The external datasets do show strong frozen-weight correlations, which is real supporting evidence and prevents a rejection. However, those datasets are evaluated with fixed weights, so they do not test the assumption that the accuracy ranking of the eleven models is stable. The paper's own acknowledgements of degraded behavior for magnetic TM23 elements and for carbon are direct indications that the coupling between U and error is not universal. The concrete test I propose would settle the concern by measuring how much of the reported performance is due to the particular OMat24-derived weighting. Since the reader already conditions acceptance on a held-out OMat24 split and the concern reinforces that request, the verdict should remain CONDITIONAL (UNCHANGED). No additional objection beyond the reader's is needed, but the test sharpens what a held-out split must show.","tokens_in":22147,"tokens_out":8327,"duration_ms":112676,"concrete_test":"Re-derive the weights w_k and the optimal ensemble size on a held-out part of OMat24 (e.g., split by space group or by elemental subset), then evaluate U with these re-derived weights on the held-out part. If the Spearman ρ drops substantially below the reported 0.87, or if the optimal ensemble size and weights change materially, the headline correlation is inflated by selection on the test set and the weighting scheme is not portable. A complementary check: for each external dataset in Fig. 3, recompute w_k from that dataset's own per-model force RMSEs and compare the resulting ρ with the frozen-weight ρ; large discrepancies would indicate that the 'universal' metric depends on the calibration set.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The definition of U in Eq. 1 depends critically on the weights w_k of Eq. 2, which are the inverse force RMSEs of each uMLIP computed on the OMat24 test set. The ensemble size (eleven models) is also selected by maximizing Spearman ρ on that same OMat24 test set in Fig. 2b. Reporting ρ=0.87 on OMat24 is therefore an in-sample statistic: the same labels that fix the metric are used to evaluate it. The external datasets in Fig. 3 test transferability with frozen weights, which is valuable but does not test whether the relative model ranking that determines w_k is stable across chemistries. If a new system changes the accuracy ordering of the ensemble members, the weighted spread is dominated by models whose errors are not representative for that system, and U becomes miscalibrated. The paper itself flags two places where this coupling degrades: Supplementary Note 1 states that for magnetic elements in TM23 (Fe, Co, Nb) the DFT reference errors can exceed U, 'obscuring the relationship between U and true model uncertainty,' and Fig. 3c identifies carbon as an exception where 'uncertainty and error are less tightly coupled.' These are precisely settings where a global OMat24-derived weighting could break down. The claim of universality requires evidence that the weighting itself, not just the frozen-weight metric, transfers to new domains.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes U, a weighted ensemble uncertainty metric for universal machine learning interatomic potentials (uMLIPs), in which eleven pretrained uMLIPs are combined with inverse-force-RMSE weights derived from the OMat24 test set. The authors report Spearman correlations of 0.87 on OMat24 and 0.82-0.92 on external datasets, and use U to define an uncertainty-aware distillation (UAMD) workflow in which high-uncertainty configurations receive DFT labels. They demonstrate that a tungsten ACE potential trained with roughly 4% DFT labels matches a fully DFT-trained model on phonons, stacking fault energies, and bicrystal tensile behavior, and that a MoNbTaW ACE potential can be trained entirely on uMLIP labels with accuracy comparable to a DFT-trained ACE.","tokens_in":22470,"tokens_out":6697,"duration_ms":71903,"significance":"If the correlation of U with true force errors is robust out of sample, the metric would give uMLIP users a practical quality monitor and a data-selection criterion that requires no additional training, while the UAMD results would substantially lower the cost of building accurate system-specific potentials. The paper's strengths include the reuse of a diverse set of pretrained models, the breadth of external validation datasets, the inclusion of physical property tests (phonons, grain-boundary tension, NEB barriers), and the public availability of code and data. The main risk is that the metric's parameters are calibrated on the OMat24 test set, so the headline in-sample correlation does not by itself certify the claim of universality; the paper's own exceptions for magnetic elements and carbon point to conditions under which the weighting could break down.","major_comments":[{"comment":"The metric's parameters—the weights w_k in Eq. (2) and the ensemble size of eleven in Fig. 2b—are determined on the OMat24 test set, and the headline Spearman rho=0.87 is computed on the same set. The external validation in Fig. 3a uses these frozen weights, which demonstrates that the fixed metric transfers, but it does not test whether the relative accuracy ordering of the individual uMLIPs, and hence the weighting, is stable across chemistries. If a new material class changes that ordering, the weighted ensemble can be dominated by models whose errors are unrepresentative, and U would become miscalibrated. Please provide an out-of-sample test of the weighting scheme itself: for example, re-derive w_k and the ensemble size on a training portion of OMat24 and evaluate on a held-out portion and on the external datasets, or report per-model force RMSEs on each external dataset to show that the ranking underlying Eq. (2) is stable. In addition, the abstract's statement that U is obtained without requiring calibration is inconsistent with the explicit calibration of w_k and the ensemble size on OMat24, and should be rephrased.","section":"Results: Universal uncertainty metric U via heterogeneous ensemble, Eqs. (1)-(2), Fig. 2b-c"},{"comment":"The paper itself identifies two exceptions to the strong U-error coupling: Supplementary Note 1 states that for magnetic elements in TM23 (Fe, Co, Nb) the DFT reference errors can exceed U, 'obscuring the relationship between U and true model uncertainty,' and Fig. 3c identifies carbon as a system where 'uncertainty and error are less tightly coupled.' These are precisely the out-of-distribution settings where a global OMat24-derived weighting could break down, and the tail behavior of U is what matters most for flagging catastrophic errors. Please quantify these deviations—the fraction of configurations affected and the magnitude of underestimation—and discuss whether the recommended U_c=1 eV/Å cutoff remains safe in these cases, or whether a materially different operating point is needed.","section":"Validation of U across diverse materials; Supplementary Note 1; Fig. 3c"},{"comment":"The claim that a W potential trained with only 4% DFT labels matches full-DFT training is qualified by Fig. 4d, which shows that elastic constants (C11, C12, C44, bulk modulus, Poisson's ratio) exhibit relative errors approaching 10% at the recommended operating point U_c=1 eV/Å. This is a substantial deviation for mechanical property prediction, and it is not acknowledged in the abstract or conclusions. Please report the actual numeric RMSE or relative-error values at the 4% operating point for all quantities in Fig. 4d, and either refine the claim of 'comparable accuracy' to state the property-dependent accuracy explicitly or evaluate whether a more conservative U_c reduces the elastic-constant errors to an acceptable level.","section":"Uncertainty-aware model distillation for W, Fig. 4d"},{"comment":"It is unclear whether the RMSE values reported in Fig. 5b are computed on the training set (the 17,654 configurations used to fit ACE_UAMD and ACE_DFT) or on a held-out test portion; the text does not describe a train/test split for this dataset. If these are training-set fitting errors, they do not support the claim that UAMD achieves accuracy comparable to ACE_DFT on unseen data. Please specify the evaluation protocol, and if the reported numbers are in-sample, provide cross-validated or held-out errors for the four scenarios in Fig. 5b.","section":"Uncertainty-aware model distillation for MoNbTaW alloys, Fig. 5b"}],"minor_comments":[{"comment":"The caption of Fig. 4 appears to mislabel the panels: the text refers to Fig. 4d for basic properties, Fig. 4e for phonon/stacking-fault comparisons, and Fig. 4f for stress-strain curves, while the caption lists these as c, d-e, and f. Please align the caption with the panel labels used in the text.","section":"Fig. 4 caption and text"},{"comment":"The phrase 'without requiring additional training or calibration' is misleading given that the weights in Eq. (2) and the ensemble size are calibrated on the OMat24 test set; consider replacing it with 'without additional training of uMLIPs' or explicitly describing the lightweight calibration step.","section":"Abstract and Introduction"},{"comment":"The 'Atom lost' error noted in Fig. 5f is not explained; please indicate whether this reflects a numerical instability of the potential, a limitation of the simulation setup, or a failure of the model, because the reader needs to interpret the premature termination of some tensile simulations.","section":"Fig. 5f and Section 'Uncertainty-aware model distillation for MoNbTaW alloys'"},{"comment":"The statement 'complete avoidance of DFT in the expanded MoNbTaW dataset' is accurate only for the UAMD labeling phase; the initial 17,654-configuration dataset from Ref. (36) was itself generated with DFT. Please add a clarifying clause to avoid overinterpretation.","section":"Discussion and Data availability"},{"comment":"The uniform-weight baseline U^(0) is not written explicitly; providing its formula would help readers compare it with U^(1) and U^(2). Also, please clarify in the text that the max over j in Eq. (1) is taken over atoms within configuration i for each model k, and that the average force in Eq. (1) is the unweighted or weighted mean as appropriate.","section":"Eqs. (1)-(3)"}],"recommendation":"major_revision","confidential_remarks":"The central idea is promising and the empirical results are extensive, but the in-sample calibration of the weights and ensemble size on OMat24, combined with the lack of a clear held-out evaluation in the MoNbTaW distillation study, makes the universality claim currently overreaching. If the authors can provide an out-of-sample test of the weighting scheme and clarify the train/test protocol, the paper would be a strong contribution to the UQ-for-foundation-models literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper has a real idea and it mostly works. The inverse-RMSE weighted heterogeneous ensemble of pretrained uMLIPs gives an uncertainty metric that tracks true force errors across a broad sweep of chemistries, and the UAMD distillation recipe cuts DFT labels to about 4% for W and zero for MoNbTaW while keeping the resulting ACE potentials close to fully DFT-trained ones. That is a practical advance, not just a rehash. The paper also does the validation honestly: multiple external datasets, property-level checks (phonons, GB tension, NEB barriers), and clear reporting of where U degrades (magnetic elements in TM23, carbon). I believe the central argument holds up.\n\nThe soft spot is the one the stress-test note flags. The headline rho=0.87 on OMat24 is in-sample: the same test labels select the ensemble weights (Eq. 2), the ensemble size (Fig. 2b), and the U_c threshold (Fig. 3d,f,h). External datasets test frozen-weight transfer, which is valuable, but they do not test whether the model ranking that defines the weights survives a change of domain. If a new chemistry reorders the uMLIP accuracies, the weighted spread is dominated by models whose errors are not representative there, and U becomes miscalibrated. The paper's own exceptions--Fe/Co/Nb in TM23, carbon--are exactly the kind of places where a global OMat24-derived weighting could break. The claim of universality is therefore stronger than the evidence. I would not call this fatal; I would call it a calibration-portability gap that needs a held-out OMat24 split or a small per-domain recalibration study.\n\nMinor issues: no error bars on the Spearman coefficients (they range 0.82-0.92; that spread matters), only one UQ baseline (Orb-confidence) is compared, and the data/parameter files are not yet public. The code repository is promised but the potential files are 'upon publication.' For a paper whose whole pitch is practical adoption, shipping the weights and datasets is part of the contribution, not an afterthought.\n\nWho should read this: anyone building or using uMLIP-based UQ, and anyone doing active-learning distillation of foundation models into lightweight potentials. The UAMD protocol and the denoising observation (smooth uMLIP labels beating noisy DFT labels in some regimes) are worth citing.\n\nI would send this to peer review. The in-sample calibration issue is serious but fixable, and the external validation plus the application results make the paper worth referee time. My own verdict would be conditional acceptance, with the portability of the weighting scheme as the main ask.","headline":"A genuinely useful uncertainty metric for uMLIPs, with a real in-sample calibration caveat that the paper's own external tests only partially address; deserves peer review.","tokens_in":814,"tokens_out":1181,"would_cite":true,"duration_ms":33112,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Eleven pretrained universal interatomic potentials, combined by inverse-RMSE weights, define a universal uncertainty metric $U$ that tracks true force errors and enables nearly DFT-free distillation.","keywords":["uncertainty quantification","universal machine learning interatomic potentials","heterogeneous ensemble","model distillation","atomistic foundation models","force-field uncertainty","high-entropy alloys","density functional theory"],"falsifier":"Take a held-out collection of spin-polarized magnetic defects, molecular crystals, and surfaces with consistent DFT labels, compute $U$ and the true force error per configuration, and check whether the Spearman correlation remains near 0.8 or higher and whether the cutoff $U_c = 1$ eV/\\AA still separates configurations with force RMSE at or below 0.1 eV/\\AA. If either fails, the universality of $U$ and the recommended threshold would need to be abandoned or revised.","tokens_in":21944,"feed_emoji":"⚛️","tokens_out":11270,"duration_ms":118830,"temperature":0.7,"pith_summary":"The paper proposes a universal uncertainty metric, $U$, for universal machine-learning interatomic potentials (uMLIPs), i.e., foundation models that predict energies and forces across the periodic table. $U$ is the inverse-RMSE-weighted spread of force predictions across an eleven-member heterogeneous ensemble of pretrained potentials, and the paper argues that it correlates strongly with true force errors on an absolute scale without retraining or calibration. The central evidence is a Spearman $\\rho = 0.87$ on the OMat24 test set and $\\rho = 0.82$--$0.92$ on diverse external datasets spanning metals, alloys, inorganic compounds, MOFs, perovskites, and battery materials. On this basis the paper builds an uncertainty-aware distillation framework in which low-uncertainty configurations receive cheap teacher labels and high-uncertainty configurations receive targeted DFT labels. The payoff is that a tungsten ACE potential trained with about 4% DFT labels matches full-DFT training, while a MoNbTaW high-entropy-alloy potential is trained with no new DFT calculations and can even surpass the accuracy of the DFT reference labels.","feed_headline":"Uncertainty from an eleven-model ensemble cuts DFT data needs by 96%","feed_subtitle":"Weighted disagreement across pretrained potentials flags bad predictions and steers near-DFT-free distillation.","key_machinery":"The central object is $U$ itself (Eq. 1), a configuration-level score built from an eleven-member heterogeneous uMLIP ensemble with inverse-RMSE weights (Eq. 2). For each configuration, each member predicts a force on every atom; $U$ takes the atom with the largest deviation from the ensemble-mean force vector, squares that deviation for each model, and sums with weights proportional to inverse force RMSE on the OMat24 test set. The same weighted ensemble mean also serves as the DFT surrogate in distillation, and the companion UAMD protocol partitions configurations by a cutoff $U_c$: low-$U$ configurations are labeled by the teacher ensemble, high-$U$ ones by DFT. This combination of weighted disagreement and cutoff-based data selection is what carries the argument.","core_discovery":"The central claim is that $U$, defined by $U_i = \\sqrt{ \\sum_k w_k [\\max_j |F_{i,j,k} - \\langle F_{i,j}\\rangle|]^2 }$ with weights $w_k = \\mathrm{RMSE}_{F,k}^{-1} / \\sum_{k'} \\mathrm{RMSE}_{F,k'}^{-1}$, measures the true force error of uMLIP predictions for general inorganic materials. Over nearly five decades of $U$ ($10^{-3}$ to $10^2$ eV/\\AA), the conditional spread around the ideal $y=x$ line stays within about one order of magnitude, so low-$U$ configurations almost never hide large errors and high-$U$ configurations reliably flag catastrophic ones. The paper further claims that this metric keeps its predictive ordering across metals, alloys, inorganic compounds, MOFs, perovskites, and battery materials, and that a universal cutoff $U_c = 1$ eV/\\AA selects configurations with force RMSE at or below 0.1 eV/\\AA. Distillation built on this metric propagates teacher accuracy into compact ACE potentials while filtering DFT numerical noise, so the student matches or in some cases surpasses the reference labels.","pith_inferences":["If $U$'s transferability holds beyond the tested datasets, it could serve as a generic acquisition function for active learning, flagging which configurations most need DFT labels across any uMLIP application.","The paper's supplementary data show that inconsistent DFT protocols can inflate apparent errors: for example, non-spin-polarized TM23 labels for magnetic Fe and Co produce reference errors that can exceed $U$. This suggests the metric's calibration is partly conditioned on reference-label consistency, and a controlled spin-polarized benchmark would sharpen the universality claim.","The claim that teacher labels are smoother than raw DFT and therefore denoise the student could be tested directly by measuring the label-noise floor of uMLIP predictions versus DFT convergence parameters, rather than only through downstream student accuracy.","Weighting by OMat24 force RMSE is a single global ranking; if future uMLIPs specialize by chemistry or structure, per-domain weight sets might improve $U$ further, but they would also trade away the metric's simplicity."],"forward_implications":["A practical deployment monitor: configurations with $U$ exceeding roughly 1 eV/\\AA can be flagged for recalculation, with force RMSE of the retained set at or below 0.1 eV/\\AA across metals, inorganic compounds, and other material classes.","Uncertainty-aware distillation: for W, about 4% DFT labels at $U_c = 1$ eV/\\AA suffice to match full-DFT-trained ACE potentials on energies, forces, phonons, and grain-boundary tensile response, and either full DFT or full uMLIP labels are worse than the optimal mixture.","Zero-DFT potentials: for MoNbTaW, an ACE potential trained entirely on uMLIP labels matches a DFT-trained ACE in force RMSE and reproduces elastic constants, vacancy migration barriers, and stress-strain behavior, with the smoothed labels filtering DFT numerical noise.","Extensible ensemble: any future uMLIP can enter the ensemble simply by evaluating its OMat24 force RMSE, improving $U$ without retraining or calibrating a new uncertainty model.","General-purpose potentials: augmenting the MoNbTaW dataset with 7,000 maximum-volume-selected defect configurations, all labeled by uMLIPs, yields a potential that reduces errors on grain-boundary deformation and fracture test sets while remaining in the interpolation regime during molecular dynamics."],"supporting_citations":[{"why":"It supplies the OMat24 test set used to assign inverse-RMSE weights and select the eleven-member ensemble, and it is the primary benchmark for $U$.","marker":"(3)"},{"why":"It provides the public catalog of pretrained uMLIPs from which the heterogeneous ensemble is drawn.","marker":"(4)"},{"why":"It gives the Orb confidence head used as the single-model comparison baseline in calibration and coverage-accuracy tests.","marker":"(7)"},{"why":"It grounds the choice of the most accurate uMLIP as the DFT surrogate and the prior observation that uMLIPs can compete with DFT on defects and alloys.","marker":"(28)"},{"why":"It provides the tungsten defect-genome configurations used for training and independent validation of the distilled ACE potential.","marker":"(29)"},{"why":"It supplies the dimer and short-range DFT configurations that test $U$ on extreme out-of-domain geometries.","marker":"(30)"},{"why":"It provides the 17,654-configuration MoNbTaW dataset that enables zero-DFT distillation and property validation.","marker":"(36)"},{"why":"It implements the maximum-volume sampling used to generate the 7,000 defect configurations for the general-purpose MoNbTaW potential.","marker":"(39)"},{"why":"It defines the atomic cluster expansion (ACE) used as the distilled student potential.","marker":"(42)"}],"fun_headline_variants":["Heterogeneous ensemble yields universal uncertainty metric for atomistic models","Uncertainty metric from ensemble cuts DFT data needs by 96%","Distillation with uncertainty metric needs only 4% of DFT labels","Universal uncertainty metric flags bad predictions across materials","One uncertainty metric ranks risk for metals, MOFs, and perovskites"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The accuracy ranking of the eleven machine-learned potentials on the large public benchmark that sets their weights continues to hold for every new material system to which the uncertainty score is applied.","fun_headline_variants_meta":{"raw":{"variants":["Heterogeneous ensemble yields universal uncertainty metric for atomistic models","Uncertainty metric from ensemble cuts DFT data needs by 96%","Distillation with uncertainty metric needs only 4% of DFT labels","Universal uncertainty metric flags bad predictions across materials","One uncertainty metric ranks risk for metals, MOFs, and perovskites"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000722,"raw_usage":{"total_tokens":3273,"prompt_tokens":1012,"completion_tokens":2261,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":2145}},"tokens_in":628,"tokens_out":2261,"duration_ms":20199,"temperature":1.0,"reasoning_tokens":2145,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:55:32.681539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out collection of spin-polarized magnetic defects, molecular crystals, and surfaces with consistent DFT labels, compute $U$ and the true force error per configuration, and check whether the Spearman correlation remains near 0.8 or higher and whether the cutoff $U_c = 1$ eV/\\AA still separates configurations with force RMSE at or below 0.1 eV/\\AA. If either fails, the universality of $U$ and the recommended threshold would need to be abandoned or revised.","supporting_citations":[],"review_version":1}