{"id":"6773a171-502f-4bd4-a176-b5a8e00bed73","arxiv_id":"2412.16551","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A 10,000-material phonon benchmark shows MatterSim-v1 and SevenNet-0 predict harmonic phonons accurately, while ORB and eqV2-M fail badly.","lead":"This paper tests seven universal machine learning interatomic potentials on more than 10,000 computed phonon datasets. It finds that only some of them, especially MatterSim-v1, reproduce vibrations accurately, while force-output models such as ORB and eqV2-M perform poorly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unspecified frozen-phonon displacement makes the poor phonon ranking of ORB and eqV2-M protocol-dependent; their unusable verdict may be an artifact.","rationale":"The reader's weakest_assumption correctly identifies the unstated frozen-phonon displacement as the key risk. The paper's methods omit the displacement amplitude, and for non-conservative models the finite-difference protocol is not a neutral probe of an energy surface. In addition, the paper's own Section III acknowledges that larger displacements partially mitigate the problem, which implies the reported numbers for ORB and eqV2-M are amplitude-dependent. The strongest claim (MatterSim-v1 ready for phonons, ORB and eqV2-M unusable) therefore rests on a single protocol choice. A controlled displacement sweep would settle this. The dataset scope (non-magnetic semiconductors) is a lesser concern because the claim is scoped to semiconductors in the same section. I do not see internal inconsistency or fabrication; the conditional verdict is appropriate and should remain unchanged pending the proposed test.","tokens_in":13177,"tokens_out":3496,"duration_ms":31067,"concrete_test":"Select a representative subset of about 100 materials from the MDR/PBE dataset spanning crystal systems and chemistries. Recompute phonon properties with phonopy for ORB and eqV2-M using displacement amplitudes of 0.005, 0.01, 0.02, and 0.05 Å, with MatterSim-v1 and direct PBE as controls. Compare MAE of maximum phonon frequency, MAE of heat capacity, and the fraction of imaginary modes against the PBE reference at each amplitude. If ORB or eqV2-M error metrics improve substantially (for example, to within a factor of two of MatterSim) at any tested amplitude, the reported ranking is a protocol artifact; if they remain poor and roughly amplitude-independent, the paper's conclusion is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central ranking (especially the categorical dismissal of ORB and eqV2-M) rests on a finite-displacement protocol whose displacement amplitude is never reported in Section IV A, which only says the force constants were obtained via the finite displacement method as implemented in phonopy. For non-conservative force models, where forces are not energy derivatives, the finite-difference force constants depend explicitly on the chosen displacement and are not guaranteed to converge. Section III itself concedes the problem can be alleviated, but far from resolved, by using larger displacements in the frozen-phonon workflow. With a single, unstated default amplitude (likely phonopy's 0.01 Å), the large imaginary-mode fractions and MAEs in Tables II and III for ORB and eqV2-M may reflect a mismatch between the displacement and these models' force-field noise, rather than an intrinsic inability to describe curvatures. Since the paper's strongest claim uses this contrast to single out MatterSim-v1 as ready for phonons, the benchmark's central comparison is only conditionally valid unless the displacement dependence is quantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks seven universal machine-learning interatomic potentials (M3GNet, CHGNet, MACE-MP-0, SevenNet-0, MatterSim-v1, ORB, eqV2-M) for harmonic phonon properties against a dataset of ~10,000 non-magnetic semiconductors. The authors recalculated the MDR phonon database with the PBE functional to match the training data of the models, and they evaluate geometry errors, maximum phonon frequency, phonon DOS, vibrational entropy, Helmholtz free energy, heat capacity, sound velocity, and dynamical stability. They report that MatterSim-v1 has the smallest phonon errors, with MAEs below the PBE–PBEsol difference, while ORB and eqV2-M produce severely distorted phonons, attributed to their non-conservative force output. The paper also releases the PBE phonon dataset.","tokens_in":13392,"tokens_out":5852,"duration_ms":47585,"significance":"The benchmark is externally grounded: the seven models are frozen public checkpoints, no parameters are fitted to the phonon reference data, and the PBE reference is consistent with the models' training sets. The finding that non-conservative force models (ORB, eqV2-M) fail for finite-displacement phonons is important and aligns with existing analyses (Ref. 45). The new PBE phonon dataset is a valuable resource. However, the central ranking, especially the categorical dismissal of ORB and eqV2-M, rests on a frozen-phonon protocol whose displacement amplitude is not reported, and the claim of DFT-comparable accuracy for MatterSim-v1 lacks statistical error bars and sensitivity analysis of the ad hoc thresholds. If the displacement dependence is quantified and the statistical claims are substantiated, the paper would provide a trustworthy benchmark and a useful guide for uMLIP development.","major_comments":[{"comment":"The amplitude of the frozen-phonon displacement is never reported in Section IV A, which only states that force constants were obtained via the finite displacement method as implemented in phonopy. For conservative models this is a minor omission, but for ORB and eqV2-M, whose forces are not energy derivatives, the finite-difference force constants depend explicitly on the chosen displacement, as the paper acknowledges in Section III ('The problem can be alleviated, but far from resolved, by using larger displacements in the frozen-phonon workflow'). The large MAEs and imaginary-mode fractions in Tables II and III for these two models may therefore be an artifact of the chosen displacement rather than evidence of intrinsic inability. Please report the displacement value (and any related convergence criteria), and provide a displacement-convergence test, e.g., varying the amplitude by an order of magnitude for a representative subset, to show that the ranking of ORB/eqV2-M is protocol-independent.","section":"IV A and III"},{"comment":"The claim that MatterSim-v1's MAEs are 'considerably smaller than the difference between PBE and PBEsol' is used to conclude that it can replace DFT for phonon calculations. However, the MAE values are reported without uncertainties, and no statistical test is provided for the comparison against the PBE-PBEsol scale. Moreover, the confusion matrix in Table III depends on the ad hoc thresholds of -50 K for imaginary acoustic modes and 0.1 states/THz for the DOS. Please provide error bars (e.g., bootstrap over the dataset or standard errors across materials), and test the sensitivity of the dynamical-stability classification to these thresholds.","section":"II B, Table II, and Table III"},{"comment":"The benchmark is restricted to non-magnetic semiconductors, yet the title claims that 'Universal Machine Learning Interatomic Potentials are Ready for Phonons' and the abstract refers to 'universal applicability.' This overstates the scope of the study. The authors do mention 'semiconductors' in the main text, but the title and abstract should be tempered to reflect the material class actually tested, or the paper should explicitly discuss the potential limitations of transferring these conclusions to metals, magnetic materials, and other systems.","section":"Title, Abstract, and II A"}],"minor_comments":[{"comment":"Refs. 29 and 45 share the identical arXiv identifier (2408.00755), but they refer to different papers; one of the identifiers must be corrected.","section":"References"},{"comment":"The caption contains a typo: 'velocities' is written as 'velocties'.","section":"Table II"},{"comment":"The word 'acoustic' is misspelled as 'accoustic' in the caption.","section":"Fig. 4 caption"},{"comment":"The code name is written as 'v asp' in the text; it should be 'VASP'.","section":"IV A"},{"comment":"The label 'T etragonal' contains an erroneous space; it should read 'Tetragonal'.","section":"Fig. 1(a) axis label"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the first systematic head-to-head of seven universal MLIPs on harmonic phonons against a consistent PBE reference, and it ships a ~10,000-material PBE phonon dataset. That alone makes it worth a look. The benchmark design is sound: same relaxation pipeline, same phonopy workflow, and the PBE-PBEsol spread as a natural error scale. Credit where due — the paper is transparent about its workflow and the data and code are available. The finding that MatterSim-v1 is exceptionally good for phonons, with MAEs well below the functional spread, is credible and useful. The contrast with ORB and eqV2-M is striking and the paper correctly identifies the likely cause: these models output forces directly, not as energy derivatives, so finite-difference force constants are noisy.\n\nThe soft spots are real but proportionate. The frozen-phonon displacement amplitude is never stated in Section IV.A. For non-conservative models, the finite-difference curvatures depend on that choice, and the paper itself concedes that larger displacements help but don't fix it. So the categorical 'unusable' verdict for ORB/eqV2-M is at least partly a statement about the protocol, not just the models. The stress-test lands on an actual omission. I would also want convergence checks on the displacement and error bars on the MAEs; as it stands, the fine ranking among the mid-tier models is not well resolved. And 'universal' is a stretch for a benchmark restricted to non-magnetic semiconductors, though the authors are upfront about the dataset composition.\n\nNone of this sinks the paper. The main claim — that MatterSim is ready for phonons while non-conservative force models are not, without additional care — holds up. I would send it to peer review. A good referee should ask for the displacement sensitivity analysis and error bars, but the work deserves their time.\n\nFor us: I would cite it for the dataset and the benchmark results, and bring it to reading group.","headline":"A valuable, mostly credible benchmark of seven uMLIPs for phonons, with the caveat that the poor showing of non-conservative models is partly protocol-dependent because the frozen-phonon displacement is never stated.","tokens_in":13921,"tokens_out":1545,"would_cite":true,"duration_ms":12406,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Universal machine-learning interatomic potentials are ready for phonons, with MatterSim-v1 matching DFT accuracy while force-only models fail.","keywords":["phonons","machine learning interatomic potentials","universal MLIPs","benchmark","harmonic phonons","MatterSim-v1","dynamical stability","PBE dataset"],"falsifier":"Repeat the ORB and eqV2-M phonon calculations with a range of frozen-phonon displacement amplitudes, from roughly 0.001 to 0.05 angstroms, and check whether their mean absolute phonon-frequency errors collapse toward the level of MatterSim-v1 or remain catastrophic; if they improve substantially, the paper's ranking of these two models is an artifact of the chosen displacement, while if they stay poor, the non-conservative-force explanation is confirmed.","tokens_in":12994,"feed_emoji":"🔬","tokens_out":7437,"duration_ms":87507,"temperature":0.7,"pith_summary":"This paper asks whether universal machine-learning interatomic potentials (uMLIPs), models trained to predict energies and forces for any chemistry, can be trusted for harmonic phonons, the small-displacement response properties that control thermal behavior and dynamical stability. To answer it, the authors recalculate about 10,000 ab initio phonon calculations with the PBE functional and test seven uMLIPs on maximum phonon frequency, phonon density of states, sound velocity, vibrational entropy, free energy, and heat capacity. The central result is that one model, MatterSim-v1, reproduces PBE phonons with mean absolute errors smaller than the difference between two DFT functionals, meaning it can replace DFT for semiconductor phonon calculations. The paper also shows that two otherwise excellent models, ORB and eqV2-M, produce largely imaginary phonons because their forces are not energy derivatives, and it releases a consistent PBE phonon dataset for future development.","feed_headline":"MatterSim-v1 predicts phonons as well as DFT","feed_subtitle":"On 10,000 semiconductors, one universal potential beats the PBE-vs-PBEsol gap while ORB and eqV2-M fail.","key_machinery":"The central object is the harmonic phonon force constant, obtained by the frozen-phonon finite-displacement method: atoms are displaced by small amounts, forces are collected, and the dynamical matrix is built from numerical second derivatives of the energy. The paper's yardstick for good enough accuracy is the difference between two DFT exchange-correlation functionals, PBE and PBEsol, which bounds the intrinsic uncertainty of the reference; any uMLIP error smaller than this functional spread is treated as DFT-level accuracy. The distinction that carries the argument is conservative versus non-conservative force models: when forces are computed as exact energy gradients, finite-displacement force constants are well-behaved, but when forces are separate outputs, the implied potential is non-conservative and the second derivatives depend on the chosen displacement, which the paper identifies as the source of the imaginary phonons.","core_discovery":"The paper establishes a clear ranking of seven uMLIPs for harmonic phonons. MatterSim-v1 is the most accurate: its mean absolute errors are 17 K for maximum phonon frequency, 15 J/K/mol for vibrational entropy, and 5 kJ/mol for Helmholtz free energy, all smaller than the corresponding PBE-versus-PBEsol differences (33 K, 25 J/K/mol, and 10 kJ/mol), so it can be used as a DFT-level calculator for phonons of non-magnetic semiconductors. SevenNet-0 is the next best, followed by MACE-MP-0, CHGNet, and M3GNet, all of which systematically soften phonon frequencies. ORB and eqV2-M, despite excellent geometry predictions, fail catastrophically on phonons: their frequency distributions peak at zero and over 80 percent of dynamically unstable systems are misclassified as stable, because they output forces as independent network predictions rather than as derivatives of the energy, making the force constants required for phonons ill-defined.","pith_inferences":["The catastrophic ORB and eqV2-M phonon errors are probably amplified by the default displacement amplitude; a displacement-size study could give these non-conservative models a fairer test, so the ranking's bottom two entries should be read with this caveat.","The paper's protocol suggests a cheap addition to any uMLIP release: report phonon density of states or force-constant Hessians on a fixed benchmark, since energy and force mean absolute errors near equilibrium do not predict response-property quality.","One untested direction is to train a conservative uMLIP with phonon-derived Hessian labels or with energy-consistent force constraints; this could lift the remaining systematic softening errors seen in M3GNet, CHGNet, MACE-MP-0, and SevenNet-0.","Because the dataset excludes magnetic and metallic materials, universal readiness is demonstrated only for non-magnetic semiconductors; extending the benchmark to those classes could change the ranking."],"forward_implications":["MatterSim-v1 can be used as a drop-in replacement for DFT in high-throughput phonon screening of non-magnetic semiconductors, making dynamical-stability and thermal-property searches orders of magnitude cheaper.","Training data and its coverage matter at least as much as model architecture: a scaled-up M3GNet-style conservative model outperforms more complex equivariant networks on response properties.","Universal potentials that output forces as separate predictions are not reliable for phonons and should be redesigned to be conservative, or used with an energy-consistent correction, before being applied to response properties.","The newly released PBE phonon dataset removes the functional mismatch that previously made benchmarking uMLIP phonons ambiguous, since all tested models were trained on PBE data.","Among the models trained on the same 1.58-million-structure dataset, SevenNet-0 is clearly the best for phonons, indicating that within a fixed training set, representation and training details still set the ceiling."],"supporting_citations":[{"why":"supplies the original phonon database of about 10,000 non-magnetic semiconductors that the paper recalculates with PBE.","marker":"[34]"},{"why":"defines the PBE functional used as the benchmark reference for all uMLIP comparisons.","marker":"[39]"},{"why":"defines the PBEsol functional whose differences from PBE provide the accuracy scale for judging uMLIP errors.","marker":"[35]"},{"why":"provides the finite-displacement phonon workflow used to compute force constants.","marker":"[48]"},{"why":"extends the phonon workflow to distribution and interpolation of dynamical properties.","marker":"[49]"},{"why":"documents the non-conservative force problem that the paper invokes to explain the ORB and eqV2-M phonon failures.","marker":"[45]"},{"why":"is the MatterSim-v1 model paper, the uMLIP identified as the most accurate for phonons.","marker":"[31]"},{"why":"is the ORB model paper, the non-conservative uMLIP that fails on phonons despite accurate geometries.","marker":"[20]"},{"why":"is the eqV2-M model paper, the second non-conservative uMLIP with catastrophic phonon errors.","marker":"[27]"}],"fun_headline_variants":["MatterSim-v1 matches DFT phonon accuracy, ORB and eqV2-M fail","Phonon benchmark: MatterSim-v1 leads, two potentials crash","Universal MLIPs tested on 10k phonons: MatterSim-v1 best","Phonon accuracy reveals uMLIP strengths: MatterSim-v1 top","MatterSim-v1 beats PBE-vs-PBEsol gap for phonons"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity rests on the assumption that a frozen-phonon finite-displacement calculation with one default displacement amplitude is a fair, model-independent test, even though for non-conservative models the resulting force constants depend on the displacement chosen, and the paper does not state what that default amplitude is.","fun_headline_variants_meta":{"raw":{"variants":["MatterSim-v1 matches DFT phonon accuracy, ORB and eqV2-M fail","Phonon benchmark: MatterSim-v1 leads, two potentials crash","Universal MLIPs tested on 10k phonons: MatterSim-v1 best","Phonon accuracy reveals uMLIP strengths: MatterSim-v1 top","MatterSim-v1 beats PBE-vs-PBEsol gap for phonons"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000413,"raw_usage":{"total_tokens":2118,"prompt_tokens":913,"completion_tokens":1205,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":1096}},"tokens_in":529,"tokens_out":1205,"duration_ms":9708,"temperature":1.0,"reasoning_tokens":1096,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:28:30.709199+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Repeat the ORB and eqV2-M phonon calculations with a range of frozen-phonon displacement amplitudes, from roughly 0.001 to 0.05 angstroms, and check whether their mean absolute phonon-frequency errors collapse toward the level of MatterSim-v1 or remain catastrophic; if they improve substantially, the paper's ranking of these two models is an artifact of the chosen displacement, while if they stay poor, the non-conservative-force explanation is confirmed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the PBE functional used as the benchmark reference for all uMLIP comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the PBEsol functional whose differences from PBE provide the accuracy scale for judging uMLIP errors."},{"cited_title":"Togo, First-principles phonon calculations with Phonopy and Phono3py, J","cited_arxiv_id":null,"evidence_quote":"extends the phonon workflow to distribution and interpolation of dynamical properties."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"documents the non-conservative force problem that the paper invokes to explain the ORB and eqV2-M phonon failures."},{"cited_title":"Batatia, D","cited_arxiv_id":null,"evidence_quote":"is the ORB model paper, the non-conservative uMLIP that fails on phonons despite accurate geometries."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the eqV2-M model paper, the second non-conservative uMLIP with catastrophic phonon errors."}],"review_version":1}