{"id":"0c4317e1-01a6-498b-93d0-29b7c4209eb9","arxiv_id":"2412.10516","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"CHIPS-FF benchmarks 16 universal machine learning force fields on 104 semiconductor materials across elastic, phonon, surface, defect, interface, and amorphous properties, finding that no model works well for interfaces.","lead":"CHIPS-FF is a new open-source platform that tests how well 16 machine learning force fields predict mechanical, thermal, and defect-related properties of 104 semiconductor materials. It gives engineers and materials scientists a standardized way to choose which AI force field works best for a given material problem.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Property-level MAEs use vdW-DF-optB88 (JARVIS) reference while most uMLFFs are PBE-trained; the text concedes this 'may result in biased or inconclusive error metrics,' so model rankings and success claims (0.16 J/m2, 0.36 eV) are not yet robust.","rationale":"The reader identified the vdW-DF-optB88 vs PBE functional mismatch as the weakest assumption, and I agree. It is the most load-bearing because it directly undermines the quantitative conclusions of the paper's central claim: that the benchmark 'identifies which uMLFFs are accurate for which properties.' The paper's own Methods paragraph contains an explicit warning that such comparisons 'may result in biased or inconclusive error metrics,' which is an in-text flag of the fragility of the headline MAEs. Table 2 and Fig. 3 show the consequences: ALIGNN-FF's lattice advantage is explicitly explained by its JARVIS training data, and the h-BN c-axis example shows how vdW-sensitive properties can change by orders of magnitude with a dispersion correction. These are not minor; they mean the reported best-performer lists could change under a consistent PBE reference. The amorphous-Si comparison (Fig. 4) is also problematic because melt/quench protocols differ between uMLFF (3500 K/10 ps, 300 K/20 ps) and AIMD (2000 K/5 ps, 300 K/5 ps), so the RDF comparison conflates protocol with model accuracy; however, that affects only one property, while the functional mismatch affects all property-level MAEs. The platform itself is a genuine contribution: it is open-source, extensible, and the workflows are clearly described; the force benchmarks on training datasets are less affected by functional mismatch because forces are less functional-sensitive. Therefore, the correct verdict is CONDITIONAL: the paper should be accepted only if the authors either (a) recompute or republish the property-level rankings with a consistent functional, or (b) reframe all property-level MAEs explicitly as 'error with respect to JARVIS-DFT vdW-DF-optB88' and qualify the success claims accordingly. My recommendation does not change the reader's CONDITIONAL verdict.","tokens_in":27850,"tokens_out":9413,"duration_ms":82918,"concrete_test":"Recompute the surface-energy and vacancy-formation-energy MAEs (Fig. 3) and lattice/volume MAEs (Table 2) for a diverse subset (e.g., 10–15 of the 104 materials including layered BN, Si, Cu, GaN) using PBE ground truth generated with the same supercells and convergence settings, then compare model rankings and MAE gaps to the vdW-DF-based results. If the Spearman rank correlation of model errors between functional references is below ~0.9, or if the best-performer sets change, the reported success claims are functional-dependent and the benchmark's 'robust evaluation' claim needs to be qualified. A cheaper intermediate check is to compare ALIGNN-FF's lattice error against PBE relaxed structures from the Materials Project for the same 104 materials; if its apparent advantage disappears, the functional mismatch is confirmed as a ranking driver.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's property-level error metrics (Tables 2, 3; Fig. 3) use vdW-DF-optB88 JARVIS-DFT values as ground truth, while 14 of 16 model variants (all except ALIGNN-FF and mace-alexandria) were trained predominantly on PBE/PBEsol data (MPTrj, Alexandria, OMat, MatterSim's mix). The paper itself concedes in the Methods: 'comparing the results of a quantity computed with a uMLFF to a DFT result with an arbitrary exchange-correlation functional and varying convergence criteria may result in biased or inconclusive error metrics.' This is precisely the situation for the headline numbers. The functional mismatch is not just a uniform offset: it is structure- and property-dependent. Layered vdW materials expose it directly—orb-v2 gives over 10% error in c for h-BN (JVASP-62940), while orb-d3-v2 cuts it to 0.02%; ALIGNN-FF, the only model trained on JARVIS, has the lowest lattice-constant errors in Table 2 (0.011 Å vs 0.015–0.068 Å for others), which the text attributes to its training data. Thus the MAEs conflate model quality with training-functional compatibility. The specific success claims quoted in the abstract context—surface energy 0.16 J/m2 and vacancy formation energy 0.36 eV for 'OMat24, ORB, MACE-MPA-0, MatterSim'—are therefore not a clean measure of those models' intrinsic accuracy; they are errors with respect to a functional most of them never saw. Because the central claim of CHIPS-FF is to provide 'robust evaluation' that 'identifies which uMLFFs are accurate for which properties,' this bias is load-bearing. The platform's flexibility is a strength, but the paper's headline rankings are only known to be rankings with respect to vdW-DF-optB88 until a consistent-functional check is done. No error bars are given, so it is also unknown whether the reported MAE gaps between models are statistically meaningful.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CHIPS-FF, an open-source workflow that benchmarks sixteen universal machine-learning force-field (uMLFF) variants on a set of 104 semiconductor-relevant materials. The platform computes structural relaxations, elastic properties, bulk moduli, phonon spectra, vacancy formation energies, surface energies, interface adhesion, and melt/quench amorphous structures, and it reports force errors on the MLEARN set and on roughly two million structures from JARVIS-DFT and Materials Project trajectory datasets. All property-level errors are measured against JARVIS-DFT reference values computed with the vdW-DF-optB88 functional, and results are integrated with the JARVIS-Leaderboard. The central claim is that CHIPS-FF provides a universal, extensible benchmark that identifies which uMLFFs are reliable for which material properties.","tokens_in":1876,"tokens_out":2115,"duration_ms":73137,"significance":"If the reported benchmarks were robust, CHIPS-FF would be a valuable community resource: it is open source, covers a broader set of properties than typical energy/force leaderboards, and it benchmarks the main current uMLFF models in a single workflow. The integration with JARVIS-Leaderboard and the inclusion of computational timings and convergence statistics are practical strengths, and the paper makes concrete, falsifiable statements about model ranking. However, the validity of those ranking statements depends on controlling for the reference DFT functional, the statistical uncertainty of the error metrics, and the overlap between training data and test data. These controls are currently missing or only partially acknowledged, so the platform's usefulness for model selection is not yet demonstrated at the level claimed.","major_comments":[{"comment":"The benchmark ground truth is JARVIS-DFT (vdW-DF-optB88), while all models except ALIGNN-FF and mace-alexandria were trained on PBE/PBEsol data. The Methods text itself concedes that comparing uMLFF results to DFT with an arbitrary exchange-correlation functional \"may result in biased or inconclusive error metrics.\" This is precisely the situation for the headline property MAEs: the low lattice-constant errors of ALIGNN-FF in Table 2 (0.011 Å versus 0.015–0.068 Å) are attributed by the text to its training on JARVIS-DFT, and the surface-energy (0.16 J/m²) and vacancy-formation-energy (0.36 eV) successes claimed in the context of Fig. 3 for OMat24, ORB, MACE-MPA-0 and MatterSim are measured against a functional these models never saw. The reported rankings therefore conflate model quality with training-functional compatibility. The paper should either add a consistent-functional comparison (e.g., PBE references for at least a subset) or explicitly re-label the metrics as \"mixed-functional MAE\" and soften the claim that the platform provides \"robust evaluation.\"","section":"Methods (reference DFT functional) and Tables 2, 3; Fig. 3"},{"comment":"The amorphous-Si comparison uses mismatched melt/quench protocols: uMLFFs at 3500 K for 10 ps followed by 300 K for 20 ps with Berendsen NVT, versus AIMD at 2000 K for 5 ps followed by 300 K for 5 ps with Nosé-Hoover. Differences in the resulting RDFs can arise from protocol (quench rate, thermostat, thermal history) as much as from model accuracy, so the MAE and R² values in Fig. 4 are not a clean benchmark of the force fields' predictive quality. The authors should either run matched protocols (same temperatures, durations, and thermostat) or restrict the conclusion to \"agreement under the specified protocols.\"","section":"Methods (amorphous Si) and Fig. 4"},{"comment":"No uncertainty estimates accompany any of the reported MAEs. With 104 materials and per-model differences as small as 0.001 Å (e.g., Table 2: eqV2 31M omat versus eqV2 31M omat mp salex for lattice constant a), the ranking statements are not statistically meaningful without standard errors, bootstrap intervals, or per-property distributions. The paper discusses uncertainty quantification for MLFFs as a future need, but for a benchmarking claim, reporting only point estimates is insufficient. Add at least standard deviations or interquartile ranges across the test set.","section":"Tables 2–4; Fig. 3–4"},{"comment":"The force-error table on ALIGNN FF DB, MPF, and MPTrj measures predictions on datasets used to train several of the benchmarked models (ALIGNN-FF, M3GNet/MatGL, CHGNet, MACE, SevenNet, ORB, OMat24). These are in-distribution checks, not held-out evaluations, and they can reflect memorization rather than transferability. The text acknowledges that these datasets were used to train uMLFFs, but the framing as a benchmark conflates reproduction with generalization. The authors should separate training-set reproduction from held-out force prediction (e.g., MLEARN) and clearly label Table 5 as a training-data consistency test.","section":"Table 5"},{"comment":"The workflow includes unconverged relaxations in subsequent property calculations: \"If a calculation did not reach convergence within 200 steps, the final structure and energy at 200 steps was logged and used for subsequent portions of the workflow.\" With ALIGNN-FF showing 44% unconverged surfaces and 35% unconverged vacancies (Table 1), the surface-energy and vacancy-formation MAEs for that model in Fig. 3 are at least partly errors on non-relaxed structures. Reporting results for the converged subset alongside the full set, or excluding unconverged entries from the MAE, would make the comparison fair and reproducible.","section":"Methods (relaxation) and Table 1 vs. Fig. 3"}],"minor_comments":[{"comment":"The abstract states \"16 graph-based MLFF models,\" but Table 2 lists 16 model variants across 8 architectures; please disambiguate the wording.","section":"Abstract"},{"comment":"Equation (1) uses the elemental solid as the chemical-potential reservoir; for compounds, this convention differs from other defect-formation definitions and should be explicitly justified or compared with the JARVIS-DFT vacancy database convention.","section":"Methods, vacancy formation"},{"comment":"The sentence \"some of these models such as MACE and ORB have explicit dispersion corrections\" is imprecise; MACE-MP-0 does not include a D3 correction by default, so specify which checkpoint or version adds dispersion.","section":"Results, vdW discussion"},{"comment":"There are minor typographical issues: \"Aprroximation\" (p. 12), \"Wycoff\" (p. 9), \"outweighing factors\" (p. 7), \"a users own\" (p. 12), and \"Wychkoff\" (Fig. S1).","section":"General"},{"comment":"The statement that data will be made available \"upon publication\" is vague; provide a persistent DOI or repository link in the manuscript.","section":"Data availability"},{"comment":"The work-of-adhesion MAE is computed against a mixed experimental/theoretical reference set; specify which entries are experimental and which are computed, since the two are not directly commensurable.","section":"Fig. 3c"}],"recommendation":"major_revision","confidential_remarks":"This is a useful infrastructure manuscript with a clear scope and an unusual breadth of property-level tests. The main issue is that the authors themselves acknowledge the functional-mismatch limitation but do not let that acknowledgment sufficiently temper the headline conclusions. A revision that adds consistent-functional comparisons, uncertainty metrics, and matched amorphous-protocol runs, and that relabels Table 5 as an in-distribution check, would make the claims defensible. The paper is well suited to a computational-materials or data-driven-materials venue if these points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dan, quick take on CHIPS-FF. It's a genuinely useful piece of infrastructure: an open-source workflow that runs 16 universal MLFFs through relaxations, EOS fits, elastic tensors, phonons, surfaces, vacancies, interfaces, and melt-quench MD on a 104-material semiconductor test set, with JARVIS-Leaderboard integration. The phonon displacement analysis is the sharpest part—showing OMat and ORB get noisy at small displacements while MACE, SevenNet, and MatterSim stay flat, consistent with Lowe et al. That alone is worth citing. The force evaluation on nearly two million structures is also a solid contribution.\n\nThe main soft spot is exactly what the stress-test note flags, and the paper concedes it: the JARVIS-DFT ground truth is vdW-DF-optB88, while most of these models were trained on PBE or PBEsol data. So the headline numbers—surface energy MAE 0.16 J/m2, vacancy MAE 0.36 eV, lattice-constant rankings—mix model error with functional mismatch. The h-BN c-axis example makes it concrete: orb-v2 gives over 10% error in c, orb-d3-v2 cuts it to 0.02%. Those are different physics, not just different model quality. ALIGNN-FF's low lattice errors are partly self-consistency, which the text acknowledges. No error bars anywhere, so we don't know if the gaps between models are statistically meaningful. The a-Si comparison also uses mismatched melt/quench protocols (3500 K for 10 ps vs 2000 K for 5 ps), which the text mentions but doesn't address. These are addressable, but they are load-bearing for the quantitative rankings. The paper's own Methods sentence—'may result in biased or inconclusive error metrics'—is the honest summary.\n\nThe interface work-of-adhesion failure and the OMat/ORB phonon noise are the most interesting scientific observations, and both are plausible despite the caveats. The platform's flexibility is real: users can plug in their own DFT reference, which mitigates the functional issue going forward. Data availability is 'upon publication' and Figshare isn't populated yet, so the artifacts aren't independently checkable right now.\n\nBottom line: this deserves a serious referee. The platform and dataset are worth having, and most of the criticism is fixable with a consistent-functional subset, error bars, and a protocol-matched a-Si run. I'd cite it for the phonon-displacement result and use the platform as a starting point, but I'd wait to quote the property-level MAEs until the functional check is done.","headline":"A genuinely useful benchmarking platform with real new data, but the headline property-level MAEs compare PBE-trained models against a vdW-DF-optB88 ground truth, so treat the rankings as conditional.","tokens_in":28830,"tokens_out":1893,"would_cite":true,"duration_ms":16247,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CHIPS-FF is an open-source benchmarking platform that evaluates universal machine-learning force fields on material properties beyond energy—lattice constants, elastic constants, phonons, vacancy formation energy, surface energy…","keywords":["Machine learning force field","deep learning","foundational models","density functional theory","high-throughput","materials discovery","semiconductors"],"falsifier":"Re-run the CHIPS-FF benchmark on the same 104 materials and 16 models using a ground truth computed with the PBE functional (matching the training data of most models), and check whether the relative rankings by MAE for lattice constants, elastic constants, surface energy, and vacancy formation energy change materially; if they do, the reported accuracy comparisons are artifacts of the functional mismatch rather than intrinsic model quality.","tokens_in":27578,"feed_emoji":"⚛️","tokens_out":12742,"duration_ms":104582,"temperature":0.7,"pith_summary":"This paper introduces CHIPS-FF, an open-source benchmarking platform that evaluates universal machine-learning force fields (uMLFFs) on material properties beyond simple energy prediction—lattice constants and volume, bulk modulus and elastic tensor, phonon band structure, vacancy formation energy, surface energy, interface work of adhesion, and amorphous-phase structure from melt/quench molecular dynamics. The platform is demonstrated on 104 semiconductor-relevant materials spanning metals, semiconductors, and insulators, using 16 pretrained graph-based models, and on force prediction for close to two million atomic structures. The paper's central claim is that such a standardized workflow gives the community a robust way to compare uMLFFs on the properties that matter for device applications, and its first benchmark run identifies OMat24, ORB, MACE-MPA-0, and MatterSim as the most accurate models for surface energy (MAE of 0.16 J/$m^{2}$) and vacancy formation energy (MAE of 0.36 eV), while also exposing weaknesses such as large phonon errors at small displacements for ORB and OMat models.","feed_headline":"Benchmark ranks 16 AI force fields on real material properties","feed_subtitle":"CHIPS-FF tests surfaces, defects, phonons and amorphous phases across 104 semiconductor materials.","key_machinery":"The load-bearing mechanism is the CHIPS-FF workflow itself: a Python pipeline that connects an atomistic simulation environment with a materials-data toolkit and drives 16 graph-based uMLFF calculators through structural relaxation using a robust cell filter, equation-of-state fitting, elastic-tensor computation, phonon band structure via finite displacements at four magnitudes, vacancy and surface supercell generation from reference databases, interface construction using a lattice-matching algorithm, and melt/quench molecular dynamics for amorphous phases. Ground truth for the bulk, elastic, phonon, vacancy, and surface benchmarks is the vdW-DF-optB88 reference data, with errors reported as mean absolute errors and formatted for direct upload to an interactive leaderboard.","core_discovery":"The paper reports a head-to-head comparison of 16 universal machine-learning force fields on properties beyond energy. ALIGNN-FF, trained on the same vdW-corrected reference data used for ground truth, captures lattice constants most accurately, while OMat24 and ORB models also relax structures well, with ORB roughly an order of magnitude cheaper. MACE-MPA-0 and MatterSim give the best simultaneous predictions of the elastic constants C11 and C44, and OMat24, ORB, MACE-MPA-0, and MatterSim reach the best surface-energy (0.16 J/$m^{2}$) and vacancy-formation-energy (0.36 eV) errors. Phonon calculations show that ORB and OMat models degrade sharply at small finite displacements, consistent with noisy forces in the low-force regime, while MACE and MatterSim remain stable. For amorphous silicon, invariant models such as ORB and MatterSim match or beat equivariant models on the radial distribution function, and no model predicts interface work of adhesion accurately, which the authors attribute to the lack of interface data in training sets.","pith_inferences":["A consistent-functional re-benchmark (e.g., PBE ground truth from the training-data source of most models) would separate genuine model quality from training-data functional effects and could shift the model rankings for elastic and surface properties.","Because CHIPS-FF is dataset-agnostic, running it against experimental reference values or multiple DFT functionals would produce functional-agnostic leaderboards, addressing the limitation the paper acknowledges about biased error metrics.","The platform's modular design (JSON input, command-line tools, leaderboard uploads) makes it straightforward to add the uncertainty-quantification layer the paper identifies as missing for most uMLFFs.","The combination of strong scaling and small-displacement phonon noise in ORB suggests a targeted benchmark on anharmonic properties such as thermal conductivity would clarify whether their speed is worth the vibrational-accuracy cost in device-thermal simulations."],"forward_implications":["New uMLFFs can be screened on semiconductor-relevant properties before large-scale deployment, since CHIPS-FF automatically records convergence, accuracy, and per-stage timing.","ORB and MatterSim emerge as cost-effective choices for relaxing large defect and surface supercells, while OMat24 and MACE-MPA-0 offer top accuracy at higher computational cost.","The large phonon errors of ORB and OMat at small displacements imply their forces are noisy in the low-force regime, which matters for vibrational and thermal-property calculations.","The consistently poor work-of-adhesion predictions mean none of the tested uMLFFs should be trusted for interface energetics without fine-tuning on interface data.","The comparable amorphous-Si accuracy of invariant and equivariant models raises the question of whether equivariance is necessary for such properties, a question the paper leaves for further benchmarking."],"supporting_citations":[{"why":"Supplies the reference structures and ground-truth bulk, elastic, and phonon data for the benchmark.","marker":"[37,38]"},{"why":"Provides the defect supercells used to compare vacancy formation energies on identical structures.","marker":"[15]"},{"why":"Generates the interface structures and supplies the experimental/theoretical work-of-adhesion reference values.","marker":"[20]"},{"why":"Provides the eqV2 model family and the OMat training data benchmarked in the workflow.","marker":"[83]"},{"why":"Supplies the MACE foundation potentials whose energy, phonon, and force performance is assessed.","marker":"[53]"},{"why":"Documents the MatterSim model whose accuracy and scaling support the cost-efficiency conclusions.","marker":"[74]"},{"why":"Documents the ORB architecture and its low computational cost, which the benchmark verifies.","marker":"[54]"},{"why":"Provides the small force-prediction test set (Cu, Ni, Li, Mo, Si, Ge) used to compare force accuracy.","marker":"[125]"},{"why":"Reports the prior phonon benchmark corroborating the small-displacement phonon errors for ORB and OMat models.","marker":"[90]"}],"fun_headline_variants":["CHIPS-FF pits 16 AI force fields on tricky material properties","16 ML force fields ranked on defects, surfaces, phonons, more","New benchmark reveals best AI force fields for real materials","AI force fields compared on elastic, phonon, and defect accuracy","Benchmark surfaces top MLFFs for semiconductor materials"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark treats DFT results computed with the vdW-DF-optB88 functional as the truth for all models, even though most of those models were trained on PBE data from a different repository, so the reported errors could be dominated by a functional mismatch rather than by model quality.","fun_headline_variants_meta":{"raw":{"variants":["CHIPS-FF pits 16 AI force fields on tricky material properties","16 ML force fields ranked on defects, surfaces, phonons, more","New benchmark reveals best AI force fields for real materials","AI force fields compared on elastic, phonon, and defect accuracy","Benchmark surfaces top MLFFs for semiconductor materials"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000241,"raw_usage":{"total_tokens":1544,"prompt_tokens":988,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":469}},"tokens_in":604,"tokens_out":556,"duration_ms":5614,"temperature":1.0,"reasoning_tokens":469,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:52:36.887342+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the CHIPS-FF benchmark on the same 104 materials and 16 models using a ground truth computed with the PBE functional (matching the training data of most models), and check whether the relative rankings by MAE for lattice constants, elastic constants, surface energy, and vacancy formation energy change materially; if they do, the reported accuracy comparisons are artifacts of the functional mismatch rather than intrinsic model quality.","supporting_citations":[],"review_version":1}