{"id":"e35099fd-02b4-4a57-8973-51a6f8416cf3","arxiv_id":"2608.06662","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Geometry-diverse fine-tuning, not single-geometry adaptation, best restores accuracy of foundation interatomic potentials on ZrO2 nanostructures.","lead":"The authors built a DFT dataset of zirconia structures spanning bulk, slabs, particles, necks, and atomic wires, and tested 26 pretrained machine learning interatomic potentials on them. They found that accuracy drops sharply on low-coordination geometries, and that fine-tuning on a mix of geometries, rather than one geometry, best reduces cross-geometry errors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fine-tuning and mixed-geometry gains may be inflated by trajectory-correlated train/test frames: the Sec. II.B split is within geometry classes, not block-wise by AIMD run.","rationale":"I read the paper as a careful benchmark whose main prescriptive claim is that adapting foundation MLIPs to geometry-changing ionic nanostructures requires geometry-diverse target data and property-specific validation. That claim rests on two pillars: supervised adaptation comparisons (FT vs FS, single- vs mixed-geometry) and property-level validations. The second pillar is honestly caveated (partial surface overlap, single MD realization per model), and the zero-shot geometry-degradation trend is robust. The first pillar, however, is only as strong as the train/test split. The data generation is dominated by quasi-continuous AIMD trajectories; random frame-level splits are a standard way to overstate generalization in such datasets, and no trajectory-block split or duplicate-removal step is reported. The conclusion explicitly lists trajectory-level data correlation as an open question, confirming the authors recognize the risk. The reader's chosen weakest assumption (DFT reference settings) is legitimate but bears mainly on absolute error magnitudes and property comparisons against mp-2858, not on the internal supervised comparison that supports the diversity recommendation. The abstract/SI energy-RMSE inconsistency is a separate reporting issue. Because the dataset and code are released, the requested reanalysis is feasible; the appropriate disposition remains CONDITIONAL, with the added condition that the adaptation results survive a trajectory-aware split.","tokens_in":27283,"tokens_out":9379,"duration_ms":98128,"concrete_test":"Re-run the 50-epoch fine-tuning/from-scratch and mixed-geometry protocols (Secs. III.C–III.E) with a trajectory-aware split: assign entire contiguous blocks of each AIMD/tensile or thermalization run—grouped by seed and elongation step—to train, validation, and test, and additionally remove any test frame whose minimum RMSD to a training frame is below about 0.1 Å. If FT<FS and the mixed-geometry disparity reduction in Fig. 4C persist under this split, the concern is settled; if the gains shrink, reverse, or move to different geometry classes, the central adaptation claim is an artifact of trajectory-level leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the geometry-stratified split in Sec. II.B produces an honest test set for the adaptation experiments. Neck and wire configurations come from continuous AIMD/tensile histories (Sec. II.A: 'progressively separating' and 'thermalizing at 300 K after each increment'), so within-class assignment almost certainly leaves near-duplicate or heavily overlapping local environments on both sides of the train/test boundary. No clustering by trajectory, seed, or elongation step is reported, and the conclusion itself lists 'trajectory-level data correlations' as an unresolved remaining question. Sections III.C and III.E and Fig. 4C—fine-tuning beats from-scratch at fixed epochs, and mixed-geometry exposure reduces cross-geometry disparity—are exactly the results most inflated by such leakage, because a model can memorize frames from the same run rather than learn transferable physics. The zero-shot geometry trends and the external property validations are less affected, but the paper's prescriptive claim that geometry-diverse target data are required rests mainly on the supervised adaptation results. This is a different and more central vulnerability than the missing DFT settings or the abstract/SI 6 vs 31.27 meV/atom discrepancy.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks 26 pretrained machine-learned interatomic potentials on a newly generated DFT dataset of ZrO2 configurations spanning bulk, slab, particle, neck, and wire geometries, motivated by an experimentally observed desintering process. The authors report geometry-dependent zero-shot error degradation, compare zero-shot inference, fine-tuning, and from-scratch training under a fixed 50-epoch protocol, evaluate geometry-specific versus mixed-geometry fine-tuning, and validate selected models on elastic, vibrational, surface-energy, and neck-dynamics properties. The central claims are that geometry-diverse target data improve cross-geometry transfer and that property-level rankings are not predicted by aggregate energy and force errors.","tokens_in":27491,"tokens_out":3138,"duration_ms":33407,"significance":"If the claims hold, the paper provides a valuable and reusable benchmark for a chemically and structurally challenging system, with a public dataset and code. The study is careful to separate reference-energy alignment from model error, to compare training strategies under controlled optimizer settings, and to acknowledge several limitations (partly interpolative surface tests, single-trajectory dynamics, incomplete optimization). These strengths make the paper a useful contribution for practitioners adapting foundation MLIPs to low-coordination ionic nanostructures. However, the headline zero-shot result contains an unresolved internal inconsistency, and the trajectory-correlated data split may materially inflate the supervised adaptation conclusions, so the central prescriptive claims should be treated as conditionally supported pending revision.","major_comments":[{"comment":"The abstract and Section III.B state that ORB-V3 achieves a zero-shot aligned energy RMSE of 6 meV/atom, but SI Table S3 lists 31.27 meV/atom for the aligned energy RMSE of the same model (and 30.36 meV/atom raw). This discrepancy concerns the paper's headline quantitative result and must be reconciled; the authors should state whether Fig. 2A and the abstract use a different metric or a different subset, and correct the text, table, or figure accordingly.","section":"Abstract and Section III.B vs SI Table S3"},{"comment":"The DFT reference dataset is described only as 'first-principles simulations based on DFT, as implemented in VASP', without specifying the exchange-correlation functional, pseudopotentials, k-point sampling, energy cutoff, or convergence thresholds. Because Sections II.F and II.G validate against Materials Project mp-2858 values, the reported property errors mix genuine model error with possible reference mismatches. These computational settings must be provided for the dataset to be reproducible and for the property-level comparisons to be interpretable.","section":"Section II.A"},{"comment":"The train/test split is stratified within geometry classes but not block-wise by AIMD trajectory, seed, or elongation step. Since neck and wire structures come from continuous AIMD and tensile histories (Section II.A), near-duplicate or highly correlated frames may appear on both sides of the split. This can inflate the fine-tuning and mixed-geometry gains reported in Fig. 4C and Sections III.C/III.E, which are the main evidence for the paper's prescriptive claim that geometry-diverse target data are necessary. The authors should add a trajectory-blocked or structurally de-duplicated split to show that the conclusions survive.","section":"Section II.B and Sections III.C/III.E"},{"comment":"Section II.H states that the tensile MD simulations were repeated with three random seeds, but Section IV states that the MD analysis is based on a single trajectory per model, and Section III.F.4 refers to 'a single initial-velocity realization'. This internal inconsistency matters because the observed failure mode (from-scratch ORB rupture) may be seed-specific. The authors should clarify how many seeds underlie Fig. 5D and whether the qualitative conclusions are stable across seeds.","section":"Section II.H vs Section IV and Section III.F.4"}],"minor_comments":[{"comment":"In Eq. (1), the sentence 'For each model, element-specific [41, 46] were determined' appears to be missing the noun 'corrections'.","section":"Section II.C"},{"comment":"The caption of Fig. 1(B) reads 'between two [26, 27]' and appears to be missing a noun such as 'grains'.","section":"Figure 1 caption"},{"comment":"Table S8 contains an unresolved cross-reference ('as summarized in Table ??') that should be fixed before publication.","section":"SI Table S8"},{"comment":"The manuscript should state explicitly that the ORB-V3 checkpoint used throughout is the Direct-20 variant, since the SI (Fig. S3) compares Direct-20 and Conservative-20 variants and the distinction is relevant for reproducibility.","section":"Section II.D"},{"comment":"The embedded text 'MatBenchDiscoveryAssessment' in the workflow diagram is unclear and should be replaced with a readable label or legend entry.","section":"Figure 1 text"}],"recommendation":"major_revision","confidential_remarks":"The dataset and code availability are a genuine asset, and the authors have been unusually candid about limitations. However, the 6 vs 31.27 meV/atom inconsistency in the headline zero-shot number and the absence of a trajectory-blocked split are substantive enough that the paper should not be accepted without revision. In my view, both are fixable within the manuscript's scope, and the reviewers should ask to see the corrected numbers and a robustness check of the supervised adaptation results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful empirical benchmark paper with a real dataset, but its headline adaptation claims rest on a train/test split that likely leaks trajectory-correlated frames, and the abstract's 6 meV/atom ORB-V3 number disagrees with the SI's 31.27 meV/atom. The geometry-resolved zero-shot trends and property tests are solid enough to be worth refereeing; the supervised fine-tuning comparisons need a block-wise split by AIMD run before the prescriptive conclusions can be trusted.\n\nWhat's new: the five-class ZrO2 dataset (bulk, slab, particle, neck, wire) tied to a real desintering process, plus the controlled comparison of zero-shot, fine-tuned, from-scratch, geometry-specific and mixed-geometry training. The negative-transfer matrices and the mixed-geometry stratification are genuinely useful. They also ship code and data (Zenodo), and they are honest about the surface-energy overlap and the single-trajectory MD caveats. The elastic/phonon validations are a good idea even if the reference mismatch with Materials Project muddies absolute numbers.\n\nSoft spots, in order of severity. First, the split in Section II.B is within geometry classes, not block-wise by AIMD/tensile run. Neck and wire frames come from continuous histories, so near-duplicate local environments almost certainly appear on both sides of the split. That means the fine-tuning-beats-from-scratch and mixed-geometry-reduces-disparity results (Figs. 3, 4C) could be memorization rather than transferable learning. The paper itself flags 'trajectory-level data correlations' as an open question, but that doesn't fix the central claim. This is the load-bearing issue, bigger than the numeric inconsistency. Second, the 6 vs 31.27 meV/atom discrepancy is unresolved and undermines the headline number; if Table S3 is correct, the abstract overstates ORB-V3 by 5x. Third, the DFT settings (functional, pseudopotentials, k-points) are never given, so comparing to Materials Project reference values mixes model error with reference mismatch. The surface-energy test is partly interpolative, but they say so; that's a minor caveat, not a flaw.\n\nSummary: this paper is for MLIP developers and benchmark builders. The dataset is a reusable contribution and the zero-shot geometry trends are credible. But the prescriptive 'geometry-diverse data required' conclusion depends on the supervised adaptation results, which need a trajectory-aware split first. I'd send it to peer review, but with a requirement to redo the adaptation comparison on block-wise splits and resolve the abstract/SI mismatch.","headline":"Useful ZrO2 geometry-diverse benchmark and zero-shot assessment, but the fine-tuning claims need a trajectory-aware split before they are believable.","tokens_in":28043,"tokens_out":2293,"would_cite":true,"duration_ms":21301,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Universal machine-learned interatomic potentials transfer poorly to low-coordination zirconia nanostructures in zero-shot use, and mixed-geometry fine-tuning on a five-class dataset reduces cross-geometry errors while single-geometry…","keywords":["machine learning interatomic potentials","transferability","fine-tuning","zero-shot inference","zirconia nanostructures","cross-geometry generalization","density functional theory dataset","molecular dynamics"],"falsifier":"Re-run the zero-shot and mixed-geometry fine-tuning evaluations on a second ionic oxide, for example HfO2, using the same five geometry classes; if cross-geometry error reduction does not reproduce, the diversity conclusion is specific to ZrO2 rather than general. Separately, recompute the reported elastic and phonon errors against a DFT reference whose exchange-correlation functional and pseudopotentials are documented, to see whether the property rankings survive a controlled reference.","tokens_in":27063,"feed_emoji":"⚛️","tokens_out":6051,"duration_ms":55245,"temperature":0.7,"pith_summary":"This paper asks whether universal machine-learned interatomic potentials, pretrained mostly on bulk crystal data, can be trusted when a simulation passes through surfaces, particles, necks, and atomically thin wires, as happens when ZrO2 grains desinter under tension. It builds a DFT dataset spanning all five geometry classes and shows that zero-shot predictions degrade sharply in the low-coordination classes: the best pretrained model reaches a force error of about 197 meV/Å overall, with the largest errors on necks and wires. Under a fixed 50-epoch training budget, fine-tuning beats training from scratch, but fine-tuning on one geometry alone often hurts other geometries, while mixed-geometry fine-tuning reduces the spread of errors across geometries. The paper's conclusion is that adapting foundation potentials to ionic nanostructures requires target data that explicitly cover the structural environments of interest, plus property-level validation, because aggregate energy and force errors do not predict elastic, vibrational, surface, or dynamical accuracy.","feed_headline":"Mixed-geometry data makes ML force fields transfer to nanowires","feed_subtitle":"Zero-shot models stumble on necks and wires; fine-tuning on all five ZrO2 geometries cuts the gap.","key_machinery":"The load-bearing object is the geometry-diverse ZrO2 dataset itself: bulk, slab, particle, neck, and wire configurations generated from DFT-based ab initio molecular dynamics and quasistatic elongation, with a geometry-stratified train/validation/test split applied before a 3 eV/Å force filter removes high-force frames. The evaluation machinery is a controlled three-way comparison of zero-shot inference, full-parameter fine-tuning from a pretrained checkpoint, and from-scratch training, all run for 50 epochs with reference-energy alignment by fitted elemental offsets so that energy errors reflect the shape of the potential-energy surface rather than reference-level shifts. Geometry-specific nested subsets and mixed-geometry subsets at controlled atomic-environment counts isolate the effect of structural diversity from dataset size. Property-level validations of elastic constants, phonons, surface energies, and tensile neck molecular dynamics then test whether improvements in supervised error metrics propagate to physical observables.","core_discovery":"The central discovery is that geometry diversity in the adaptation data, not just the amount of data, is what makes pretrained MLIPs transfer across a structural process like the ZrO2 desintering. On the authors' dataset, geometry-specific fine-tuning improves in-domain accuracy but frequently produces negative transfer, with wire-only fine-tuning degrading energy errors on several other classes, whereas mixed-geometry fine-tuning at controlled total atomic environments lowers cross-geometry errors for both MACE and ORB checkpoints. The paper also establishes that pretrained initialization provides a data-efficiency advantage: at equal optimization epochs, fine-tuned models beat from-scratch models with comparable wall-clock time. Property-level tests show that no single strategy is uniformly best: zero-shot models retain the lowest phonon errors and competitive elasticity, while adaptation helps most on surface energies and neck dynamics, and average error rankings fail to predict these outcomes.","pith_inferences":["The negative-transfer pattern is a practical instance of catastrophic forgetting: full-parameter fine-tuning on a narrow geometry overwrites representations that remain useful elsewhere, suggesting that replay-based or parameter-efficient strategies may be needed for multi-geometry deployment.","The dataset's design could be reused as an out-of-distribution benchmark for other ionic oxides, and replacing manual geometry labels with automated active-learning structure generation would test whether explicit diversity is strictly necessary or merely sufficient.","Because long-range electrostatics are not isolated in the experiments, the residual wire and neck errors may partly reflect missing charge physics; charge-aware or polarizable potentials offer a direct test of that hypothesis.","The property-level hierarchy suggests a practical validation recipe: check phonon curvature and surface energetics before trusting molecular dynamics on far-from-equilibrium nanostructures."],"forward_implications":["Foundation MLIPs should not be used zero-shot for low-coordination ionic nanostructures; even the best zero-shot force error exceeds commonly cited accuracy thresholds for stable dynamics.","Adapting to a process that spans bulk, surface, particle, neck, and wire motifs requires training data drawn from all of those classes, not just the target class.","Geometry-specific fine-tuning can backfire outside its domain, so deployment scope should determine whether narrow or diverse fine-tuning is appropriate.","Model rankings based on average energy and force RMSE are insufficient; independent physical validation is needed before trusting elastic, phonon, or dynamical predictions.","Pretrained initialization is a data-efficiency win at a fixed optimization budget, not a training-cost reduction, since fine-tuning and from-scratch training take similar wall-clock time."],"supporting_citations":[{"why":"Provides the fine-tuning-versus-from-scratch and reference-energy consistency findings that frame the training protocol.","marker":"[12]"},{"why":"Reports the experimentally observed ZrO2 desintering and ultrathin-wire formation that motivates the five-class dataset.","marker":"[26]"},{"why":"Supplies a materials-science benchmark for MLIPs that the authors position their out-of-distribution dataset as complementing.","marker":"[36]"},{"why":"Documents that bulk-dominated pretraining datasets underrepresent low-dimensional environments, explaining the observed zero-shot degradation.","marker":"[37]"},{"why":"Establishes the surface-focused assessment approach that the paper extends to necks and wires.","marker":"[38]"},{"why":"Gives the cross-functional transferability and reference-energy consistency evidence used to justify training shifts.","marker":"[41]"},{"why":"Supplies the cross-domain transfer optimization approach that mixed-geometry fine-tuning builds on.","marker":"[46]"},{"why":"Defines the Matbench Discovery MP-compliant versus MP-non-compliant grouping used to organize the 26-model benchmark.","marker":"[50]"},{"why":"Is the MACE foundation checkpoint selected as the representative MP-compliant architecture.","marker":"[55]"},{"why":"Is the ORB-V3 checkpoint that achieves the best zero-shot energy and force errors and serves as the representative MP-non-compliant model.","marker":"[57]"}],"fun_headline_variants":["Mixed-geometry fine-tuning beats single-geometry for MLIP transfer","Zero-shot MLIPs fail on nanowires; mixed-geometry tuning recovers","Geometry diversity, not data volume, drives MLIP transfer success","Cross-geometry transfer needs diverse training, not just more data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes the density-functional-theory settings used to generate the ZrO2 reference data are consistent with the Materials Project reference used for elastic and phonon validation, but those settings are not reported in the paper.","fun_headline_variants_meta":{"raw":{"variants":["Mixed-geometry fine-tuning beats single-geometry for MLIP transfer","Zero-shot MLIPs fail on nanowires; mixed-geometry tuning recovers","Geometry diversity, not data volume, drives MLIP transfer success","Cross-geometry transfer needs diverse training, not just more data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000274,"raw_usage":{"total_tokens":1668,"prompt_tokens":1004,"completion_tokens":664,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":588}},"tokens_in":620,"tokens_out":664,"duration_ms":6753,"temperature":1.0,"reasoning_tokens":588,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:55:12.806019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the zero-shot and mixed-geometry fine-tuning evaluations on a second ionic oxide, for example HfO2, using the same five geometry classes; if cross-geometry error reduction does not reproduce, the diversity conclusion is specific to ZrO2 rather than general. Separately, recompute the reported elastic and phonon errors against a DFT reference whose exchange-correlation functional and pseudopotentials are documented, to see whether the property rankings survive a controlled reference.","supporting_citations":[{"cited_title":"Batatia, D","cited_arxiv_id":null,"evidence_quote":"Provides the fine-tuning-versus-from-scratch and reference-energy consistency findings that frame the training protocol."},{"cited_title":"Pa¸ sca, H","cited_arxiv_id":null,"evidence_quote":"Reports the experimentally observed ZrO2 desintering and ultrathin-wire formation that motivates the five-class dataset."},{"cited_title":"Kabylda, J","cited_arxiv_id":null,"evidence_quote":"Documents that bulk-dominated pretraining datasets underrepresent low-dimensional environments, explaining the observed zero-shot degradation."},{"cited_title":"Batatia, W","cited_arxiv_id":null,"evidence_quote":"Establishes the surface-focused assessment approach that the paper extends to necks and wires."},{"cited_title":"Jacobs, D","cited_arxiv_id":null,"evidence_quote":"Gives the cross-functional transferability and reference-energy consistency evidence used to justify training shifts."},{"cited_title":"Alampara, M","cited_arxiv_id":null,"evidence_quote":"Supplies the cross-domain transfer optimization approach that mixed-geometry fine-tuning builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Matbench Discovery MP-compliant versus MP-non-compliant grouping used to organize the 26-model benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the MACE foundation checkpoint selected as the representative MP-compliant architecture."}],"review_version":1}