{"id":"56cab90d-f82a-415d-99c8-82ac90ef308f","arxiv_id":"2608.11812","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Fine-tuning a pretrained DPA-4 force field on about 170 NiO PBE+U calculations reverses an incorrect phase ordering learned from no-U data, with accuracy comparable to direct fine-tuning.","lead":"The authors show that machine-learned force fields pretrained on the wrong electronic theory can be corrected with a small amount of higher-accuracy data. Using nickel oxide, they demonstrate that fine-tuning on about 170 DFT+U calculations reverses the predicted phase ordering, even when the model was previously trained on the opposite energetics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing PBE+U relaxation of Oct/Sqr endpoints: the demonstrated sign reversal may be an artifact of the fixed PBE interpolation path.","rationale":"The reader's weakest assumption identifies exactly the most load-bearing gap: the Oct–Sqr phase ordering used to define 'opposite energetics' is measured only on PBE-relaxed geometries, with no PBE+U relaxation test. This is not a minor technicality—it is the foundation on which the entire transfer experiment is built. If PBE+U relaxation reverses or eliminates the ordering, then the source and target surfaces are not actually 'opposing' in the relevant sense, and the demonstration of recovery becomes an artifact of the interpolation path. The paper's own caveat ('rather than a minimum-energy path') flags this limitation but does not resolve it. A single, well-posed DFT check—relaxing both endpoints at the PBE+U level—would settle the issue. I agree with the reader's conditional verdict because the qualitative claim is otherwise well supported: independent splits, HSE06 corroboration along the fixed path, and the two-stage transfer isolating prior wrong energetics. The quantitative 'nearly same data efficiency' clause is also under-tested, but that is secondary. The required testing is straightforward and does not undermine the paper's internal logic; it only determines whether the strategic conclusion generalizes beyond the PBE-path geometry. Therefore the appropriate verdict remains CONDITIONAL, pending the PBE+U relaxation check and release of code for reproducibility.","tokens_in":10937,"tokens_out":4557,"duration_ms":49934,"concrete_test":"Perform FM PBE+U geometry optimizations of the Oct and Sqr endpoint structures (same U_eff = 6.2 eV, 520 eV cutoff, same k-point sampling), starting from both the PBE-relaxed geometries and representative PBE+U AIMD snapshots. Compute ΔE = E(Sqr) − E(Oct) at the relaxed geometries, and also build a linear interpolation between the PBE+U-relaxed endpoints and evaluate the fine-tuned MLFFs on it. If Sqr remains above Oct by a comparable margin, the path artifact concern is resolved; if the ordering flips or becomes negligible, the central claim must be reframed as path-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that fine-tuning can reverse incorrect source-level phase energetics—rests on the premise that the PBE+U target surface truly has the opposite Oct–Sqr ordering. In Section III A and Figure 2, this ordering is established only along a fixed linear interpolation between endpoints relaxed with non-spin-polarized PBE. The paper explicitly states the path is not a minimum-energy path, yet it never checks whether PBE+U structural relaxation changes the relative stability of the endpoints. If PBE+U relaxation substantially stabilizes Oct or destabilizes Sqr (or vice versa), the 0.936 eV/f.u. Sqr–Oct difference on the PBE path could shrink, vanish, or even reverse, making the 'phase reversal' a property of the chosen interpolation rather than of the PBE+U target surface. This matters directly for the transfer experiment: the no-U-adapted model is said to encode the wrong preference, and the target dataset is said to encode the right one. If the right one only exists on a non-equilibrium path, the practical multi-fidelity conclusion is weakened. The AIMD trajectories do sample thermal relaxation, but the evaluation and the headline phase-ordering comparison remain tied to the fixed PBE geometries, so the concern is not addressed by the training data alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript asks whether a foundation machine-learned force field that was pretrained, or even previously fine-tuned, on electronic-structure data with qualitatively incorrect phase energetics can be efficiently corrected by fine-tuning on a compact target-level dataset. Using NiO as a case study, the authors build a common structural interpolation between an octahedral (Oct) and a square-planar (Sqr) phase and show that non-spin-polarized PBE favors Sqr while ferromagnetic PBE+U favors Oct. They fine-tune DPA-3 and DPA-4 foundation models on 50--290 PBE+U configurations, with and without an intermediate no-U fine-tuning stage, reporting energy and force RMSEs of roughly 0.5 meV/atom and 30 meV/Å with about 170 target labels. For DPA4-Neo, they further show that a model first fine-tuned on the no-U surface, when subsequently fine-tuned on PBE+U data, predicts a positive Sqr--Oct energy difference and thus recovers the qualitative PBE+U phase ordering. The paper concludes that incorrect source-level phase energetics can be reversed by target-level fine-tuning and proposes a multi-fidelity strategy in which broad, consistent pretraining is combined with system-specific fine-tuning.","tokens_in":11134,"tokens_out":5269,"duration_ms":56254,"significance":"The question addressed is timely and practically important for the deployment of foundation machine-learned force fields: if a small target-level dataset can override a qualitatively wrong source-level phase preference, then pretraining can prioritize label consistency and affordability over reproducing every target-level ordering. The study is well controlled in several respects: four pretrained checkpoints are compared, three independent train--test splits are used, the target-level labels are generated with consistent DFT+U settings, and HSE06 provides an external consistency check for the PBE+U ordering. The data and fine-tuned models are made publicly available. The main caveats, which I detail below, are that the target phase ordering is established only on a fixed PBE-relaxed interpolation path, and that the 'same data efficiency' claim for phase-ordering recovery is demonstrated at a single training-set size and for a single model.","major_comments":[{"comment":"The claim that PBE+U reverses the Oct--Sqr ordering is established only along a fixed linear interpolation between endpoints relaxed with non-spin-polarized PBE. Because the central transfer experiment is framed as correcting 'incorrect source-level phase energetics,' the target surface's ordering should be confirmed at PBE+U-relaxed endpoints. If PBE+U relaxation substantially stabilizes Sqr relative to Oct, the demonstrated sign reversal could be an artifact of the chosen path rather than a property of the target surface. Please add PBE+U structural relaxations of the Oct and Sqr endpoints (and ideally a small number of intermediate configurations) and report whether the ordering survives; the HSE06 check should also be extended to the relaxed endpoints if feasible. This is a load-bearing robustness test, not a presentation issue.","section":"Section III A, Fig. 2"},{"comment":"The abstract claims that no-U-adapted models recover the PBE+U phase ordering 'with nearly the same target-data efficiency as models fine-tuned directly.' However, Fig. 5 shows the phase-ordering reversal only for DPA4-Neo at N_train = 170, while Fig. 4 supports data-efficiency equivalence using aggregate energy and force RMSEs. These RMSEs are dominated by Oct-like configurations because the target dataset has much greater Oct coverage, as acknowledged in Section III C, so low RMSE does not by itself establish correct Sqr--Oct energetics. Please add a panel showing the predicted Sqr--Oct energy difference (at least its sign) versus N_train for direct and two-stage fine-tuning, including both DPA-4 models and the split-to-split spread, or alternatively adjust the central claim to state explicitly that ordering recovery was verified at one training-set size for one model.","section":"Section III C, Figs. 4 and 5"}],"minor_comments":[{"comment":"The recovered endpoint difference is 0.765 eV/f.u. versus the PBE+U target value of 0.936 eV/f.u.; the authors attribute this to uneven Sqr coverage. It would strengthen the practical message to report how many of the 170 training configurations are Sqr-like and whether targeted enrichment of the Sqr region changes the endpoint error.","section":"Section III C, Fig. 5"},{"comment":"The HSE06 agreement is described as qualitative, but the final two HSE06 points are omitted and the number of consistently converged points is not stated. Please report the range of λ for which HSE06 results were obtained and how many configurations were discarded.","section":"Section II A and Fig. 2"},{"comment":"The r2SCAN FM value of Sqr below Oct by 0.018 eV/f.u. is very small; before concluding that r2SCAN is in 'qualitative disagreement' with PBE+U and HSE06, please give an estimate of numerical uncertainty, for example from convergence tests or k-point sensitivity.","section":"Section III A"},{"comment":"The text uses informal model names such as DPA4-Air and DPA4-Neo after defining full names in Section II C; adding a table with the checkpoint identifiers and pretraining datasets would improve reproducibility.","section":"General"},{"comment":"There are minor typographical issues, including a missing space in '5×10−6eV/atom' and in 'PBE+Upredict' in the abstract; I also recommend giving a version or DOI for the QUESTS documentation in Ref. [27].","section":"Section II A"}],"recommendation":"major_revision","confidential_remarks":"The fixed-interpolation issue in Section III A is the main correctness risk; if the authors can supply PBE+U endpoint relaxations and a plot of the predicted Sqr--Oct ordering versus N_train, the paper's central claim would be adequately supported. I see no concerns about attribution or scope: Ref. [14] is a self-citation to the authors' prior fine-tuning work and is used appropriately for context."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the two-stage fine-tuning experiment is real, clean, and novel, but the 'phase reversal' is demonstrated on a fixed PBE-relaxed path, not on relaxed PBE+U structures. That narrows the practical conclusion more than the abstract admits.\n\nThe good stuff: this is the first clear demonstration I know of that a model deliberately fine-tuned to the wrong phase preference (no-U PBE) can be redirected to the opposite PBE+U ordering with roughly the same data cost as direct fine-tuning. The control is solid—three independent splits, an HSE06 cross-check, consistent FM PBE+U labels, and the two-stage comparison isolates the effect of prior wrong energetics. The Figshare deposit is a plus, and the citation pattern is clean: the prior work on label consistency [13] and fine-tuning efficiency [14] is properly acknowledged and extended. The paper is also honest about its limits: the path is not an MEP, the Sqr basin is under-sampled, and the discussion lists what could break in other materials.\n\nThe soft spots, in order. First, the path issue. Everything hinges on PBE and PBE+U disagreeing about Oct versus Sqr along a path whose endpoints were relaxed at non-spin-polarized PBE. If PBE+U relaxation flips the ordering, then the target surface the models are corrected toward isn't the real PBE+U phase stability; it's just that one path. The authors don't test this. It's not fatal for the internal logic—the AIMD data also lives on that geometry region, so the fine-tuning works as stated—but it does limit the practical offer: you can correct your model to reproduce the target surface you chose, but we don't know if that surface is the relevant one. A few PBE+U relaxations would settle it. Second, the under-sampling of Sqr is real but minor, and the authors already flag it; it explains the 0.765 vs 0.936 eV underestimation. Third, 'nearly the same data efficiency' is a visual claim; there's no statistical test. Minor. Fourth, no commit hash or exact code version; the hyperparameters are clear, but full reproducibility would need the exact training script. Minor.\n\nOverall, a well-executed case study with a genuinely new result. The reader's take is about right. I'd send it to peer review, but ask for a relaxation check at the target level before acceptance.","headline":"A well-executed two-stage fine-tuning case study whose central demonstration holds, but the phase-reversal test rests on a fixed PBE-relaxed path, so the practical claim is narrower than the abstract implies.","tokens_in":11684,"tokens_out":6907,"would_cite":false,"duration_ms":60970,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A machine-learned force field carrying the wrong NiO phase ordering can recover the correct ordering with roughly 170 DFT+U labels, almost as efficiently as direct fine-tuning.","keywords":["machine-learned force fields","DPA-4","DFT+U","NiO","fine-tuning","phase stability","multi-fidelity learning","Hubbard correction"],"falsifier":"Relax the Oct and Sqr endpoints with ferromagnetic PBE+U using consistent magnetic initialization and compare their relaxed energies; if Sqr remains below Oct at the PBE+U-relaxed geometries, the source-to-target reversal shown along the PBE-relaxed path does not reflect the true target surface.","tokens_in":10705,"feed_emoji":"⚛️","tokens_out":8158,"duration_ms":77279,"temperature":0.7,"pith_summary":"This paper asks whether a machine-learned force field can be rescued cheaply when its pretraining data gave the wrong relative stability for two competing crystal phases. In nickel oxide (NiO), plain PBE favors a square-planar phase while PBE+U—density-functional theory with an on-site Hubbard correction—favors an octahedral phase, a reversal also seen in HSE06 calculations. The authors show that DPA-4, a graph-neural-network foundation force field, reaches energy errors near 0.5 meV/atom with roughly 170 PBE+U training labels, and that even after first being fine-tuned to the wrong no-U surface it recovers the PBE+U phase ordering almost as efficiently as direct fine-tuning. If this holds, pretraining for universal force fields can prioritize cheap, consistent, broad data, leaving application-specific physics to compact fine-tuning sets.","feed_headline":"NiO force field flips its phase error with 170 DFT+U labels","feed_subtitle":"Even a model first taught the wrong NiO phase order recovers the right one as fast as direct fine-tuning.","key_machinery":"The load-bearing mechanism is two-stage fine-tuning of pretrained DPA-4 foundation models on target-level data selected for information diversity, with a fixed Oct–Sqr interpolation as the diagnostic. DPA-4 is a graph-neural-network interatomic potential pretrained on broad materials datasets; the target dataset consists of about 170 ferromagnetic PBE+U configurations chosen by an entropy-based selection criterion from constant-pressure ab initio molecular dynamics trajectories. The fixed interpolation between PBE-relaxed endpoints isolates the electronic-structure method's effect on phase ordering, and fine-tuning transfers the model from predicting Sqr 0.796 eV per formula unit below Oct to 0.765 eV per formula unit above, correcting the inherited preference.","core_discovery":"The central discovery is a demonstration of recoverable phase energetics. Along a fixed structural interpolation between the octahedral (Oct) and square-planar (Sqr) phases of NiO, non-spin-polarized PBE places Sqr 0.793 eV per formula unit below Oct, whereas ferromagnetic PBE+U places Sqr 0.936 eV per formula unit above Oct, with HSE06 in qualitative agreement. Pretrained DPA-4 models fine-tuned on roughly 170 ferromagnetic PBE+U configurations reach energy and force root-mean-square errors of about 0.5 meV/atom and 30 meV/Å; a two-stage model first fine-tuned to no-U PBE and then to PBE+U predicts Sqr 0.765 eV per formula unit above Oct, matching the sign of the target surface and nearly matching direct fine-tuning in data efficiency. The paper treats this as evidence that foundation force fields are transferable priors rather than zero-shot calculators.","pith_inferences":["The same two-stage recipe likely extends to other correlated oxides where generalized-gradient and DFT+U surfaces disagree; testing CoO or MnO along analogous paths would show whether the fast reversal is specific to NiO or generic.","A stronger implication for dataset design is that universal pretraining corpora might deliberately exclude U corrections across all transition-metal compounds, trading zero-shot accuracy for label consistency, provided downstream fine-tuning is always expected.","The paper's residual 0.765-versus-0.936 eV per formula unit underestimate suggests a testable prediction: adding more Sqr-like configurations to the target set should close most of the gap, because the authors attribute the error to under-coverage of the Sqr region.","Because the model learns whichever magnetic branch its labels encode, multi-branch targets may need explicit magnetic-state labeling; an experiment mixing FM and AFM PBE+U labels would reveal whether a single potential can represent both."],"forward_implications":["A foundation force field can be valuable even when its pretraining data encode the wrong phase ordering for a specific material.","With roughly 170 PBE+U configurations, DPA-4 reaches energy RMSE near 0.5 meV/atom and force RMSE near 30 meV/Å on NiO.","Prior fine-tuning to the opposing no-U surface does not noticeably increase the amount of target data needed to recover the PBE+U ordering.","Foundation models should be compared by how efficiently and reproducibly they adapt, not only by zero-shot error.","Pretraining data can be chosen for consistency and affordability even if they omit Hubbard-U corrections, as long as user-side fine-tuning can impose target energetics."],"supporting_citations":[{"why":"argues that consistently generated data without U is better for pretraining than mixed-U data, the premise this paper leverages.","marker":"[13]"},{"why":"shows that foundation MLFFs can be fine-tuned for specific materials with high data efficiency, the baseline capability this paper tests.","marker":"[14]"},{"why":"supplies the DPA-3 pretrained foundation models used as a comparison branch.","marker":"[15]"},{"why":"supplies the DPA-4 architecture and pretrained checkpoints used in the central transfer experiments.","marker":"[16]"},{"why":"defines the rotationally invariant DFT+U form with an effective Hubbard U of 6.2 eV that generates the target-level labels.","marker":"[22]"},{"why":"provides the information-entropy criterion used to select compact, diverse training configurations.","marker":"[28]"},{"why":"provides the fine-tuning and evaluation workflow used to train and test all models.","marker":"[30]"}],"fun_headline_variants":["170 labels flip NiO force field's phase ordering","Wrong prior, right result: NiO fine-tuning fixes phase energetics","Fine-tuning reverses NiO phase error with 170 DFT+U points","NiO force field corrected: 170 labels reverse phase ordering","Data-efficient fix flips NiO phase energetics in force field"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the premise that the fixed linear interpolation between PBE-relaxed endpoints faithfully represents the target PBE+U phase energetics; if PBE+U relaxation changed the relative stability, the sign reversal could be a path artifact.","fun_headline_variants_meta":{"raw":{"variants":["170 labels flip NiO force field's phase ordering","Wrong prior, right result: NiO fine-tuning fixes phase energetics","Fine-tuning reverses NiO phase error with 170 DFT+U points","NiO force field corrected: 170 labels reverse phase ordering","Data-efficient fix flips NiO phase energetics in force field"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000668,"raw_usage":{"total_tokens":3078,"prompt_tokens":1005,"completion_tokens":2073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":1984}},"tokens_in":621,"tokens_out":2073,"duration_ms":14260,"temperature":1.0,"reasoning_tokens":1984,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:26:24.049178+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Relax the Oct and Sqr endpoints with ferromagnetic PBE+U using consistent magnetic initialization and compare their relaxed energies; if Sqr remains below Oct at the PBE+U-relaxed geometries, the source-to-target reversal shown along the PBE-relaxed path does not reflect the true target surface.","supporting_citations":[{"cited_title":"Warford, F","cited_arxiv_id":null,"evidence_quote":"argues that consistently generated data without U is better for pretraining than mixed-U data, the premise this paper leverages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"shows that foundation MLFFs can be fine-tuned for specific materials with high data efficiency, the baseline capability this paper tests."},{"cited_title":"Schwalbe-Koda, S","cited_arxiv_id":null,"evidence_quote":"provides the information-entropy criterion used to select compact, diverse training configurations."}],"review_version":1}