{"id":"aa94966e-6327-470d-ba8e-68836c8412bb","arxiv_id":"2411.14608","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper presents the first machine learning interatomic potentials for uranium mononitride and validates them against DFT and experiment for thermal, defect, and diffusion properties.","lead":"Researchers built two machine-learning models that can simulate uranium mononitride, a promising nuclear fuel material, at temperatures up to 2700 K. The models match density-functional-theory calculations for many properties and could enable larger simulations of fuel damage and gas release.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Defect validation is in-sample: defect configurations used for Table III were seeded into the AL training data, so the 'usable for defects' claim lacks out-of-sample support until a defect-excluded retraining test is done.","rationale":"The reader's weakest assumption was PBE accuracy; that is a real limitation, but it is harder to falsify from the paper alone and is partially mitigated by the explicit PBE-vs-PBE+U phonon benchmark and by agreement with experiment for lattice parameter and heat capacity. The more immediately load-bearing weakness is internal: the defect-formation validation is seeded by the same defect types used for training. This directly affects the paper's strongest application claim. It does not invalidate the paper—finite-temperature properties, migration barriers, and cascades provide partial support—but it makes the current 'excellent agreement' for point defect energies insufficient as evidence of transferability. My recommendation remains CONDITIONAL (same as reader), with the added condition that the defect-excluded retraining test be performed or clearly reported. Hence UNCHANGED.","tokens_in":12888,"tokens_out":6219,"duration_ms":67098,"concrete_test":"After releasing the dataset and potentials, retrain both ANI and HIP-NN with identical hyperparameters on a training set from which all configurations containing each defect type (U Frenkel, N Frenkel, bound/unbound Schottky) are removed, using a leave-one-defect-type-out scheme; then recompute the Table III formation energies and Table I RMSE. If any formation energy shifts by more than ~0.2 eV from the DFT reference, or if training RMSE degrades sharply, the published defect validation is in-sample and should be downgraded from independent confirmation. In addition, report a held-out RMSE using a temporal/iteration split so that active-learning selection does not contaminate the error estimate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central 'defect modeling' claim is supported mainly by Table III, but Section II.B explicitly states that neutral Schottky defects and U/N Frenkel pairs were 'manually added to the dataset in the early iterations of AL and subsequently used for sampling,' and Fig. 1(c) confirms an U interstitial from the last AL iteration is in the final dataset. Table III then reports UFP, NFP, and bound/unbound Schottky formation energies from ANI/HIP-NN versus DFT. These are the same defect classes the models were trained on, so close agreement demonstrates that the models can fit labeled in-sample configurations; it does not establish transferability to unseen defect geometries, which is what 'can be used for modeling defects' requires. Consistently, Table I reports only 'training RMSE' and no held-out test RMSE, so generalization error is not quantified. The circularity is partial: NEB migration barriers, cascades, and Xe incorporation are not directly targeted by training and provide some transfer evidence. But the principal defect-energy validation is not independent, and the conclusion's strongest defect statement rests on it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript develops two neural-network interatomic potentials (ANI and HIP-NN) for uranium mononitride using a DFT-PBE training set enriched by an active-learning procedure. The potentials are tested against DFT and experiment for lattice parameters, elastic constants, phonon spectra, finite-temperature thermodynamic properties, stoichiometric point-defect formation energies, defect migration barriers, nitrogen self-diffusion, xenon incorporation energies, and a 1 keV collision cascade. The central claim is that these are the first machine-learning interatomic potentials for reliable atomic-scale modeling of UN at finite temperatures and that they reproduce DFT and experimental data closely enough to be used for defect and radiation-damage studies.","tokens_in":13123,"tokens_out":5039,"duration_ms":49609,"significance":"If the claims hold, the two potentials would be a valuable community resource for atomistic simulations of UN, enabling larger-scale and longer-time studies than DFT alone. The active-learning pipeline, the breadth of validation (phonon dispersion, thermal expansion, heat capacity, bulk modulus, and a four-million-atom cascade test), and the explicit comparison with classical potentials are notable strengths. The main gaps are that defect validation is performed on configurations that were seeded into the training set, no held-out test error is reported, and some numerical results (Xe interstitial incorporation, C44) contradict the strength of the conclusions drawn from them.","major_comments":[{"comment":"The defect validation is in-sample. Neutral Schottky defects and U/N Frenkel pairs were manually added to the training set in the early active-learning iterations, and Fig. 1(c) shows an uranium interstitial that entered in the final iteration. Table III then reports formation energies for exactly these defect classes. Agreement with DFT is therefore an interpolation test and does not support the conclusion that the potentials 'can be used for modeling defects' in unseen geometries. Please provide a defect-excluded retraining test (or at least a held-out set of defect configurations never used in training) and report those errors, or explicitly reframe the claim as in-sample reproduction.","section":"Section II.B, Fig. 1(c), Table III"},{"comment":"Only training RMSE values are reported for energy and forces. Since the active-learning ensemble uses randomized train/validation/test splits, held-out metrics should be available and should be reported alongside training RMSE. Without out-of-sample error, the abstract's claim that the potentials are 'reliable' for atomic-scale modeling is not quantitatively supported.","section":"Table I"},{"comment":"The conclusion states that xenon incorporation energies are 'in close agreement with DFT', but Table V reports Xe_int = 16.85 eV versus the DFT reference of 14.68 eV, a discrepancy of more than 2 eV that the text itself acknowledges. This overstatement should be corrected. The Xe_int result should either be removed from the claim or be discussed explicitly as a known limitation of combining the MLIP with the Buckingham potential for Xe.","section":"Table V and Conclusions"},{"comment":"The C44 values from both MLIPs differ from experiment by 39-43%. The manuscript attributes this discrepancy to the PBE functional, but the DFT-PBE value in Table II is 52 GPa while the MLIPs give 43-46 GPa, so the potentials also deviate from their own DFT reference by 12-17%. Please quantify this additional error and temper the elastic-constant validation claim, or provide evidence that the C44 discrepancy is entirely inherited from the training labels.","section":"Table II"},{"comment":"All reference labels use ferromagnetic PBE, and the choice is justified by the absence of imaginary phonons compared with PBE+U. However, no sensitivity check is provided for the properties most relevant to the conclusions: defect formation energies and migration barriers could depend on magnetic ordering or on the Hubbard U. A limited check (for example, recomputing the Table III defect energies with an AFM or PBE+U setup) would address the main correctness risk in transferring the potentials to real UN, whose experimental ground state is antiferromagnetic.","section":"Section II.C and Appendix C"}],"minor_comments":[{"comment":"The caption assigns '(b) highly disordered structure' and '(c) structure containing an uranium interstitial', while the main text says Fig. 1(b) shows an uranium interstitial and Fig. 1(c) shows a highly disordered configuration. The labeling should be made consistent.","section":"Figure 1"},{"comment":"The legend in Figure 6 shows 'Nvac 1.5%', while the text and caption describe concentrations of 0.5% and 1%. Please reconcile these numbers.","section":"Figure 6"},{"comment":"The reported uncertainty thresholds appear inconsistent: the text gives a maximum-force threshold of 0.024 eV/Å, which is smaller than the mean-force threshold of 0.088 eV/Å. This is likely a typo and should be corrected.","section":"Appendix B"},{"comment":"The surname Tseplyaev is spelled 'Tseplayaev' in some places, and Figure 3's caption spells Kocevski as 'Koceveski'. Please unify the spelling.","section":"Throughout"},{"comment":"The data-availability statement says the dataset and potentials 'will be available after completing LANL reviewing process'. For reproducibility, please provide a repository link or DOI at publication time (or at minimum state 'available upon reasonable request').","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially publishable after the defect-validation circularity and the held-out test reporting are addressed, and after the overstatement about Xe incorporation is corrected. The active-learning dataset and the breadth of validation are genuinely useful, but the strongest claims currently rest on in-sample evidence. A defect-excluded retraining test would be the most convincing fix and should be within reach given the existing active-learning infrastructure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what's new: this is the first MLIP for uranium mononitride, and it comes with two architectures (ANI and HIP-NN), a documented active-learning dataset of ~12k configurations, and a broad validation sweep: lattice parameter, elastic constants, phonons, heat capacity, bulk modulus vs temperature, migration barriers, nitrogen diffusion, Xe incorporation, and a 1 keV cascade. That is a lot of useful material for the nuclear fuels community, and the comparison against classical potentials is genuinely informative. The PBE vs PBE+U phonon check in Appendix C is a good methodological decision, and the authors are transparent about where the PBE reference is limiting (C44, lattice offset).\n\nWhere it gets soft. The concern about circularity is right: Table III's defect formation energies (UFP, NFP, Schottky) are computed for defect classes that were seeded into the active-learning dataset in early iterations (Section II.B). So 'excellent agreement' there is interpolation, not prediction. The paper would be stronger with a defect-excluded retraining test, or at least a held-out set of unseen defect geometries. Related: Table I reports only training RMSE, no test RMSE, so generalization error is not quantified for the final models. The Xe interstitial incorporation is off by >2 eV and the authors attribute that to the Buckingham hybrid; fair, but it limits the claim about noble gas accuracy. The migration barriers are more convincing because they weren't directly trained, though the text admits sampling with interstitials may have explored those regions.\n\nMinor stuff: the uncertainty threshold values in Appendix B look swapped (mean force threshold 0.088 vs max 0.024 eV/Å) — likely a typo. And the data availability says potentials and dataset are pending LANL review; that's a practical block for anyone wanting to build on this now.\n\nBottom line: the central defect claim is partly circular, and the paper overstates it in the abstract. But the finite-temperature MD, diffusion, and cascade tests are genuinely out-of-sample, and the resource itself (first UN MLIPs, training pipeline, hyperparameters) is valuable. This deserves a serious referee. I'd send it to review with a request for artifact release and an out-of-sample defect test, or at least explicit acknowledgment of the in-sample nature of Table III.","headline":"First MLIPs for UN with a solid active-learning pipeline; defect validation is partly in-sample, but the finite-temperature and transfer tests are real progress and the paper deserves refereeing.","tokens_in":13662,"tokens_out":2119,"would_cite":true,"duration_ms":20566,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper develops the first machine learning interatomic potentials for uranium mononitride, showing they reproduce DFT-level thermophysical properties, defect energetics, and radiation-damage behavior at molecular dynamics cost.","keywords":["uranium mononitride","machine learning interatomic potential","active learning","neural network potential","density functional theory","nuclear fuel","defect energetics","molecular dynamics"],"falsifier":"Compute the same stoichiometric defect formation energies and migration barriers with a DFT method that reproduces UN's experimental antiferromagnetic ordering and lattice parameter (e.g., PBE+U with a Hubbard U whose phonons remain real). If those values differ from the PBE reference by more than the MLIPs' training error (roughly 0.5 eV for defects), the potentials' agreement with PBE would not transfer to the more accurate reference, and a retrained potential would be needed.","tokens_in":12698,"feed_emoji":"⚛️","tokens_out":10518,"duration_ms":88581,"temperature":0.7,"pith_summary":"Uranium mononitride is a leading candidate for accident-tolerant nuclear fuel, but its atomistic modeling has been stuck between expensive density functional theory and less accurate classical potentials. This paper develops the first machine learning interatomic potentials for UN, one ANI-type and one HIP-NN-type, trained on a 12,336-configuration DFT dataset assembled through active learning. Both potentials reproduce the DFT reference for lattice parameter, heat capacity, bulk modulus, phonons, and point-defect energies, and they show promise for diffusion, noble-gas impurities, and radiation cascades. The result matters because it gives fuel performance models a fast, validated way to simulate UN under realistic reactor conditions.","feed_headline":"Machine learning potentials model uranium nitride near DFT accuracy","feed_subtitle":"Two neural network potentials match DFT for defect and thermophysical properties of UN fuel.","key_machinery":"The central machinery is an active-learning pipeline that builds a 12,336-configuration DFT (PBE) training set, combined with two local neural network potentials. The ANI potential is a Behler–Parrinello-style network that sums atomic energies with species-specific radial and angular symmetry function descriptors. The HIP-NN potential is a graph-based, message-passing network whose atomic-environment features are learned end-to-end; its core identity is $A_{i,b} = \\sum_{\\nu,j,b} V_{\\nu}\\,\\mathrm{abs}_\\nu(r_{ij}) Z_{j,b}$, which sums learned sensitivity functions of pairwise distances over all neighbors. Both potentials conserve energy by construction (forces are gradients of the total energy via automatic differentiation) and use only local atomic environments within cutoff radii of 5–7 Å. The active-learning loop, driven by a query-by-committee ensemble of eight ANI models under oscillating temperature and density perturbations, selects only high-uncertainty configurations for DFT labeling, covering crystalline, defective, and disordered states.","core_discovery":"We have developed two machine learning interatomic potentials for uranium mononitride — an ANI neural network potential and a HIP-NN message-passing potential — trained on a DFT (PBE) dataset that was enriched by an active learning loop to cover crystalline, defective, and disordered configurations. Both potentials reproduce the reference DFT energies and forces to within about 2.5 meV/atom and 0.1 eV/Å respectively, and they capture temperature-dependent lattice parameter, specific heat, and bulk modulus from 300 K to 2700 K in close agreement with DFT-MD and experiment. They also match DFT for stoichiometric point-defect formation energies and migration barriers, and the HIP-NN potential demonstrates credible nitrogen diffusion trends, xenon incorporation energies (with a hybrid classical description for Xe), and a million-atom collision cascade. This establishes the first reliable machine learning interatomic potentials for atomic-scale modeling of UN at finite temperatures, bridging the gap between accurate but costly DFT and the larger-scale simulations needed for fuel performance models.","pith_inferences":["Because all training labels are ferromagnetic PBE, the potentials are effectively surrogates for that specific functional; users should check against antiferromagnetic or PBE+U data before trusting defect energetics near magnetic transitions.","The active-learning recipe of oscillating temperature and density, seeding defects, and query-by-committee selection could transfer to other actinide fuels, where reference data are scarce and expensive.","The hybrid treatment of xenon with a classical Buckingham potential introduces more than 2 eV of error at interstitial sites; retraining a multi-species MLIP that explicitly includes Xe would likely fix this and make fission-gas modeling fully consistent.","The potentials' ability to reproduce the sharp rise in heat capacity above 1500 K suggests they capture defect-generation physics, so free-energy or thermodynamic-integration studies could use them to predict melting and phase stability."],"forward_implications":["Both MLIPs reproduce the PBE phonon dispersion including the optical modes that classical potentials miss, so the potentials support lattice dynamics and thermal property predictions.","Defect formation energies for uranium and nitrogen Frenkel pairs and Schottky defects match DFT within a few tenths of an eV, enabling large-scale defect kinetics simulations.","The HIP-NN potential runs a 1 keV collision cascade in a four-million-atom cell, demonstrating that radiation damage simulations beyond DFT's reach are feasible.","Nitrogen diffusion in hypo-stoichiometric UN is predicted to be vacancy-assisted at low temperature and interstitial-dominated at higher temperature, consistent with experimental mechanisms.","The ~0.7% lattice underestimate and the 39–43% error in the C44 elastic constant are attributed to the PBE functional rather than the MLIP fitting, so the potentials would improve automatically with a better reference functional."],"supporting_citations":[{"why":"Supplies the query-by-committee active learning method used to build the training dataset.","marker":"[29]"},{"why":"Defines the ANI neural network architecture used for one of the two potentials.","marker":"[21]"},{"why":"Defines the original HIP-NN hierarchical neural network architecture.","marker":"[22]"},{"why":"Provides the improved tensor-sensitivity HIP-NN variant with vector message passing used here.","marker":"[23]"},{"why":"Classical UN potential used as the main baseline for comparison and as the source of Buckingham parameters for Xe interactions.","marker":"[16]"},{"why":"Prior UO2 machine learning potential whose hyperparameters and active-learning setup were adapted for the ANI potential.","marker":"[19]"},{"why":"Finite-temperature DFT-MD data for UN used as reference for lattice parameter and bulk modulus comparisons.","marker":"[40]"},{"why":"Experimental lattice parameter data used to benchmark thermal expansion predictions.","marker":"[39]"},{"why":"Prior DFT study of UN magnetic ordering and exchange-correlation functionals that motivates the PBE choice over PBE+U.","marker":"[6]"}],"fun_headline_variants":["First ML interatomic potentials for uranium mononitride","Two neural nets bring DFT accuracy to UN modeling","Machine-learned potentials match DFT for uranium nitride","Active learning yields accurate ML potentials for UN fuel"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire training set is labeled with PBE density functional theory using ferromagnetic ordering on uranium, and the paper assumes this reference energy surface is accurate for the high-temperature, defective, and disordered configurations the potentials are meant to explore, so any PBE error is inherited by the potentials.","fun_headline_variants_meta":{"raw":{"variants":["First ML interatomic potentials for uranium mononitride","Two neural nets bring DFT accuracy to UN modeling","Machine-learned potentials match DFT for uranium nitride","Active learning yields accurate ML potentials for UN fuel"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000653,"raw_usage":{"total_tokens":2959,"prompt_tokens":879,"completion_tokens":2080,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":2019}},"tokens_in":495,"tokens_out":2080,"duration_ms":15756,"temperature":1.0,"reasoning_tokens":2019,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:05:48.596312+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the same stoichiometric defect formation energies and migration barriers with a DFT method that reproduces UN's experimental antiferromagnetic ordering and lattice parameter (e.g., PBE+U with a Hubbard U whose phonons remain real). If those values differ from the PBE reference by more than the MLIPs' training error (roughly 0.5 eV for defects), the potentials' agreement with PBE would not transfer to the more accurate reference, and a retrained potential would be needed.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the query-by-committee active learning method used to build the training dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the ANI neural network architecture used for one of the two potentials."},{"cited_title":"Lubbers, J","cited_arxiv_id":null,"evidence_quote":"Defines the original HIP-NN hierarchical neural network architecture."},{"cited_title":"Chigaev, J","cited_arxiv_id":null,"evidence_quote":"Provides the improved tensor-sensitivity HIP-NN variant with vector message passing used here."},{"cited_title":"Kocevski, M","cited_arxiv_id":null,"evidence_quote":"Classical UN potential used as the main baseline for comparison and as the source of Buckingham parameters for Xe interactions."},{"cited_title":"Stippell, L","cited_arxiv_id":null,"evidence_quote":"Prior UO2 machine learning potential whose hyperparameters and active-learning setup were adapted for the ANI potential."},{"cited_title":"Kocevski, D","cited_arxiv_id":null,"evidence_quote":"Finite-temperature DFT-MD data for UN used as reference for lattice parameter and bulk modulus comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Experimental lattice parameter data used to benchmark thermal expansion predictions."},{"cited_title":"Kocevski, D","cited_arxiv_id":null,"evidence_quote":"Prior DFT study of UN magnetic ordering and exchange-correlation functionals that motivates the PBE choice over PBE+U."}],"review_version":1}