{"id":"f1ed0947-a83e-4582-932a-750046f01535","arxiv_id":"2501.11191","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Neuroevolution potentials, trained with an evolutionary strategy and running on GPUs, match or approach quantum-accurate energies and forces while simulating systems with millions of atoms, at speeds far beyond competing machine-learned potentials.","lead":"Neural network potentials called NEP, implemented in the open-source GPUMD package, promise near-quantum accuracy with the speed of classical force fields, enabling million-atom molecular dynamics simulations. This review benchmarks NEP against other machine-learned potentials on carbon and demonstrates new applications in surfaces, friction, and alloys.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Speed benchmark compares NEP in GPUMD against non-native LAMMPS runs and unmatched model sizes; the headline efficiency claim needs a controlled, accuracy-matched benchmark.","rationale":"I read the paper as a method review with illustrative new applications, and the strongest, most consequential claim is the efficiency one because it appears in the abstract, Sec. III.C, and Sec. VII. The benchmark evidence is suggestive but not controlled: different MD engines, different hardware configurations, and different model sizes are compared simultaneously. The reader's weakest_assumption identified the implementation/hardware mismatch; I agree and sharpen it by noting that accuracy is not matched, so the speed comparison is not a like-for-like test of NEP as a method. The new case studies (Pt reconstruction, CNT growth, alloy fatigue) are demonstrations rather than decisive evidence; even if their validation is thin, they do not carry the central claim as much as the speed benchmark does. The NEP accuracy results on carbon are credible, the models and data are deposited, and other independent works support NEP's practical competitiveness, so the paper is not fatally flawed. However, the headline efficiency claim should remain conditional on a fair, controlled speed comparison that separates method from implementation and accounts for accuracy differences.","tokens_in":50130,"tokens_out":7746,"duration_ms":82750,"concrete_test":"Reproduce Fig. 8 under a controlled protocol: (1) train or take all models on the Rowe carbon set with hyperparameters tuned so that their force RMSEs on the same held-out test set are within 10% (e.g., use a smaller MACE/NequIP variant if needed); (2) run each model through its native engine on the same single V100 and compute atom-step/s at N = 64k, 512k, and 4M atoms; (3) additionally run the same NEP model through the NEP_CPU/LAMMPS interface on the same 64 CPU cores used for GAP and compare directly in that CPU-only setting. If NEP remains fastest at matched accuracy and retains a large margin in the CPU comparison, the central claim is supported. If the margin disappears or shrinks in these controlled runs, the claim should be reworded to 'NEP in GPUMD is fastest in this implementation setup' rather than claiming a general method-level speed advantage.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing claim is Section III.B.4 / Fig. 8: NEP is significantly faster than DP, GAP, MACE, and NequIP, especially for large systems. What this figure actually compares is software stacks: NEP runs in native GPUMD v3.9.5 on a V100, DP/MACE/NequIP run in LAMMPS on the same GPU, and GAP runs on 64 Xeon CPU cores with no GPU. The paper itself notes in Sec. II.B.6 that NEP also has a separate LAMMPS CPU implementation, but no speed test of that implementation is reported. Thus the observed speed gap conflates the NEP architecture with GPUMD's optimized kernels, neighbor lists, and memory layout. A second confound is without accuracy matching: MACE is more accurate on the carbon training set (Figs. 3 and 4) but slower, so a smaller MACE or NequIP model could narrow the speed gap substantially. The conclusion that NEP achieves significantly higher computational speeds is therefore not yet established as a property of the NEP method per se; it is established for one NEP model running in one package against particular non-native implementations. This matters because the abstract and Sec. VII use this efficiency as the central reason NEP enables routine million-atom near-first-principles MD.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review provides a comprehensive overview of the neuroevolution potential (NEP) method, including its neural-network architecture, descriptor construction, loss function, and practical workflows implemented in the GPUMD package. The authors benchmark NEP against DP, GAP, MACE, and NequIP on a general-purpose carbon dataset, evaluating training accuracy, physical-property predictions (bilayer graphene binding and sliding energies, amorphous carbon sp3 fractions), and computational speed. They then survey NEP applications to structural properties, phase transitions, and mechanical behavior, and introduce several new case studies: a Pt(001) surface reconstruction model, a FeC model for carbon nanotube growth, an hBN sliding model for tribology, and UNEP-v1-based simulations of a complex alloy under compression, impact, and fatigue. The central conclusion is that NEP combines near first-principles accuracy with exceptional computational efficiency, enabling million-atom MD simulations on a single GPU.","tokens_in":50378,"tokens_out":5180,"duration_ms":50968,"significance":"If the efficiency claim is correct and the benchmark is fair, the paper makes a strong case for NEP as a practical tool for large-scale accurate atomistic simulations, with implications for materials discovery and mechanistic studies. The paper is strengthened by the open availability of code, training data, and scripts, and by the inclusion of new, reproducible case studies. However, the speed comparison conflates the NEP model with the GPUMD implementation, and the accuracy claims are tempered by the paper's own binding-energy results; these issues must be resolved before the headline claims about NEP's 'exceptional efficiency' and 'near-first-principles accuracy' are fully supported.","major_comments":[{"comment":"The speed benchmark is not a controlled comparison of the NEP method against other MLP methods. NEP runs in its native GPUMD v3.9.5 on a single V100 GPU, while DP, MACE, and NequIP run in LAMMPS on the same GPU, and GAP runs on 64 CPU cores. The paper itself notes in Sec. II.B.6 that NEP has a separate CPU/LAMMPS implementation, but no speed test of that implementation is reported. As a result, the observed speed gap conflates the NEP architecture with GPUMD's optimized kernels, neighbor-list handling, and memory layout. In addition, no accuracy matching is performed: the models differ in accuracy (e.g., MACE is more accurate than NEP on the carbon training set, Figs. 3 and 4), so a smaller MACE or NequIP model could substantially narrow the speed gap. The conclusion in Sec. VII that NEP offers 'empirical potential-like efficiency' and enables routine million-atom near-first-principles MD rests on this uncontrolled comparison. We request either (a) a benchmark that includes NEP running in LAMMPS or the competing methods running in their native packages, with accuracy-matched model sizes, or (b) a revision of the claims to state explicitly that the speed advantage is demonstrated for the NEP implementation in GPUMD, not for the NEP method per se.","section":"Sec. III.B.4, Fig. 8"},{"comment":"The paper repeatedly claims that NEP provides near-first-principles accuracy (abstract and Sec. VII), but its own Fig. 5 shows that NEP significantly overestimates the equilibrium interlayer spacing of AB-stacked bilayer graphene and deviates from the DFT binding energy curve. The text acknowledges that none of the MLPs accurately locate the global minimum, but this caveat is not carried through to the summary claims of 'near-first-principles accuracy' made elsewhere. Since this is a fundamental property (van der Waals binding) that is relevant to many of the applications discussed in the review, the authors should either temper the accuracy claim throughout the paper or provide evidence that this failure is an isolated outlier, not representative of typical NEP performance on the benchmarked systems.","section":"Sec. III.B.2, Fig. 5"},{"comment":"The Pt(001) surface reconstruction is presented as a new case study demonstrating the capability of the NEP approach, but the only validation is qualitative consistency with one prior DP simulation (Qian et al.). No direct comparison is made to DFT energies of the reconstructed and unreconstructed surfaces, nor to any experimental characterization of the reconstruction. Given that the paper explicitly states that the NEP model's test RMSEs (7.76 meV/atom energy, 145.46 meV/Å force) are 'relatively higher' than the previous DP model, the claim that the NEP model reliably captures the subtle energetics of surface reconstruction would be more convincing with a quantitative benchmark, such as comparing surface energies of different reconstructions against DFT. As written, the case study illustrates that NEP can produce a similar trajectory to DP on one system, but it does not independently establish the accuracy of the NEP prediction.","section":"Sec. V.B, Fig. 20"}],"minor_comments":[{"comment":"The phrase 'int the output layer' should read 'in the output layer'.","section":"Sec. II.B.1"},{"comment":"The word 'trainig' appears in the sentence describing RMSE values ('... for the trainig dataset'); correct to 'training'.","section":"Sec. II.B.7"},{"comment":"The section title uses 'Pt(001)' while the text and Fig. 20 consistently use 'Pt(100)'; the Miller-index notation should be unified.","section":"Sec. V.B"},{"comment":"The caption refers to a 'neuroevolution potential (ENP) model' but the method is NEP; correct the abbreviation.","section":"Fig. 23"},{"comment":"The phrase 'As a hindsight' is ungrammatical; suggest rewording to something like 'To improve diversity, we augmented the dataset with 300 liquid structures...'.","section":"Sec. V.B"},{"comment":"The Data availability statement mentions a 'zenodo repository' but provides no URL or DOI; please include the direct identifier.","section":"Sec. VII (Data availability)"},{"comment":"The sentence 'simulations consisting of 100 steps were run' is ambiguous about whether timings include the 100 steps only or are averaged over a longer run after equilibration; please clarify the timing protocol.","section":"Sec. III.B.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is written by the developers of NEP, which is appropriate for a method review, but the benchmark asymmetry (pre-trained NEP and GAP models vs. newly trained DP/MACE/NequIP models) and the software-stack confound in the speed test might be perceived as favoring NEP. The authors should make the limitations of the benchmark explicit and ideally add an accuracy-matched comparison. The new case studies are interesting but need stronger validation, as noted in the major comments. No concerns about novelty disclosure beyond the heavy self-citation pattern, which is typical for this type of review and not a reason for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a well-organized and honest review of the NEP method by the people who developed it, and it includes several genuinely new case studies: Pt(001) surface reconstruction, CNT growth on an Fe55 cluster, twisted hBN superlubricity, and a small suite of mechanical simulations of a compositionally complex alloy. Second, the headline efficiency claim—that NEP is dramatically faster than DP, GAP, MACE, and NequIP—is not actually established by the benchmark they run. The speed comparison pits NEP in its native GPUMD package against DP/MACE/NequIP running inside LAMMPS and GAP on 64 CPU cores. That conflates the NEP architecture with GPUMD's optimized kernels and neighbor lists. The paper itself notes that NEP has a separate LAMMPS CPU implementation, but they never benchmark that version. As it stands, Fig. 8 shows GPUMD is fast, not that NEP as a regression model is inherently faster than the others.\n\nWhat the paper does well: the carbon benchmark is genuinely useful. It compares training accuracy, physical observables (bilayer graphene binding/sliding, amorphous carbon sp3 fraction), and speed across five MLP families. The new models are trained on published datasets, full hyperparameters are given in the text, and the authors deposited data and models on Zenodo. That is reproducible and deserves credit. The review also surveys a broad range of applications and is candid about NEP's limitations, including the lack of explicit electrostatics and the absence of a full periodic-table foundation model.\n\nThe soft spots are proportionate. The selection bias the reader flagged (pre-trained NEP and GAP vs. self-trained others) is real but minor, since the training choices look reasonable and accuracy is not the load-bearing claim. More important is the speed figure, which the abstract and conclusions lean on heavily. The new case studies are valid demonstrations, not deep validations: Pt(001) is compared only to a prior DP study, the CNT growth model has a force RMSE of about 385 meV/Å, and the fatigue simulations run 10 cycles with no error bars. That is acceptable for illustrative purposes, but the framing should say so.\n\nWho should read this: anyone considering NEP for large-scale MD or wanting a compact map of its application space. It deserves a serious referee, not a desk reject. The main revision needed is to either run a controlled speed benchmark (same engine, accuracy-matched models) or substantially soften the efficiency claim. I'd also ask for error bars on the mechanical case studies.","headline":"A useful, reproducible review of the NEP method with several new case studies, but the headline speed advantage is a software-stack comparison, not a settled property of the NEP architecture.","tokens_in":50957,"tokens_out":1658,"would_cite":true,"duration_ms":18261,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["71.15.Pd","34.20.-b"],"model":"deepseek-v4-flash","headline":"This review argues that a neuroevolution-trained machine-learned potential can deliver near-first-principles accuracy at speeds that make million-atom molecular dynamics routine.","keywords":["neuroevolution potential","machine-learned interatomic potentials","molecular dynamics","GPUMD","computational efficiency","amorphous carbon","short-range order","high-entropy alloys"],"falsifier":"Re-run the diamond-system speed benchmark of Figure 8 with all five methods on identical hardware, one GPU or NEP through its LAMMPS CPU interface, each with its most optimized native implementation; if NEP is not clearly the fastest at $10^6$ atoms, the headline efficiency claim collapses. On the accuracy side, train a fresh NEP on the same carbon dataset and quench amorphous carbon at 3.0 g cm$^{-3}$; if the resulting sp$^3$ fraction falls outside the 60-80% range reported for the other methods, the adequate-accuracy half of the claim fails.","tokens_in":49916,"feed_emoji":"⚛️","tokens_out":13049,"duration_ms":109176,"temperature":0.7,"pith_summary":"This review argues that the neuroevolution potential (NEP), a machine-learned interatomic potential trained on density-functional-theory data and implemented in the open-source GPUMD package, has reached the point where it can serve as a near-first-principles engine for routine molecular dynamics. The central evidence is a head-to-head benchmark against four other machine-learned potentials on a shared carbon dataset, in which NEP is roughly an order of magnitude faster than its nearest GPU-based rival on large systems while still reproducing the energetics and bonding statistics that matter for materials physics. If that claim holds, simulations that once required massive computing campaigns, such as short-range order in alloys, radiation cascades, fracture, and tribology, become accessible on desktop GPUs. The review covers about three dozen NEP applications and contributes new case studies of its own: platinum surface reconstruction, carbon-nanotube growth on an iron cluster, superlubric sliding of twisted bilayer hexagonal boron nitride, and mechanical loading of a quinary complex alloy.","feed_headline":"Simulate millions of atoms at near-first-principles accuracy on one GPU","feed_subtitle":"The review benchmarks the neuroevolution potential against four rivals and shows the million-atom problems it unlocks.","key_machinery":"The load-bearing object is the NEP descriptor-network pair. Every atom contributes a site energy $U_i = \\mathcal{N}(\\mathbf{q}_i)$, where the descriptor vector $\\mathbf{q}_i$ is assembled from radial functions expanded in Chebyshev polynomials with trainable, species-pair-dependent coefficients and from angular components built by summing products of radial functions with spherical harmonics; total energy is the sum of site energies. Deliberately cheap descriptors are what make the speed claim concrete, and a species-pair-dependent coefficient set is what lets multi-element models cost almost the same as single-element ones. Training uses a separable natural evolution strategy, a derivative-free, black-box optimizer, on a loss that jointly fits energy, force, and virial with $L_1$ and $L_2$ penalties. The generalization argument behind the 16-element UNEP-v1 model is that one- and two-component structures already outline the descriptor space, so higher-component alloys are interpolation points rather than new territories; that is the mechanism that lets a model trained on binaries describe a quinary alloy with no retraining.","core_discovery":"The paper's central claim is that NEP occupies a genuinely new point in the accuracy-cost plane of interatomic potentials: training accuracy on par with the best machine-learned potentials, a descriptor whose evaluation cost is nearly independent of the number of chemical species, and inference speed that lets a single GPU drive million-atom systems. The speed advantage is quantified in the benchmark: on a common 6088-structure carbon dataset, NEP is the fastest of the five methods on one V100 GPU and can sustain simulations of about six million atoms on that card, MACE and NequIP are the slowest, DP sits in between, and GAP on 64 CPU cores matches DP on one GPU. Accuracy is presented as adequate rather than best-in-class: MACE posts the lowest root-mean-square errors, while NEP and DP are comparable to each other and ahead of GAP; in physically decisive tests, namely the sp$^3$ fraction of quenched amorphous carbon and the binding and sliding energy landscapes of bilayer graphene, NEP is among the models that agree with DFT and experiment. The review's application chapters are then offered as evidence of what this combination unlocks: million-atom radiation cascades, nano-tribology of incommensurate interfaces, short-range order sampled at experimental length scales, and phase transitions followed at near-experimental heating rates.","pith_inferences":["A direct test of the speed claim's portability would be a five-model comparison inside a single engine, such as running NEP and its rivals through their LAMMPS interfaces on the same GPU, which would separate the method's intrinsic cost from the GPUMD engine's optimizations.","If the UNEP-v1 descriptor-space interpolation argument is right, a NEP foundation model trained only on elemental and binary data could be stress-tested first on ternary and quaternary high-entropy alloys outside the original 16 elements.","The force re-weighting trick used for GeSn, which emphasizes small forces in the loss to improve energy-landscape minimization, could be exported to surface-reconstruction and phase-transition models where near-equilibrium forces dominate the physics.","The water result suggests a general recipe, train a NEP on a many-body-corrected reference and add nuclear quantum effects through path integrals, that could be carried to other hydrogen-bonded or proton-transferring systems."],"forward_implications":["Near-first-principles molecular dynamics of million-atom systems becomes routine on a single desktop GPU, and multi-GPU runs reach tens of millions to one hundred million atoms.","The NEP-ZBL combination puts primary radiation damage within reach at sizes where defect-cluster statistics become physically meaningful, for both elemental metals and high-entropy alloys.","Short-range order in alloy systems can be sampled by Monte Carlo and molecular dynamics at volumes matching characterization tools such as atom-probe tomography, not just at DFT-cell sizes.","A periodic-table foundation model becomes a data-efficient target: training on elemental and binary structures may be enough for accurate multicomponent predictions.","Disordered and hydrogen-bonded materials, including amorphous carbon and water, can be modeled at near-quantum-chemical accuracy with empirical-potential-level cost when nuclear quantum effects are added."],"supporting_citations":[{"why":"Introduces the neuroevolution potential method whose descriptor, evolutionary training, and efficiency are the subject of this review.","marker":"[9]"},{"why":"Supplies the GPUMD engine that performs the GPU-based training and inference on which the speed and memory claims rest.","marker":"[10]"},{"why":"Documents the NEP3/GPUMD implementation and hyperparameter conventions used throughout the review's case studies.","marker":"[62]"},{"why":"Provides the general-purpose carbon dataset and the pretrained GAP model that anchor the five-method benchmark in Section III.","marker":"[67]"},{"why":"Provides the pretrained NEP carbon model evaluated alongside the newly trained DP, MACE, and NequIP models in the benchmark.","marker":"[92]"},{"why":"Deep potential, the main rival whose accuracy is shown comparable and whose speed is shown slower than NEP.","marker":"[64]"},{"why":"Gaussian approximation potential, the baseline that the carbon benchmark and surface-reconstruction discussion assess NEP against.","marker":"[63]"},{"why":"MACE, the accuracy-leading baseline that posts the lowest errors in the benchmark.","marker":"[66]"},{"why":"NequIP, the equivariant message-passing baseline whose accuracy and speed are contrasted with NEP.","marker":"[65]"},{"why":"UNEP-v1, the 16-element general-purpose NEP whose construction and multicomponent applications anchor the alloy and mechanical-property sections.","marker":"[33]"}],"fun_headline_variants":["Neuroevolution potentials: near-DFT accuracy at million-atom speed","GPUMD's NEP: one GPU drives millions of atoms with ML accuracy","Machine-learned potentials that scale to million-atom simulations on one GPU","NEP: the ML potential that unlocks million-atom MD on one GPU"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The speed comparison assumes that the chosen software setups are typical: GAP runs on 64 CPU cores rather than a GPU, and every non-NEP model runs inside LAMMPS while NEP runs in its native GPUMD engine, so a rival model with a differently optimized or differently hosted implementation could close the measured gap.","fun_headline_variants_meta":{"raw":{"variants":["Neuroevolution potentials: near-DFT accuracy at million-atom speed","GPUMD's NEP: one GPU drives millions of atoms with ML accuracy","Machine-learned potentials that scale to million-atom simulations on one GPU","NEP: the ML potential that unlocks million-atom MD on one GPU"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000322,"raw_usage":{"total_tokens":1858,"prompt_tokens":1043,"completion_tokens":815,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":733}},"tokens_in":659,"tokens_out":815,"duration_ms":7336,"temperature":1.0,"reasoning_tokens":733,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:32:02.503957+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the diamond-system speed benchmark of Figure 8 with all five methods on identical hardware, one GPU or NEP through its LAMMPS CPU interface, each with its most optimized native implementation; if NEP is not clearly the fastest at $10^6$ atoms, the headline efficiency claim collapses. On the accuracy side, train a fresh NEP on the same carbon dataset and quench amorphous carbon at 3.0 g cm$^{-3}$; if the resulting sp$^3$ fraction falls outside the 60-80% range reported for the other methods, the adequate-accuracy half of the claim fails.","supporting_citations":[{"cited_title":"Fan , author Y","cited_arxiv_id":null,"evidence_quote":"Provides the pretrained NEP carbon model evaluated alongside the newly trained DP, MACE, and NequIP models in the benchmark."}],"review_version":1}