{"id":"71d244b8-e7d8-4205-85fc-0e501467cdb9","arxiv_id":"2505.06462","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"MPNICE matches near-best accuracy on a wide range of organic and inorganic benchmarks while aiming for 5-20x faster inference.","lead":"A new machine learning force field called MPNICE predicts atomic charges during simulation and includes long-range electrostatics, running several times faster than comparable models. The authors benchmark it on organic liquids, molecular crystals, inorganic materials, and platinum-iridium complexes it was never trained on.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No measured speedup: MPNICE timings are reported but both comparison models OOM, so the 5–20x faster claim is unsupported.","rationale":"The reader's weakest_assumption correctly flags xTB charge bias, and the authors themselves note reduced dielectric and Born-charge magnitudes for NaCl. That is a real limitation, but it does not defeat the central claim: the paper's charge-dependent demonstrations (LO-TO splitting via cancellation, vertical IP/EA) still hold. The efficiency claim, by contrast, is a quantitative headline assertion that is empirically unverified in the main text. Section 3.4 does not contain a single measured speedup; both comparison models OOM, and the conclusion's 'order of magnitude' overstates Section 3.4's 'competitive speed.' Because the paper's contribution is explicitly 'efficient long-range MLFFs,' the missing controlled timing comparison is the most load-bearing concern. My check would settle it directly. I therefore keep the reader's CONDITIONAL verdict: the accuracy benchmarks are substantial, but the headline speed claim needs a direct measurement against the named models in a controlled setting before acceptance as stated.","tokens_in":30923,"tokens_out":4534,"duration_ms":45640,"concrete_test":"Run a controlled benchmark on the same L4 GPU: (1) install SevenNet-0 and MatterSim-v1.0.0-1M as ASE calculators and in LAMMPS; (2) build the water, diamond, and aluminum boxes used in Section 3.4 at experimental density; (3) run 500 NVT steps with each model, recording per-atom wall time per timestep; (4) run MPNICE in Desmond on the same boxes, and if possible also as an ASE/LAMMPS calculator to isolate engine effects; (5) compute speedups. If the median MPNICE speedup over these comparable-accuracy models is below 5x, the abstract's '5-20x faster' claim should be revised to the measured range.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central efficiency claim ('5-20x faster inference versus models with comparable accuracy') is not supported by the data in Section 3.4. The section reports only MPNICE absolute timings in Desmond (76/36/20 µs/atom for diamond/water/aluminum). The two named comparison models, MatterSim-v1.0.0-1M and SevenNet-0, run out of memory on the smallest diamond cell (2744 atoms) and no timing is reported for them at any system size. The conclusion repeats 'an order of magnitude faster than comparable models,' but no speedup factor or comparative timing appears in the main text. The comparison is also not apples-to-apples: MPNICE runs in the specialized Desmond engine while the competitors run as ASE calculators, which can differ in neighbor-list construction, batching, and memory footprint. Since 'efficient' is half of the paper's contribution, the unmeasured speedup is a load-bearing gap. If the real speed advantage is smaller, or contingent on engine-specific optimizations, the central claim is materially overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MPNICE, an invariant message-passing machine-learned force field that iteratively predicts atomic partial charges through an approximate Qeq scheme and includes explicit electrostatics and optional D3 dispersion. The authors train a suite of models for organic molecules (direct, delta-learned, and crystal-aware), for inorganic crystals (MPtrj, OMAT24a, and a sequential inorganic model), and for hybrid organic/inorganic multi-task and shared-force training. The accuracy benchmarks are extensive: torsion scans (Genentech, TorsionNet500, and a new Torsion2000 set), tautomer energies, organic crystal lattice energies and polymorph ranking, liquid densities, water properties, ionization energies and electron affinities, monoelemental crystal ranking, Matbench Discovery, QMOF, elastic moduli, phonon MDR benchmarks, the NaCl non-analytic phonon correction, Li diffusion in amorphous LiAlO2, and Pt/Ir organometallic structure optimization. The paper also claims in the abstract and conclusions that MPNICE achieves 5-20x faster inference than comparable models, a claim that Section 3.4 does not presently substantiate with comparative timing data.","tokens_in":31189,"tokens_out":6091,"duration_ms":63759,"significance":"If the benchmark results and the efficiency claim both hold, MPNICE would be a significant contribution: it combines near-best-class accuracy on several organic and inorganic tasks with a charge-aware architecture that enables charge-dependent properties without an equivariant backbone. The manuscript also provides reusable test data, notably DLPNO-CCSD(T) reference energies for TorsionNet500, a new Torsion2000 set, tautomer and crystal test geometries, and a large set of liquid-density benchmarks; these are valuable community assets. The shared-force hybrid training idea is original and is tested on a demanding organometallic zero-shot task. The authors are also appropriately explicit in Section 4 about limitations such as spin-state insensitivity, extrapolative use on reactive surfaces and defects, and untested organic/inorganic interfaces. The central efficiency claim, however, is currently unsupported by comparative measurements, and the charge-dependent property demonstrations inherit a documented bias from training to GFN1-xTB charges.","major_comments":[{"comment":"The headline claim of \"5-20x faster inference\" and the conclusion's \"order of magnitude faster than comparable models\" are not supported by the evidence in Section 3.4. That section reports absolute MPNICE timings (76, 36, and 20 microseconds per atom for diamond, water, and aluminum) but provides no timing for MatterSim-v1.0.0-1M or SevenNet-0; both competitors are described as running out of memory on the smallest diamond cell, so no speedup ratio is measured anywhere. The comparison is also not apples-to-apples because MPNICE runs in the Desmond MD engine while the competitors run as ASE calculators, which can differ in neighbor-list handling, batching, and memory footprint. The authors should measure all models under a common driver at matched system sizes, report timings for at least one system where every model fits in memory, and state the resulting speedup factors; if this is not possible, the abstract and conclusions should be revised to claim only competitive or engine-specific performance.","section":"Abstract; Section 3.4, Fig. 12"},{"comment":"The inorganic models are trained to reproduce GFN1-xTB partial charges because MPtrj has no charge labels, and Section 3.2.3 explicitly states that the magnitudes of the dielectric tensor and Born effective charges are reduced relative to PBEsol and that unphysical off-diagonal terms appear for NaCl. Since the abstract and Section 4 present \"direct prediction of charge-dependent properties\" as a headline capability, the current evidence supports only qualitative response tensors for the studied NaCl case, with the successful LO-TO splitting partly due to error cancellation between Z* and the dielectric tensor. The authors should either validate the charge model against a higher-quality reference for at least one representative system, or explicitly qualify the charge-dependent-property claim in the abstract and conclusions as qualitative.","section":"Section 3.2; Section 3.2.3"},{"comment":"Table 9 excludes elastic-moduli outliers (values below -50 or above 600 GPa) from the reported MAE and R2 statistics, and the number of excluded outliers is not negligible in absolute terms: Inorganic MPNICE has 82 excluded shear-modulus outliers, and its reported GVRH R2 is only 0.52 even after exclusion. Because Section 3.2.2 concludes that \"all models achieve reasonable performance, comparable to previous reports of MACE MP 0,\" the authors should report outlier counts as fractions of the test set, give statistics both with and without outlier exclusion, and justify the exclusion threshold with a sensitivity analysis.","section":"Section 3.2.2, Table 9"}],"minor_comments":[{"comment":"The parameters eta, r_a, theta_a, and zeta in the AEV definitions are not all defined near the equations; please add explicit definitions and state the numerical values used for these parameters.","section":"Section 2, Eqs. (1)-(3)"},{"comment":"The timings are averaged over 500 steps, but no variance or error bars are reported; please report standard errors and state whether the quoted times include only neural-network inference or also engine overhead such as neighbor-list construction and integration.","section":"Section 3.4, Fig. 12"},{"comment":"The new torsion benchmark is referred to as \"Torsion2000\" in the text and Figure 2 but as \"TorsionTest2000\" in Table 12 and the Data Availability section; please standardize the name.","section":"Section 3.1.1 vs. Table 12 and Data Availability"},{"comment":"There is a typo in the sentence beginning \"Interstingly, Kovacs et al. report...\"; \"Interstingly\" should be \"Interestingly\".","section":"Section 3.1.5"},{"comment":"Please specify the exact definitions used for the fixed-ion dielectric tensor and Born effective charges, including the Ewald summation convention and whether clamped-ion conditions are imposed, so that the automatic-differentiation procedure can be reproduced.","section":"Section 3.2.3, Eqs. (20)-(21)"}],"recommendation":"major_revision","confidential_remarks":"The efficiency claim is repeated in the abstract and conclusions but rests on a single section with no comparative timing, so I would ask the editor to judge the revision specifically on whether real speedup measurements are added or the claim is removed. Note also that the models themselves are only available through the commercial Schrodinger Suite, while the test data are public; this asymmetry is a reproducibility consideration that the authors should address. The accuracy benchmarks themselves are extensive and, with the above qualifications, appear sound."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core of this paper is a large, well-designed benchmark suite for a charge-aware invariant message-passing force field, and on that front it mostly delivers. What's genuinely new: the iterative Qeq charge equilibration combined with message passing, the shared-force hybrid training that lets one output head serve two levels of theory, the external-field response formalism with dielectric tensors and Born charges, and two new reference datasets (Torsion2000 and DLPNO-CCSD(T) TorsionNet500). The benchmarks are extensive and the results are competitive: MPNICE matches or beats existing transferable organic potentials on rotamers, tautomers, and crystal ranking, and the inorganic models sit near SevenNet and MatterSim on Matbench Discovery while using fewer parameters. The zero-shot LiAlO2 diffusion barrier is a nice applications-level validation.\n\nThe soft spots are real but localized. The biggest is the speed claim. The abstract says \"5-20x faster inference versus models with comparable accuracy,\" but Section 3.4 gives only MPNICE timings; MatterSim and SevenNet both OOM on the smallest diamond cell and no comparative number ever appears. The conclusion repeats \"an order of magnitude faster,\" which is not supported by data in the paper. Given that efficiency is half the contribution, this needs fixing--either by reporting timings on systems where the competitors fit, or by softening the claim to \"competitive speed\" and discussing memory trade-offs. The Desmond-vs-ASE engine mismatch also makes the comparison hard to interpret.\n\nOther concerns are more minor. The elastic moduli statistics exclude outliers (>600 or <-50 GPa); that's worth reporting, but the outlier counts are given, so it's transparent. The MACE-OFF23 torsion comparison uses a different metric (mean barrier height vs RMSD), as the authors note. The inorganic models train to GFN1-xTB charges, and the authors honestly acknowledge the resulting bias in the dielectric tensor. That's an external approximate target, not a circularity.\n\nAll told, this is a solid, honest engineering-plus-benchmarks paper with one load-bearing gap in its central efficiency claim. It deserves peer review--the architecture, datasets, and response-property formalism are valuable even if the speedup needs to be substantiated. I'd cite the Torsion2000 and DLPNO-CCSD(T) benchmarks in my own work, and I'd bring it to a reading group focused on MLFF charge models.","headline":"A broad, honest benchmark paper for a charge-aware MLFF; the architecture and new datasets are real contributions, but the headline 5-20x speedup is not actually measured.","tokens_in":669,"tokens_out":861,"would_cite":true,"duration_ms":19584,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By predicting atomic partial charges inside the network, the MPNICE force field reaches near-best accuracy at 5–20x lower inference cost.","keywords":["machine learning force fields","message passing neural networks","atomic partial charges","charge equilibration","long-range electrostatics","molecular dynamics","inorganic crystals","ionization energies"],"falsifier":"Run the inorganic model on a set of roughly 50 polar crystals for which reference values from density-functional perturbation theory are available for the dielectric tensor, Born effective charges, and LO-TO splitting, then compare magnitudes. The paper's NaCl example already shows reduced magnitudes and small off-diagonal artifacts; if systematic underestimation or large errors in crystals with asymmetric unit cells appears broadly, the claim that the charge representation supports reliable electric response properties is false.","tokens_in":30765,"feed_emoji":"⚡","tokens_out":9459,"duration_ms":90977,"temperature":0.7,"pith_summary":"This paper claims that a machine-learned force field can have electrostatics built in without sacrificing speed. The proposed architecture, MPNICE, iteratively predicts atomic partial charges at every message-passing step, so long-range interactions are represented explicitly rather than absorbed into a purely local energy. Across organic and inorganic benchmarks, the authors find accuracy close to the best available models while running 5–20 times faster at inference. The same charge representation also makes charge-dependent properties—vertical ionization energies, dielectric tensors, Born effective charges, and the long-range correction to phonons—computable from the trained model without retraining. If the claim holds, large-scale molecular dynamics of liquids and materials no longer needs to choose between fidelity and tractability.","feed_headline":"Charge-aware force field matches accuracy at 5–20x lower cost","feed_subtitle":"Predicting atomic partial charges inside the network unlocks long-range electrostatics and charge-dependent properties.","key_machinery":"The load-bearing object is the iterative charge-equilibration loop inside the message-passing stack. With the total charge fixed, each atom's charge is $q_i = -(\\chi_{\\mathrm{eff},i} - \\lambda)/J_{ii}$, with $\\lambda$ chosen so that $\\sum_i q_i = Q_{\\mathrm{tot}}$, where $\\chi_{\\mathrm{eff},i}$ is an MLP-predicted electronegativity and $J_{ii}$ a self-interaction term; this is an approximation to the Qeq method. The loop makes charges internal variables rather than fixed inputs, so long-range electrostatics enter the energy while the messages themselves stay within a finite cutoff, and it gives the model a channel through which an external electric field can be applied as a linear perturbation of the electronegativities. That field perturbation is what turns the same trained model into a predictor of dielectric tensors, Born charges, and non-analytic phonon corrections.","core_discovery":"MPNICE is an invariant message-passing potential in which each atom carries a node-feature vector, a local environment expansion (atomic environment vectors), and a running atomic partial charge. After each interaction block, an MLP predicts an effective electronegativity for each atom, and charges are updated through an analytic charge-equilibration step that enforces global charge conservation and makes the predicted charges feed the next block and the final energy. The final energy is a sum of per-atom terms plus an explicit electrostatic energy from the equilibrated point charges and, optionally, a dispersion correction; multiple output heads let one body of shared features serve different reference levels of theory. The authors demonstrate that this design yields rotamer, tautomer, crystal-ranking, liquid-density, and elastic/phonon accuracies at or near the level of the best existing transferable models, while the charge variables give direct access to ionization energies and to response tensors such as dielectric constants and Born charges.","pith_inferences":["The point-charge representation is the main pressure point: the paper's own NaCl example shows dielectric and Born-charge magnitudes running low relative to reference values, with small off-diagonal artifacts in symmetric cells. A natural extension is to give each atom a learned polarizability or higher multipoles, and the architecture's derivative machinery could then be tested against density-fu","If the 5–20x inference advantage persists for larger models and longer cutoffs, the practical simulation envelope shifts: the reported 17,000-atom water box at roughly 0.07 ns/day would scale to multi-nanosecond simulations of electrolytes, interfaces, and amorphous battery coatings on a single mid-range GPU.","The shared-force hybrid trick—fixing the energy scale to one dataset and letting forces carry the other—looks like a general recipe for multi-fidelity training that could be tested on other incompatible pairs of reference theories beyond the organic/inorganic split.","The dependence on semi-empirical charges suggests an explicit experiment: retrain the inorganic model on density-functional-theory-quality partial charges and compare dielectric and Born-charge predictions. If the errors shrink, the architecture's response claims are limited by the charge labels rather than by the model form."],"forward_implications":["Stable, zero-shot liquid simulations: one Organic MPNICE model simulates 62 common solvents with roughly 4% average density error and reproduces water's diffusion coefficient and hydration free energy closely, suggesting transferable models can be dropped into condensed-phase molecular dynamics without task-specific fitting.","Charge-dependent observables become available from a single trained model: vertical ionization potentials and electron affinities for small molecules, and the LO-TO splitting in polar crystals such as NaCl, can be computed directly from the predicted charge representation.","Cheaper inorganic screening: Inorganic MPNICE ranks monoelemental crystals and thousands of metal-organic framework structures with energy mean absolute errors comparable to much larger general models, and can be used to estimate bulk and shear moduli near equilibrium.","Delta learning with a cheap baseline: the delta-learned organic model reaches roughly 0.2 kcal/mol on torsion scans and 0.3 kcal/mol on tautomer energies, indicating that semi-empirical corrections inside this architecture recover much of the cost of high-level reference data.","Multi-task and shared-force training give a usable single model: a hybrid trained on both organic and inorganic data keeps inorganic total energies accurate and predicts qualitatively correct geometries for unseen Pt/Ir organometallic complexes, though with reduced domain-specific peak accuracy."],"supporting_citations":[{"why":"It supplies the symmetry-function representation of local geometry used to build the radial and angular messages.","marker":"[8]"},{"why":"It is the prior charge-aware neural potential whose charge-equilibration approximation is extended, and it provides one of the main baseline comparisons.","marker":"[10]"},{"why":"It is a comparable transferable model supporting charged species, used as a benchmark for rotamer and tautomer accuracy.","marker":"[11]"},{"why":"It provides the trajectory dataset with PBE energies, forces, and stress used to train the inorganic models.","marker":"[20]"},{"why":"It provides the off-equilibrium AIMD subset used to pretrain the inorganic model before subsequent training to the main PBE trajectory dataset.","marker":"[19]"},{"why":"They supply the organic dataset of drug-like molecules and peptides used in training the organic models.","marker":"[41,42]"},{"why":"It supplies the supplementary organic DFT-energy dataset included in organic training and used as a benchmark reference.","marker":"[43]"},{"why":"It defines the semi-empirical target for the delta-learned organic model and the tight-binding charge reference that informs the inorganic charge labels.","marker":"[44]"},{"why":"It provides the dispersion-correction scheme optionally added to total energies for molecular crystals and inorganic frameworks.","marker":"[40]"},{"why":"It supplies the benchmark numbers and elastic-tensor procedure used to compare inorganic model performance on framework energies and moduli.","marker":"[22]"}],"fun_headline_variants":["Charge-aware MLFF: 5-20x faster, long-range accuracy","MPNICE: charge-equilibrated force fields for liquids and solids","5-20x faster MLFF with long-range charge prediction","Charge-aware MPNICE: long-range MLFF at 5-20x speed","Predicting partial charges inside network gives 5-20x speedup"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inorganic models inherit whatever bias the semi-empirical tight-binding charges carry, because the paper trains them to reproduce those charges rather than true or density-functional-theory charges and then uses the charges to build electrostatic energies and response tensors.","fun_headline_variants_meta":{"raw":{"variants":["Charge-aware MLFF: 5-20x faster, long-range accuracy","MPNICE: charge-equilibrated force fields for liquids and solids","5-20x faster MLFF with long-range charge prediction","Charge-aware MPNICE: long-range MLFF at 5-20x speed","Predicting partial charges inside network gives 5-20x speedup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000706,"raw_usage":{"total_tokens":3182,"prompt_tokens":945,"completion_tokens":2237,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":2138}},"tokens_in":561,"tokens_out":2237,"duration_ms":15800,"temperature":1.0,"reasoning_tokens":2138,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:41:52.542144+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the inorganic model on a set of roughly 50 polar crystals for which reference values from density-functional perturbation theory are available for the dielectric tensor, Born effective charges, and LO-TO splitting, then compare magnitudes. The paper's NaCl example already shows reduced magnitudes and small off-diagonal artifacts; if systematic underestimation or large errors in crystals with asymmetric unit cells appears broadly, the claim that the charge representation supports reliable electric response properties is false.","supporting_citations":[{"cited_title":"P.; Chodera, J","cited_arxiv_id":null,"evidence_quote":"It is a comparable transferable model supporting charged species, used as a benchmark for rotamer and tautomer accuracy."}],"review_version":1}