Pith. sign in

REVIEW 19 cited by

Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.07237 v2 pith:H6XEB5T2 submitted 2022-10-13 physics.comp-ph cs.LGphysics.chem-ph

Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations

classification physics.comp-ph cs.LGphysics.chem-ph
keywords benchmarkforcesimulationmodelsbenchmarkedevaluationforceslearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for learned MD simulation. We curate representative MD systems, including water, organic molecules, a peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate future work.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

    cond-mat.mtrl-sci 2024-10 conditional novelty 8.0

    OMat24 releases a new open dataset of 110M+ DFT calculations and EquiformerV2 models achieving SOTA on Matbench Discovery with F1>0.9 for stability and 20 meV/atom accuracy for formation energies.

  2. Geometric Algebra Meets Cartesian Tensors: Higher-Order Equivariance for Interatomic Potentials

    physics.chem-ph 2026-06 unverdicted novelty 7.0

    CliffordSTF couples Clifford multivectors to rank-2 and rank-3 symmetric-traceless tensor tracks through bilinear cross-track contractions, lifting force cosine similarity from 0.055 to 0.551 on rMD17 while outperform...

  3. Girsanov Reweighting for Uncertainty Propagation in Rare-Event Kinetics

    physics.chem-ph 2026-07 conditional novelty 6.0

    Girsanov reweighting of AMS-sampled reactive trajectories propagates machine-learned interatomic potential parameter uncertainty to committor probabilities and, under extra assumptions, to reaction rates.

  4. Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials

    cs.LG 2026-07 unverdicted novelty 6.0

    SOAP and SOAP-Muon optimizers deliver faster convergence and higher final accuracy than Adam for NequIP and Allegro MLIPs, with the largest gains under partial force supervision.

  5. From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

    cs.LG 2026-02 conditional novelty 6.0

    A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.

  6. Knowledge Distillation of a Protein Language Model Yields a Foundational Implicit Solvent Model

    physics.bio-ph 2026-01 conditional novelty 6.0

    A 45,000-parameter GNN distilled from a protein language model's secondary-structure predictions can drive MD simulations and, combined with GBn2 electrostatics, approximately reproduces explicit-solvent folding profi...

  7. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

    cond-mat.mtrl-sci 2024-10 accept novelty 6.0

    OMat24 provides over 110 million DFT calculations and EquiformerV2 models that reach state-of-the-art performance on material stability and formation energy prediction.

  8. MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures

    cond-mat.mtrl-sci 2024-05 unverdicted novelty 6.0

    MatterSim delivers a single deep learning force field that simulates inorganic materials across elements, 0-5000 K, and up to 1000 GPa with near first-principles accuracy for lattice dynamics, mechanics, and Gibbs fre...

  9. A general-purpose atomic cluster expansion interatomic potential for niobium

    cond-mat.mtrl-sci 2026-07 unverdicted novelty 5.0

    A new general-purpose ACE interatomic potential for niobium is trained on diverse DFT data and validated on phonons, high-pressure behavior, dislocation barriers, and a near-million-atom fracture simulation.

  10. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 conditional novelty 5.0

    An agentic HPC skill automates NEB microkinetics, recovers from common failures, and benchmarks ~12 universal MLIPs against DFT for CO2 sublimation on graphite.

  11. Universal Interatomic Potentials as Configuration-Space Generators for One-Shot and Iterative Fine-Tuning of Ab Initio-Accurate Material-Specific Models

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 5.0

    Universal MLIPs serve as configuration generators whose DFT-relabeled subsamples enable one-shot or iterative training of material-specific MLIPs that recover accurate reactive energy profiles with 600-2000 DFT calculations.

  12. Dynamical properties of ab initio water from machine-learning potentials

    cond-mat.soft 2026-06 unverdicted novelty 5.0

    Machine-learning interatomic potentials trained on prior ab initio data for multiple DFT functionals are used to compare water dynamical properties, with RPBE-D3/zd identified as best matching experiment and further v...

  13. Harnessing AtomisticSkills for Agentic Atomistic Research

    physics.chem-ph 2026-05 unverdicted novelty 5.0

    AtomisticSkills is a new harness framework with 100+ human-curated skills that lets general AI agents perform atomistic research tasks including simulations, screening, and analysis, shown on electrolyte design, CO2 c...

  14. Benchmarking empirical and machine-learned interatomic potentials using phase diagram predictions for Lead

    cond-mat.mtrl-sci 2026-05 unverdicted novelty 5.0

    The EDDP machine-learned potential for lead predicts the observed FCC-HCP phase transition at ~15 GPa, unlike EAM and MEAM models, when paired with nested sampling.

  15. NEPMaker: Active learning of neuroevolution machine learning potential for large cells

    physics.comp-ph 2026-04 unverdicted novelty 5.0

    NEPMaker uses D-optimality active learning to identify and locally embed extrapolative atomic environments from large simulations into periodic structures for training neuroevolution potentials, aiming to cut extrapol...

  16. Differentiable hybrid force fields support scalable autonomous electrolyte discovery

    cond-mat.mtrl-sci 2026-04 unverdicted novelty 4.0

    Differentiable hybrid force fields combine physical models with neural corrections to enable fast, accurate, and calibratable simulations for scalable autonomous electrolyte discovery.

  17. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 unverdicted novelty 3.0

    Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.

  18. Deep Learning for Protein Complex Prediction and Design

    cs.LG 2026-05 unverdicted novelty 3.0

    Deep learning architectures tailored to protein hierarchy combined with sequence-space search algorithms are used to improve prediction of protein complex structures and to design new interacting sequences.

  19. Differentiable hybrid force fields support scalable autonomous electrolyte discovery

    cond-mat.mtrl-sci 2026-04 unverdicted novelty 3.0

    Differentiable hybrid force fields fuse physical skeletons with neural corrections to deliver fast, accurate, and top-down calibratable simulation for autonomous electrolyte discovery.