Pith. sign in

REVIEW 3 cited by

The Importance of Being Scalable: Improving the Speed and Accuracy of Neural Network Interatomic Potentials Across Chemical Domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.24169 v1 pith:JZZHSL4Q submitted 2024-10-31 cs.LG

classification cs.LG
keywords modelscalingnnipsattentionescaipperformanceefficientlyinteratomic
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Scaling has been critical in improving model performance and generalization in machine learning. It involves how a model's performance changes with increases in model size or input data, as well as how efficiently computational resources are utilized to support this growth. Despite successes in other areas, the study of scaling in Neural Network Interatomic Potentials (NNIPs) remains limited. NNIPs act as surrogate models for ab initio quantum mechanical calculations. The dominant paradigm here is to incorporate many physical domain constraints into the model, such as rotational equivariance. We contend that these complex constraints inhibit the scaling ability of NNIPs, and are likely to lead to performance plateaus in the long run. In this work, we take an alternative approach and start by systematically studying NNIP scaling strategies. Our findings indicate that scaling the model through attention mechanisms is efficient and improves model expressivity. These insights motivate us to develop an NNIP architecture designed for scalability: the Efficiently Scaled Attention Interatomic Potential (EScAIP). EScAIP leverages a multi-head self-attention formulation within graph neural networks, applying attention at the neighbor-level representations. Implemented with highly-optimized attention GPU kernels, EScAIP achieves substantial gains in efficiency--at least 10x faster inference, 5x less memory usage--compared to existing NNIPs. EScAIP also achieves state-of-the-art performance on a wide range of datasets including catalysts (OC20 and OC22), molecules (SPICE), and materials (MPTrj). We emphasize that our approach should be thought of as a philosophy rather than a specific model, representing a proof-of-concept for developing general-purpose NNIPs that achieve better expressivity through scaling, and continue to scale efficiently with increased computational resources and training data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BoostMD: Accelerating molecular sampling by leveraging ML force field features from previous time-steps

    physics.chem-ph 2024-12 conditional novelty 7.0 of 10

    BoostMD accelerates MLFF molecular dynamics by predicting energy changes from previous-step node features and positional displacements, reporting 8x speedup and matching the reference model's sampled free energy surfa...

  2. Platonic Transformers: A Solid Choice For Equivariance

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Platonic Transformers achieve exact equivariance to translations plus discrete Platonic-solid rotations by lifting features into multiple reference frames and sharing one RoPE attention across them, with a linear-time...

  3. Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials

    cond-mat.mtrl-sci 2025-02 conditional novelty 4.0 of 10

    A perspective arguing that current MLIP training on DFT data is insufficient, and proposing CCSD(T)-quality data, metrology, and efficient inference as research priorities.

Pith tools