SLayerGen generates crystals invariant to any space or layer group via autoregressive lattice and Wyckoff sampling plus equivariant diffusion, achieving gains over bulk models on diperiodic materials after correcting a prior loss inconsistency for hexagonal groups.
hub Mixed citations
Orb-v3: atomistic simulation at scale
Mixed citation behavior. Most common role is background (67%).
abstract
We introduce Orb-v3, the next generation of the Orb family of universal interatomic potentials. Models in this family expand the performance-speed-memory Pareto frontier, offering near SoTA performance across a range of evaluations with a >10x reduction in latency and > 8x reduction in memory. Our experiments systematically traverse this frontier, charting the trade-off induced by roto-equivariance, conservatism and graph sparsity. Contrary to recent literature, we find that non-equivariant, non-conservative architectures can accurately model physical properties, including those which require higher-order derivatives of the potential energy surface. This model release is guided by the principle that the most valuable foundation models for atomic simulation will excel on all fronts: accuracy, latency and system size scalability. The reward for doing so is a new era of computational chemistry driven by high-throughput and mesoscale all-atom simulations.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
Curvature revives suppressed charge density wave order localized to non-Euclidean regions in a TiSe2-NbSe2 heterostructure despite global Fermi surface disruption.
LSD extends speculative sampling to second-order Langevin dynamics, achieving 3-9x speedup in MD while exactly sampling from the target distribution without relative error.
Learned functional perturbations plus CRPS training convert deterministic ML interatomic potentials into probabilistic ones, improving CRPS by 19-32% on N-body benchmarks and uncertainty-error correlation from 0.75 to 0.84 on silica.
Lang2MLIP is an LLM multi-agent framework that automates end-to-end development of machine learning interatomic potentials from natural language input for heterogeneous materials systems.
Kernels from pretrained MLIP latent spaces outperform standard acquisition methods in active learning for reactive chemistry, reducing required labels by 38% for energy error and 28% for force error.
Ca extraction from post-spinel CaV2O4 is kinetically limited to at most half theoretical capacity at 298 K by high Ca migration barriers in the ordered γ phase and a persistent γ–δ two-phase region.
Large MLIPs deliver marginal accuracy gains over lightweight models while sacrificing orders of magnitude in throughput and scalability, making lightweight models the practical Pareto-optimal choice.
MLIP Studio unifies 60+ universal machine learning interatomic potentials in a free web platform for interactive atomistic simulations, benchmarking, and MLIP-accelerated DFT workflows.
Enerzyme framework trains electrostatics-aware NNPs on under 1,000 system-specific points to reproduce MTase reaction energetics and transition states for clusters up to 545 atoms.
High-throughput ML interatomic potentials plus phonon calculations discover and experimentally confirm CsTlI4 as having record-low intrinsic room-temperature thermal conductivity of 0.14 W m^{-1} K^{-1}.
High-throughput screening combining Voronoi polyhedral volumes and foundational ML models identifies 37 promising Ca cathode candidates from the Materials Project database.
A reweighting method with mean energy-gap approximation transfers PMFs between MLIPs to recover target reaction and activation free energies at low cost for a 601-atom Li+ transport system across DFT levels.
Force-aware Neural Tangent Kernels combined with chunked acquisition provide scalable and distribution-robust active learning for MLIPs, outperforming baselines on OC20 and remaining competitive on other benchmarks.
Machine learning models, especially certain deep neural networks, can predict lattice thermal conductivity with useful accuracy across different generalization tests while being orders of magnitude faster than first-principles calculations.
CrystalREPA closes the representation gap between crystal generators and universal MLIPs via contrastive alignment, yielding more stable and valid generated crystals while revealing that MLIP teacher quality is better predicted by representation distinguishability than by leaderboard accuracy.
Structural pruning of SO(3) equivariant atomistic models from large checkpoints yields 1.5-4x fewer parameters and 2.5-4x less pre-training compute than small models trained from scratch, while outperforming them on most Matbench Discovery metrics and downstream tasks.
Torched-TACAW plus ORB MD and z-partitioned supercells enables near-ab-initio-quality atomic-resolution vibrational STEM-EELS for thick TiO2 models with tractable memory and data flow.
Universal MLIPs serve as configuration generators whose DFT-relabeled subsamples enable one-shot or iterative training of material-specific MLIPs that recover accurate reactive energy profiles with 600-2000 DFT calculations.
PRISMat generates crystal slabs with mean absolute errors of 0.188 eV/A² for cleavage energy and 2.79 eV for work function, reducing error by 4× versus the next best model while using less inference time.
SevenNet-Nano is a lightweight universal ML interatomic potential distilled from a larger multi-task foundation model, delivering high accuracy, transferability, and over 10x computational speedup for scalable atomistic simulations.
Benchmarks of 15 MLIPs show parameter count and training set size correlate with accuracy, architecture drives speed and memory, and explicit Coulomb terms provide no benefit.
Different uMLIPs encode chemical space in distinct ways, with high cross-model feature reconstruction errors, and fine-tuning preserves strong pre-training bias in the latent features.
A hierarchical screening protocol using PBE phase diagrams, ML interatomic potentials, and SCAN refinement reduces 894 computationally stable materials to 25 high-confidence experimental synthesis targets.
citing papers explorer
-
SLayerGen: a Crystal Generative Model for all Space and Layer Groups
SLayerGen generates crystals invariant to any space or layer group via autoregressive lattice and Wyckoff sampling plus equivariant diffusion, achieving gains over bulk models on diperiodic materials after correcting a prior loss inconsistency for hexagonal groups.
-
Curvature-driven revival of charge density waves in non-Euclidean space
Curvature revives suppressed charge density wave order localized to non-Euclidean regions in a TiSe2-NbSe2 heterostructure despite global Fermi surface disruption.
-
Speculative Sampling For Faster Molecular Dynamics
LSD extends speculative sampling to second-order Langevin dynamics, achieving 3-9x speedup in MD while exactly sampling from the target distribution without relative error.
-
Uncertainty-aware Machine Learning Interatomic Potentials via Learned Functional Perturbations
Learned functional perturbations plus CRPS training convert deterministic ML interatomic potentials into probabilistic ones, improving CRPS by 19-32% on N-body benchmarks and uncertainty-error correlation from 0.75 to 0.84 on silica.
-
Lang2MLIP: End-to-End Language-to-Machine Learning Interatomic Potential Development with Autonomous Agentic Workflows
Lang2MLIP is an LLM multi-agent framework that automates end-to-end development of machine learning interatomic potentials from natural language input for heterogeneous materials systems.
-
Pretrained Model Representations as Acquisition Signals for Active Learning of MLIPs
Kernels from pretrained MLIP latent spaces outperform standard acquisition methods in active learning for reactive chemistry, reducing required labels by 38% for energy error and 28% for force error.
-
Phase stability and ionic transport in post-spinel CaV$_2$O$_4$ cathode
Ca extraction from post-spinel CaV2O4 is kinetically limited to at most half theoretical capacity at 298 K by high Ca migration barriers in the ordered γ phase and a persistent γ–δ two-phase region.
-
Are Machine Learning Interatomic Potentials Truly Practical? A Benchmark of 23 Mainstream Models
Large MLIPs deliver marginal accuracy gains over lightweight models while sacrificing orders of magnitude in throughput and scalability, making lightweight models the practical Pareto-optimal choice.
-
MLIP Studio: An Open Platform for Interactive Benchmarking and Atomistic Simulations Using Machine Learning Interatomic Potentials
MLIP Studio unifies 60+ universal machine learning interatomic potentials in a free web platform for interactive atomistic simulations, benchmarking, and MLIP-accelerated DFT workflows.
-
Enerzyme: A Framework for Efficient Training of Reactive Neural Network Potentials for Enzyme Catalysis with Application to Methyltransferases
Enerzyme framework trains electrostatics-aware NNPs on under 1,000 system-specific points to reproduce MTase reaction energetics and transition states for clusters up to 545 atoms.
-
Approaching the Limit of Intrinsic Crystalline Thermal Insulation
High-throughput ML interatomic potentials plus phonon calculations discover and experimentally confirm CsTlI4 as having record-low intrinsic room-temperature thermal conductivity of 0.14 W m^{-1} K^{-1}.
-
Geometry-based Discovery of Calcium Battery Cathodes Accelerated by Foundational Machine-Learned Models
High-throughput screening combining Voronoi polyhedral volumes and foundational ML models identifies 37 promising Ca cathode candidates from the Materials Project database.
-
Reweighting free energy profiles between universal machine learning interatomic potentials for fast consensus building
A reweighting method with mean energy-gap approximation transfers PMFs between MLIPs to recover target reaction and activation free energies at low cost for a 601-atom Li+ transport system across DFT levels.
-
Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs
Force-aware Neural Tangent Kernels combined with chunked acquisition provide scalable and distribution-robust active learning for MLIPs, outperforming baselines on OC20 and remaining competitive on other benchmarks.
-
Fast and Accurate Prediction of Lattice Thermal Conductivity via Machine Learning Surrogates
Machine learning models, especially certain deep neural networks, can predict lattice thermal conductivity with useful accuracy across different generalization tests while being orders of magnitude faster than first-principles calculations.
-
CrystalREPA: Transferring Physical Priors from Universal MLIPs to Crystal Generative Models
CrystalREPA closes the representation gap between crystal generators and universal MLIPs via contrastive alignment, yielding more stable and valid generated crystals while revealing that MLIP teacher quality is better predicted by representation distinguishability than by leaderboard accuracy.
-
Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning
Structural pruning of SO(3) equivariant atomistic models from large checkpoints yields 1.5-4x fewer parameters and 2.5-4x less pre-training compute than small models trained from scratch, while outperforming them on most Matbench Discovery metrics and downstream tasks.
-
Efficient Large-Scale STEM-EELS Simulations With Torched-TACAW
Torched-TACAW plus ORB MD and z-partitioned supercells enables near-ab-initio-quality atomic-resolution vibrational STEM-EELS for thick TiO2 models with tractable memory and data flow.
-
Universal Interatomic Potentials as Configuration-Space Generators for One-Shot and Iterative Fine-Tuning of Ab Initio-Accurate Material-Specific Models
Universal MLIPs serve as configuration generators whose DFT-relabeled subsamples enable one-shot or iterative training of material-specific MLIPs that recover accurate reactive energy profiles with 600-2000 DFT calculations.
-
PRISMat: Policy-Driven, Permutation-Invariant Autoregressive Material Generation
PRISMat generates crystal slabs with mean absolute errors of 0.188 eV/A² for cleavage energy and 2.79 eV for work function, reducing error by 4× versus the next best model while using less inference time.
-
A Lightweight Universal Machine-Learning Interatomic Potential via Knowledge Distillation for Scalable Atomistic Simulations
SevenNet-Nano is a lightweight universal ML interatomic potential distilled from a larger multi-task foundation model, delivering high accuracy, transferability, and over 10x computational speedup for scalable atomistic simulations.
-
Accuracy and Efficiency Benchmarks of Pretrained Machine Learning Potentials for Molecular Simulations
Benchmarks of 15 MLIPs show parameter count and training set size correlate with accuracy, architecture drives speed and memory, and explicit Coulomb terms provide no benefit.
-
Comparing the latent features of universal machine-learning interatomic potentials
Different uMLIPs encode chemical space in distinct ways, with high cross-model feature reconstruction errors, and fine-tuning preserves strong pre-training bias in the latent features.
-
Predicting Novel Stable Materials for Experimental Synthesis
A hierarchical screening protocol using PBE phase diagrams, ML interatomic potentials, and SCAN refinement reduces 894 computationally stable materials to 25 high-confidence experimental synthesis targets.
-
Additive binding energies in asphalt on a quantum processor via quantum-selected configuration interaction (QSCI)
Hybrid quantum workflow on IQM Emerald processor computes -3.52 kcal/mol binding energy for pyridine-phenol complex via QSCI in (10e,10o) space, matching CASCI but underbinding relative to CCSD(T) benchmark of -8.5 to -9.5 kcal/mol.
-
Six Open Questions in Machine-Learned Interatomic Potential Foundation Models
This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.
-
From Knowledge to Action: Outcomes of the 2025 Large Language Model (LLM) Hackathon for Applications in Materials Science and Chemistry
Hackathon submissions indicate LLMs are moving from general assistants toward composable multi-agent systems for structuring scientific knowledge and automating tasks in materials science and chemistry.