Pith. sign in

REVIEW 16 cited by

Matbench Discovery -- A framework to evaluate machine learning crystal stability predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14920 v3 pith:3HOYXYA7 submitted 2023-08-28 cond-mat.mtrl-sci cs.LG

Matbench Discovery -- A framework to evaluate machine learning crystal stability predictions

classification cond-mat.mtrl-sci cs.LG
keywords discoverymaterialsevaluationframeworkrandomstabilitybenchmarkingcgcnn
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

The rapid adoption of machine learning (ML) in domain sciences necessitates best practices and standardized benchmarking for performance evaluation. We present Matbench Discovery, an evaluation framework for ML energy models, applied as pre-filters for high-throughput searches of stable inorganic crystals. This framework addresses the disconnect between thermodynamic stability and formation energy, as well as retrospective vs. prospective benchmarking in materials discovery. We release a Python package to support model submissions and maintain an online leaderboard, offering insights into performance trade-offs. To identify the best-performing ML methodologies for materials discovery, we benchmarked various approaches, including random forests, graph neural networks (GNNs), one-shot predictors, iterative Bayesian optimizers, and universal interatomic potentials (UIP). Our initial results rank models by test set F1 scores for thermodynamic stability prediction: EquiformerV2 + DeNS > Orb > SevenNet > MACE > CHGNet > M3GNet > ALIGNN > MEGNet > CGCNN > CGCNN+P > Wrenformer > BOWSR > Voronoi fingerprint random forest. UIPs emerge as the top performers, achieving F1 scores of 0.57-0.82 and discovery acceleration factors (DAF) of up to 6x on the first 10k stable predictions compared to random selection. We also identify a misalignment between regression metrics and task-relevant classification metrics. Accurate regressors can yield high false-positive rates near the decision boundary at 0 eV/atom above the convex hull. Our results demonstrate UIPs' ability to optimize computational budget allocation for expanding materials databases. However, their limitations remain underexplored in traditional benchmarks. We advocate for task-based evaluation frameworks, as implemented here, to address these limitations and advance ML-guided materials discovery.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

    cond-mat.mtrl-sci 2024-10 conditional novelty 8.0

    OMat24 releases a new open dataset of 110M+ DFT calculations and EquiformerV2 models achieving SOTA on Matbench Discovery with F1>0.9 for stability and 20 meV/atom accuracy for formation energies.

  2. DPA4: Pushing the Accuracy-Cost Frontier of Interatomic Potentials with EMFA SO(2) Convolution

    physics.chem-ph 2026-06 unverdicted novelty 7.0

    DPA4 is a new SE(3)-equivariant interatomic potential with EMFA SO(2) convolution that sets new accuracy-cost records on Matbench Discovery and SPICE benchmarks using fewer parameters than prior models.

  3. JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

    cs.DC 2026-05 unverdicted novelty 7.0

    JanusPipe introduces SymFold and WaveK to enable efficient 3D-parallel training for conservative MLIPs, reporting 1.51x and 1.45x average throughput gains over 1F1B and Hanayo baselines on 32 GPUs.

  4. Pushing the limits of unconstrained machine-learned interatomic potentials

    physics.chem-ph 2026-01 conditional novelty 7.0

    Unconstrained non-equivariant and direct-force neural interatomic potentials scale to 730M parameters and match or beat equivariant state-of-the-art models on several atomistic benchmarks.

  5. MatMind: A Structure-Activity Knowledge-Driven Generative Foundation Model for Materials Science

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 6.0

    MatMind is a unified LLM-based generative model for crystals that reports lowest MAE on energy above hull, bulk modulus and band gap while achieving 65.3% S.U.N. rate on unconditional generation.

  6. MatFormBench: A Benchmarking Evaluation Framework for Target-Driven Materials Formulation

    cond-mat.mtrl-sci 2026-05 unverdicted novelty 6.0

    MatFormBench introduces a synthetic data generator with five difficulty levels and MatFormScore metric to benchmark 39 inverse design algorithms for target-driven materials formulation.

  7. JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials

    cs.DC 2026-05 unverdicted novelty 6.0

    JanusPipe is a new 3D-parallel training system for conservative MLIPs that uses SymFold and WaveK to achieve 1.51x and 1.45x average throughput gains over 1F1B and Hanayo on 32 GPUs.

  8. Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning

    cs.LG 2026-05 unverdicted novelty 6.0

    Structural pruning of SO(3) equivariant atomistic models from large checkpoints yields 1.5-4x fewer parameters and 2.5-4x less pre-training compute than small models trained from scratch, while outperforming them on m...

  9. From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

    cs.LG 2026-02 conditional novelty 6.0

    A bond-deformation benchmark plus a force-smoothness metric is proposed to detect PES artifacts and guide MLIP architecture design, with improvements shown on a new Transformer-style model.

  10. AI-Driven Expansion and Application of the Alexandria Database

    cond-mat.mtrl-sci 2025-12 accept novelty 6.0

    A combined generative model, ML potential, and graph neural network pipeline expands the Alexandria database by 1.3 million DFT-validated compounds with 99% success near the convex hull and releases training data for ...

  11. MiAD: Mirage Atom Diffusion for De Novo Crystal Generation

    cs.LG 2025-11 conditional novelty 6.0

    Mirage infusion lets crystal diffusion models vary atom counts during generation and raises the S.U.N. rate on MP-20 to 8.2%.

  12. OpenCSP: A Deep Learning Framework for Crystal Structure Prediction from Ambient to High Pressure

    cond-mat.mtrl-sci 2025-09 conditional novelty 6.0

    OpenCSP is an open pressure-diverse dataset and model suite that matches or beats larger universal atomistic models on high-pressure crystal structure prediction with far fewer training data.

  13. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

    cond-mat.mtrl-sci 2024-10 accept novelty 6.0

    OMat24 provides over 110 million DFT calculations and EquiformerV2 models that reach state-of-the-art performance on material stability and formation energy prediction.

  14. MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures

    cond-mat.mtrl-sci 2024-05 unverdicted novelty 6.0

    MatterSim delivers a single deep learning force field that simulates inorganic materials across elements, 0-5000 K, and up to 1000 GPa with near first-principles accuracy for lattice dynamics, mechanics, and Gibbs fre...

  15. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 conditional novelty 5.0

    An agentic HPC skill automates NEB microkinetics, recovers from common failures, and benchmarks ~12 universal MLIPs against DFT for CO2 sublimation on graphite.

  16. Toward Exascale AI for Science: A Scalable AI Skill for Autonomous Microkinetics Discovery

    cs.CE 2026-06 unverdicted novelty 3.0

    Introduces a scalable AI skill framework for autonomous microkinetics discovery that automates workflows and evaluates surrogate reliability.