Pith. sign in

REVIEW 3 cited by

Massive Atomic Diversity: a compact universal dataset for atomistic machine learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.19674 v1 pith:JKJ3MVMT submitted 2025-06-24 cond-mat.mtrl-sci physics.chem-phphysics.comp-ph

classification cond-mat.mtrl-sciphysics.chem-phphysics.comp-ph
keywords datasetstructuresmodelsaccurateatomicdatabasesdetailsmaterials
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of machine-learning models for atomic-scale simulations has benefited tremendously from the large databases of materials and molecular properties computed in the past two decades using electronic-structure calculations. More recently, these databases have made it possible to train universal models that aim at making accurate predictions for arbitrary atomic geometries and compositions. The construction of many of these databases was however in itself aimed at materials discovery, and therefore targeted primarily to sample stable, or at least plausible, structures and to make the most accurate predictions for each compound - e.g. adjusting the calculation details to the material at hand. Here we introduce a dataset designed specifically to train machine learning models that can provide reasonable predictions for arbitrary structures, and that therefore follows a different philosophy. Starting from relatively small sets of stable structures, the dataset is built to contain massive atomic diversity (MAD) by aggressively distorting these configurations, with near-complete disregard for the stability of the resulting configurations. The electronic structure details, on the other hand, are chosen to maximize consistency rather than to obtain the most accurate prediction for a given structure, or to minimize computational effort. The MAD dataset we present here, despite containing fewer than 100k structures, has already been shown to enable training universal interatomic potentials that are competitive with models trained on traditional datasets with two to three orders of magnitude more structures. We describe in detail the philosophy and details of the construction of the MAD dataset. We also introduce a low-dimensional structural latent space that allows us to compare it with other popular datasets and that can be used as a general-purpose materials cartography tool.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Score-based diffusion models for accurate crystal-structure inpainting and reconstruction of hydrogen positions

    cond-mat.mtrl-sci 2026-01 conditional novelty 6.0 of 10

    Adapting TD-Paint to crystal diffusion models reconstructs hydrogen positions with a LES success rate above 97%, beating unconditioned diffusion and DFT-based inpainting.

  2. VASP Plugins: Linking the Vienna ab-initio Simulation Package with Python

    cond-mat.mtrl-sci 2026-07 accept novelty 5.5 of 10

    A C++/pybind11 shared-memory plugin layer exposes VASP SCF and ionic data as NumPy arrays so Python can modify structure, forces, local potential, and occupancies in place.

  3. Six Open Questions in Machine-Learned Interatomic Potential Foundation Models

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 2.0 of 10

    This perspective article develops a definition of foundational MLIPs and poses six open questions that the authors believe will define future research in machine-learned interatomic potentials.

Pith tools