Pith. sign in

REVIEW 4 major objections 4 minor 93 references

A delta-learning model built on the latent features of a foundational interatomic potential corrects DFT formation energies to within 49 meV/atom of experiment, matching experimental uncertainty.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 16:00 UTC pith:DMWCYML7

load-bearing objection Solid, transparent benchmark of latent-feature delta-learning with PET-OMATPES, hitting <50 meV/atom on a filtered experimental set; the main caveat is evaluation scope. the 4 major comments →

arxiv 2607.18092 v1 pith:DMWCYML7 submitted 2026-07-20 cond-mat.mtrl-sci

Correcting DFT formation energies towards experimental accuracy using foundational MLIPs and latent-feature delta-learning

classification cond-mat.mtrl-sci
keywords formation energiesdelta learningfoundational MLIPlatent featuresr2SCANGGA correctionsthermodynamic stabilitykernel ridge regression
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to show that formation energies in GGA-level crystal-structure databases can be pushed to near-experimental accuracy without any new density-functional calculations. Its recipe is two-stage: evaluate a foundational machine-learning interatomic potential trained on r2SCAN reference data at existing GGA-relaxed geometries (this alone cuts the mean absolute error versus experiment by more than 40%), then train a small kernel-ridge model on the potential's latent structural features to learn and subtract the remaining residual. On the merged experimental reference set, the best model reaches a mean absolute error below 50 meV/atom, which the authors equate with experimental uncertainty itself. The paper also introduces formation-energy values for its underlying database and shows that differences between public DFT databases are dominated by their empirical correction schemes rather than by DFT codes or settings. A sympathetic reader would care because formation energies and convex-hull stability are the standard filter in computational materials discovery, and this recipe promises to upgrade large existing databases at almost no cost.

Core claim

On the authors' own terms, the central discovery is that the latent-space representation of a foundational machine-learning interatomic potential is a better feature set for delta-learning DFT-to-experiment residuals than composition-only descriptors. Kernel ridge regression with a Laplacian kernel trained on those latent features predicts the residual between the potential's zero-shot r2SCAN formation energies and experimental formation enthalpies with a held-out MAE of 49 meV/atom, and the same approach applied to the directly corrected GGA target reaches 52 meV/atom. The structure-aware latent features lower the MAE and, once regularization is tuned, lower the stability flip rate compared

What carries the argument

The load-bearing mechanism is delta-learning on the backbone latent features of the PET-OMATPES foundational machine-learning interatomic potential, a representation extracted after the message-passing layers and before the readout layers, which the paper finds to be richer than the final-layer features. A kernel ridge regression model with a tuned Laplacian kernel learns the difference between the potential's zero-shot r2SCAN formation energy and the experimental formation enthalpy, and that learned correction is added on top. Regularization strength plays a double role: it controls both overfitting of the small reference set and the distortion of relative phase stability, letting the model

Load-bearing premise

The load-bearing premise is that the filtered set of experimental formation enthalpies used for training and testing (2,356 compounds after filtering, 1,384 after merging with the studied database) is a fair and unbiased sample of a GGA database's contents, so the sub-50 meV/atom error measured on that set transfers to new materials; if the reference set skews toward well-measured, easier compounds, the claim does not generalize.

What would settle it

Collect new experimental formation enthalpies for chemistries and structure families underrepresented in the current reference set (for example, chalcogenides beyond oxides or compounds with large DFT-experiment disagreements) and evaluate the trained delta model on them. An MAE clearly above 50 to 75 meV/atom, or a stability flip rate outside the 3 to 4 percent band, would show that the corrections do not transfer beyond the reference distribution.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Existing GGA databases can have their formation energies brought to about 50 meV/atom agreement with experiment without any additional DFT, making their stability filters materially more reliable.
  • Pairing PBEsol-relaxed geometries with r2SCAN-level MLIP energies is almost as accurate as fully relaxing with the MLIP, extending a known DFT practice to foundational potentials.
  • Structure-aware latent features outperform composition-only features for delta corrections, lowering both the formation-energy MAE and the stability flip rate.
  • With tuned regularization, the learned correction keeps stability flip rates at 3 to 4 percent at typical energy windows above the convex hull, with most flips inherited from the physically grounded MLIP rather than introduced by the correction model.
  • The ML correction improves more compounds than fitted elemental-reference corrections and degrades the compounds it misses less severely.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the error transfers beyond the curated reference distribution, the same two-step recipe could re-score entire GGA databases into a meta-GGA-quality stability layer for the cost of one MLIP evaluation and a small kernel model.
  • The structure-aware latent features are not limited to formation energies; the same delta-learning construction could be pointed at other experimental targets such as band gaps, magnetic ordering temperatures, or adsorption energies wherever a labelled reference set exists.
  • The corrections cluster into a few discrete modes across train-test splits, which suggests that a committee or majority-vote ensemble of delta models would be a more stable production choice than any single split; this is a testable extension implied by the paper's variance analysis.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces formation energies for the Materials Cloud three-dimensional crystals database (MC3D), compares them with OQMD and Materials Project, and proposes a two-step correction towards experimental formation enthalpies: first, zero-shot r2SCAN energies from the PET-OMATPES foundational MLIP, and second, learned delta corrections using classical ML models (RF, GPR, KRR) with compositional (Magpie) or MLIP latent features. On a filtered experimental reference set of 1384 PBEsol compounds, the best model (KRR-LAP with latent features) reaches a test MAE of 49 meV/atom for the r2SCAN target, down from ~160 meV/atom for uncorrected GGA, and the paper analyses stability flip rates and regularization trade-offs. The authors also validate the MLIP zero-shot energies against the Alexandria database.

Significance. If the <50 meV/atom accuracy transfers beyond the filtered experimental subset, the method provides a practical route to correct existing GGA databases to near-experimental accuracy without additional DFT. The paper is transparent about its evaluation: 30 train-test splits, 5-fold CV, two feature sets, four model classes, and a comparison with FERE. The introduction of MC3D formation energies is a useful resource. However, the central claim rests on a small, filtered, experimentally characterized set with random splits, and the FERE comparison is confounded by in-sample fitting; these issues need to be resolved before the general claim is supported.

major comments (4)
  1. [Methods, 'Reference data a'; 'Machine learning models and training'] The evaluation protocol uses 30 random 80/20 splits stratified only by element count on the 1384-compound filtered experimental subset. This subset passes four exclusion filters (elemental phases; uncertainty >10%; DFT–experiment disagreement >0.5 eV/atom; cross-source disagreement >150 meV/atom). Random splits keep test structures close to training structures, so the reported test MAE is an in-distribution estimate. The central claim of the paper—that GGA/MC3D formation energies can be corrected to <50 meV/atom—is used to motivate correcting the full MC3D database, for which most entries lack experimental references. The paper does not provide a composition-based or chemical-family hold-out, nor a validation on compounds excluded by the filters. The SI (§S4.D) itself concedes that 'few residual structure-specific artifacts persist depending on which structures appear in the training set
  2. [Comparing with existing approaches; SI §S5] The FERE corrections are fitted on the full dataset (1511 compounds, MC3D PBE version; Table S2) and then evaluated on the same test splits used for ML, which are disjoint from the ML training data. Thus the FERE results in Fig. 6 and Figs. S15–S19 are in-sample (or at least use the test labels for fitting), while the ML results are out-of-sample. Moreover, the FERE corrections are fitted on PBE data but applied to the PBEsol baseline in the comparison. This confounds the claimed superiority of the ML approach. Refit FERE within each training split (or a nested CV) and on the same functional/target as the ML models before comparing.
  3. [Balance accurate formation energies...; Methods] The computation of energies above the convex hull and stability flip rates is not described. It is unclear whether the hull is constructed from all MC3D structures or only from the 1384-compound experimental subset. If only the subset is used, the hull is incomplete and the flip rates (Fig. 3b, 4, 5) may be artifacts of a truncated hull. Specify the hull construction, including how corrected formation energies of non-experimental MC3D entries are handled.
  4. [Machine learning models and training; Fig. 4] The final KRR-LAP-LF model uses α=0.1 selected based on the test-set trade-off in Fig. 4a. This is a post-selection choice; the reported MAE of 49 meV/atom is therefore not a purely out-of-sample estimate. Use a validation set for α selection or report the MAE across all α values and discuss the selection bias.
minor comments (4)
  1. [Abstract] The phrase 'reducing the mean absolute error by more than 40% relative to GGA' should specify which comparison (zero-shot vs. pure DFT) and which GGA (PBE/PBEsol), since the 40% reduction refers to the zero-shot step and the later ML correction is additional.
  2. [Discussion / Fig. 3a] The claim that the MAE is 'comparable to the experimental uncertainty itself' is somewhat overstated: the achieved 49 meV/atom is above the commonly cited ~25 meV/atom typical uncertainty, although within the 70 meV/atom range across databases. Qualify the statement accordingly.
  3. [Fig. 6] The main-text figure uses the test split with the best ML MAE; the worst split is only in the SI. Label the figure as 'best split' or show both splits in the main text to avoid cherry-picking concerns.
  4. [Data/Code Availability] The statement says 'will be made publicly available upon publication'. For a fully reproducible claim, provide a repository link or temporary access during review.

Circularity Check

0 steps flagged

No significant circularity: the <50 meV/atom claim rests on supervised delta-learning against external experimental labels with held-out test splits, and the MLIP baseline is an externally trained model.

full rationale

The central derivation chain is an empirical benchmark, not a closed analytic loop. The delta-learning models are trained to predict the residual between DFT or MLIP formation energies and experimental formation enthalpies, using either Magpie compositional features or PET-OMATPES latent features as inputs. The experimental values enter only as training labels, never as features, and performance is reported on 80/20 train-test splits repeated over 30 seeds, so the reported test MAE is not the training objective evaluated on the training set. The zero-shot improvement comes from PET-OMATPES, a foundational MLIP trained externally on OMAT and MATPES r2SCAN data, not fitted to the experimental enthalpies used here; therefore the >40% MAE reduction is an independent transfer result rather than a fitted parameter renamed as a prediction. The MC3D formation energies themselves are produced by Quantum ESPRESSO DFT calculations, and the FERE corrections are used only as a comparison baseline, not as the source of the headline ML claim. The paper does cite earlier work by the same group (e.g., MC3D [17], Materials Cloud [18], and the computational-parameter reference [80]), but these citations identify data and infrastructure; they are not invoked as a uniqueness theorem, an ansatz, or the justification for the fitted corrections. The SI honestly notes residual structure-specific artifacts across data splits and the need for larger experimental reference datasets (SI S4.D); this is a generalizability and dataset-size limitation concerning the 1384-compound experimental reference set, not a circularity. The random stratified splits may overstate transfer to the full MC3D database, but that is a correctness/robustness concern distinct from circular reasoning. In sum, no load-bearing step reduces by construction to its own inputs.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 0 invented entities

The headline numbers rest on experimental reference data, a pretrained MLIP, and fitted correction models. The only parameters fitted inside this paper are the ML hyperparameters and FERE corrections; the MLIP weights come from prior work.

free parameters (5)
  • Delta-learning model parameters (KRR/GPR/RF) = trained on ~1100 compounds (80% of 1384)
    The central <50 meV/atom result is the test performance of these fitted models; it is not a parameter-free prediction.
  • KRR regularization alpha = 0.1
    Selected by cross-validation as the best MAE/flip-rate trade-off; the headline model is KRR-LAP-LF at alpha=0.1.
  • Kernel and model hyperparameters (RBF/LAP widths, GPR/RF settings) = not tabulated
    Chosen by 5-fold CV per split; not fully listed in the paper, so reproduction requires unpublished choices.
  • FERE-part elemental corrections (13 elements) = Table S2 (e.g., O -0.357 eV/atom)
    Fitted by least squares to experimental formation energies; used as a baseline comparison, not the central claim.
  • FERE-all elemental corrections = Table S2
    Same fit over all elements; used to compare degradation behavior.
axioms (5)
  • domain assumption DFT formation energy is a 0 K total-energy difference; finite-temperature corrections are negligible.
    Methods 'DFT formation energies'; standard but enters the comparison to room-temperature experimental enthalpies.
  • domain assumption The filtered experimental formation-enthalpy set (Wang/Kirklin) is an accurate ground truth, including compounds with no reported uncertainty.
    Methods 'Reference data a'; all MAE numbers are measured against this set.
  • domain assumption PET-OMATPES zero-shot energies faithfully approximate r2SCAN DFT for these materials.
    Results 'MLIPs to improve DFT versus experiment'; validated against Alexandria in SI Fig. S9, but no meta-GGA MC3D exists.
  • standard math Stratified 80/20 splits with 30 seeds give unbiased estimates of out-of-sample error.
    Methods 'Machine learning models and training'; standard CV assumption, but does not fix distribution shift from the experimental subset to the full database.
  • domain assumption Latent features of PET-OMATPES encode structure-dependent correction-relevant information without leaking the target.
    SI Fig. S8 KPCovR analysis supports informativeness; no leakage because the MLIP was trained on DFT energies, not experimental residuals.

pith-pipeline@v1.3.0-alltime-deepseek · 29113 in / 12934 out tokens · 126169 ms · 2026-08-01T16:00:48.786374+00:00 · methodology

0 comments
read the original abstract

Crystal structure databases curated by high-throughput density functional theory calculations typically serve as the starting point for computational materials discovery efforts. Thermodynamic stability data, such as formation energies and the energy above the convex hull, are important quantities to guide the search for novel materials, enabling filtering for (meta)stable structures. Here, we present the thermodynamic stability of the fully open-source, reproducible, and experimentally focused Materials Cloud three-dimensional crystals database (MC3D). We compare against two other DFT databases, the Open Quantum Materials Database (OQMD) and the Materials Project (MP), as well as against experimental formation enthalpies. We then demonstrate how recent foundational machine learning interatomic potentials (MLIPs) trained at the r$^2$SCAN level (specifically, we test PET-OMATPES here) can be leveraged to improve the agreement of formation energies with experiment, reducing the mean absolute error by more than 40% relative to GGA without requiring any additional DFT calculation. Our results validate and extend the established practice of combining PBEsol geometries with meta-GGA energies to the era of foundational MLIPs. Finally, we train classical machine learning models to further correct the formation energies in a delta-learning framework, where we use the information-rich latent features of the foundational MLIP. These models further reduce the mean absolute error below 50 meV/atom, bringing it down to values comparable with the experimental uncertainty itself. Notably, compared to purely compositional features, the latent features (combined with carefully tuned regularization) simultaneously reduce the prediction error and limit the impact of the learned corrections on the relative phase stability.

Figures

Figures reproduced from arXiv: 2607.18092 by Giovanni Pizzi, Marnik Bercx, Timo Reents.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

93 extracted references · 3 canonical work pages

  1. [1]

    S. P. Huber, S. Zoupanos, M. Uhrin, L. Talirz, L. Kahle, R. H¨ auselmann, D. Gresch, T. M¨ uller, A. V. Yakutovich, C. W. Andersen, F. F. Ramirez, C. S. Adorf, F. Gargiulo, S. Kumbhar, E. Passaro, C. Johnston, A. Merkys, A. Ce- pellotti, N. Mounet, N. Marzari, B. Kozinsky, and G. Pizzi, Scientific Data7, 300 (2020), number: 1

  2. [2]

    Uhrin, S

    M. Uhrin, S. P. Huber, J. Yu, N. Marzari, and G. Pizzi, Computational Materials Science187, 110086 (2021)

  3. [3]

    A. S. Rosen, M. Gallant, J. George, J. Riebesell, H. Sa- hasrabuddhe, J.-X. Shen, M. Wen, M. L. Evans, G. Pe- tretto, D. Waroquiers, G.-M. Rignanese, K. A. Persson, A. Jain, and A. M. Ganose, Journal of Open Source Soft- ware9, 5995 (2024)

  4. [4]

    A. Jain, S. P. Ong, W. Chen, B. Medasani, X. Qu, M. Kocher, M. Brafman, G. Petretto, G.-M. Rignanese, G. Hautier, D. Gunter, and K. A. Persson, Concurrency and Computation: Practice and Experience27, 5037 (2015), eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/cpe.3505

  5. [5]

    Mathew, J

    K. Mathew, J. H. Montoya, A. Faghaninia, S. Dwarakanath, M. Aykol, H. Tang, I.-h. Chu, T. Smidt, B. Bocklund, M. Horton, J. Dagdelen, B. Wood, Z.-K. Liu, J. Neaton, S. P. Ong, K. Persson, and A. Jain, Computational Materials Science139, 140 (2017)

  6. [6]

    Ganose, J

    A. Ganose, J. Riebesell, J. George, J.-X. Shen, A. S. Rosen, A. Ashok Naik, N. Winner, M. Wen, R. Guha, M. Kuner, G. Petretto, Z. Zhu, M. Hor- ton, H. Sahasrabuddhe, A. Kaplan, J. Schmidt, C. Er- tural, R. Kingsbury, M. McDermott, R. Goodall, A. Bonkowski, T. Purcell, D. Z¨ ugner, and J. Qi, ato- mate2 (2024)

  7. [7]

    A. Jain, G. Hautier, C. J. Moore, S. Ping Ong, C. C. Fischer, T. Mueller, K. A. Persson, and G. Ceder, Com- putational Materials Science50, 2295 (2011)

  8. [8]

    A. Jain, S. P. Ong, G. Hautier, W. Chen, W. D. Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. A. Persson, APL Materials1, 011002 (2013)

  9. [9]

    J. E. Saal, S. Kirklin, M. Aykol, B. Meredig, and C. Wolverton, JOM65, 1501 (2013)

  10. [10]

    Kirklin, J

    S. Kirklin, J. E. Saal, B. Meredig, A. Thompson, J. W. Doak, M. Aykol, S. R¨ uhl, and C. Wolverton, npj Com- putational Materials1, 1 (2015), number: 1

  11. [11]

    Curtarolo, W

    S. Curtarolo, W. Setyawan, G. L. W. Hart, M. Jahnatek, R. V. Chepulskii, R. H. Taylor, S. Wang, J. Xue, K. Yang, O. Levy, M. J. Mehl, H. T. Stokes, D. O. Demchenko, and D. Morgan, Computational Materials Science58, 218 (2012)

  12. [12]

    Curtarolo, W

    S. Curtarolo, W. Setyawan, S. Wang, J. Xue, K. Yang, R. H. Taylor, L. J. Nelson, G. L. W. Hart, S. Sanvito, M. Buongiorno-Nardelli, N. Mingo, and O. Levy, Com- putational Materials Science58, 227 (2012)

  13. [13]

    C. E. Calderon, J. J. Plata, C. Toher, C. Oses, O. Levy, M. Fornari, A. Natan, M. J. Mehl, G. Hart, M. Buon- giorno Nardelli, and S. Curtarolo, Computational Mate- rials Science108, 233 (2015)

  14. [14]

    Schmidt, N

    J. Schmidt, N. Hoffmann, H.-C. Wang, P. Borlido, P. J. M. A. Carri¸ co, T. F. T. Cerqueira, S. Botti, and M. A. L. Marques, Advanced Materials35, 2210788 (2023)

  15. [15]

    Schmidt, H.-C

    J. Schmidt, H.-C. Wang, T. F. T. Cerqueira, S. Botti, and M. A. L. Marques, Scientific Data9, 64 (2022)

  16. [16]

    Schmidt, T

    J. Schmidt, T. F. Cerqueira, A. H. Romero, A. Loew, F. J¨ ager, H.-C. Wang, S. Botti, and M. A. Marques, Ma- terials Today Physics48, 101560 (2024)

  17. [17]

    S. P. Huber, M. Minotakis, M. Bercx, T. Reents, K. Eimre, N. Paulish, N. H¨ ormann, M. Uhrin, N. Marzari, and G. Pizzi, Digital Discovery5, 1114 (2026)

  18. [18]

    Talirz, S

    L. Talirz, S. Kumbhar, E. Passaro, A. V. Yakutovich, V. Granata, F. Gargiulo, M. Borelli, M. Uhrin, S. P. Huber, S. Zoupanos, C. S. Adorf, C. W. Andersen, O. Sch¨ utt, C. A. Pignedoli, D. Passerone, J. VandeVon- dele, T. C. Schulthess, B. Smit, G. Pizzi, and N. Marzari, Scientific Data7, 299 (2020)

  19. [19]

    Hohenberg and W

    P. Hohenberg and W. Kohn, Physical Review136, B864 (1964)

  20. [20]

    Kohn and L

    W. Kohn and L. J. Sham, Physical Review140, A1133 (1965)

  21. [21]

    Cavignac, J

    T. Cavignac, J. Schmidt, P.-P. D. Breuck, A. Loew, T. F. T. Cerqueira, H.-C. Wang, A. Bochkarev, Y. Lyso- gorskiy, A. H. Romero, R. Drautz, S. Botti, and M. A. L. Marques, arXiv preprint arXiv:2512.09169 (2025)

  22. [22]

    A. S. Parackal, F. Trybel, F. A. Faber, and R. Armiento, arXiv preprint arXiv:2601.21393 (2026)

  23. [23]

    W. Sun, S. T. Dacek, S. P. Ong, G. Hautier, A. Jain, W. D. Richards, A. C. Gamst, K. A. Persson, and G. Ceder, Science Advances2, e1600225 (2016)

  24. [24]

    J. Meng, M. S. Sheikh, R. Jacobs, J. Liu, W. O. Nachlas, X. Li, and D. Morgan, Nature Materials , 1 (2024)

  25. [25]

    H.-C. Wang, S. Botti, and M. A. L. Marques, npj Com- putational Materials7, 1 (2021)

  26. [26]

    J. Sun, A. Ruzsinszky, and J. Perdew, Physical Review Letters115, 036402 (2015). 12

  27. [27]

    A. P. Bart´ ok and J. R. Yates, The Journal of Chemical Physics150, 161101 (2019)

  28. [28]

    J. W. Furness, A. D. Kaplan, J. Ning, J. P. Perdew, and J. Sun, The Journal of Physical Chemistry Letters11, 8208 (2020)

  29. [29]

    Zhang, D

    Y. Zhang, D. A. Kitchaev, J. Yang, T. Chen, S. T. Dacek, R. A. Sarmiento-P´ erez, M. A. L. Marques, H. Peng, G. Ceder, J. P. Perdew, and J. Sun, npj Computational Materials4, 1 (2018)

  30. [30]

    Kingsbury, A

    R. Kingsbury, A. S. Gupta, C. J. Bartel, J. M. Munro, S. Dwaraknath, M. Horton, and K. A. Persson, Physical Review Materials6, 013801 (2022)

  31. [31]

    E. B. Isaacs and C. Wolverton, Physical Review Materials 2, 063801 (2018)

  32. [32]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Physical Re- view Letters77, 3865 (1996)

  33. [33]

    J. P. Perdew, A. Ruzsinszky, G. I. Csonka, O. A. Vydrov, G. E. Scuseria, L. A. Constantin, X. Zhou, and K. Burke, Physical Review Letters100, 136406 (2008)

  34. [34]

    R. S. Kingsbury, A. S. Rosen, A. S. Gupta, J. M. Munro, S. P. Ong, A. Jain, S. Dwaraknath, M. K. Horton, and K. A. Persson, npj Computational Materials8, 1 (2022)

  35. [35]

    M. C. Kuner, A. D. Kaplan, K. A. Persson, M. Asta, and D. C. Chrzan, npj Computational Materials11, 352 (2025)

  36. [36]

    Malosso, F

    C. Malosso, F. Bigi, P. Pegolo, J. W. Abbott, P. Loche, M. Rossi, M. Ceriotti, and A. Mazitov, arXiv preprint arXiv:2603.02089 (2026)

  37. [37]

    A. D. Kaplan, R. Liu, J. Qi, T. W. Ko, B. Deng, J. Riebe- sell, G. Ceder, K. A. Persson, and S. P. Ong, arXiv preprint arXiv:2503.04070 (2025)

  38. [38]

    Bosoni, L

    E. Bosoni, L. Beal, M. Bercx, P. Blaha, S. Bl¨ ugel, J. Br¨ oder, M. Callsen, S. Cottenier, A. Degomme, V. Dikan, K. Eimre, E. Flage-Larsen, M. Fornari, A. Garcia, L. Genovese, M. Giantomassi, S. P. Huber, H. Janssen, G. Kastlunger, M. Krack, G. Kresse, T. D. K¨ uhne, K. Lejaeghere, G. K. H. Madsen, M. Marsman, N. Marzari, G. Michalicek, H. Mirhosseini, T...

  39. [39]

    Prandini, A

    G. Prandini, A. Marrazzo, I. E. Castelli, N. Mounet, and N. Marzari, npj Computational Materials4, 1 (2018)

  40. [40]

    Grindy, B

    S. Grindy, B. Meredig, S. Kirklin, J. E. Saal, and C. Wolverton, Physical Review B87, 075150 (2013)

  41. [41]

    L. Wang, T. Maxisch, and G. Ceder, Physical Review B 73, 195107 (2006)

  42. [42]

    Stevanovi´ c, S

    V. Stevanovi´ c, S. Lany, X. Zhang, and A. Zunger, Phys- ical Review B85, 115104 (2012)

  43. [43]

    A. Wang, R. Kingsbury, M. McDermott, M. Horton, A. Jain, S. P. Ong, S. Dwaraknath, and K. A. Persson, Scientific Reports11, 15496 (2021), number: 1

  44. [44]

    Friedrich, D

    R. Friedrich, D. Usanmaz, C. Oses, A. Supka, M. Fornari, M. Buongiorno Nardelli, C. Toher, and S. Curtarolo, npj Computational Materials5, 1 (2019), number: 1

  45. [45]

    Friedrich and S

    R. Friedrich and S. Curtarolo, The Journal of Chemical Physics160, 042501 (2024)

  46. [46]

    Hautier, S

    G. Hautier, S. P. Ong, A. Jain, C. J. Moore, and G. Ceder, Physical Review B85, 155208 (2012)

  47. [47]

    G. G. C. Peterson and J. Brgoch, Journal of Physics: Energy3, 022002 (2021)

  48. [48]

    S. Gong, S. Wang, T. Xie, W. H. Chae, R. Liu, Y. Shao- Horn, and J. C. Grossman, JACS Au2, 1964 (2022)

  49. [49]

    C. J. Bartel, A. Trewartha, Q. Wang, A. Dunn, A. Jain, and G. Ceder, npj Computational Materials6, 1 (2020)

  50. [50]

    V. I. Hegde, C. K. H. Borg, Z. Del Rosario, Y. Kim, M. Hutchinson, E. Antono, J. Ling, P. Saxe, J. E. Saal, and B. Meredig, Physical Review Materials7, 053805 (2023)

  51. [51]

    R. A. Gouvˆ ea, P.-P. De Breuck, T. Pretto, G.-M. Rig- nanese, and M. J. L. Santos, npj Computational Materi- als12, 67 (2026)

  52. [52]

    S. Y. Kim, Y. J. Park, and J. Li, npj Computational Materials 10.1038/s41524-026-02167-x (2026)

  53. [53]

    Adhikari, C

    S. Adhikari, C. J. Bartel, and C. Sutton, arXiv preprint arXiv:2307.07609 (2023)

  54. [54]

    M. K. Horton, P. Huck, R. X. Yang, J. M. Munro, S. Dwaraknath, A. M. Ganose, R. S. Kingsbury, M. Wen, J. X. Shen, T. S. Mathis, A. D. Kaplan, K. Berket, J. Riebesell, J. George, A. S. Rosen, E. W. C. Spotte- Smith, M. J. McDermott, O. A. Cohen, A. Dunn, M. C. Kuner, G.-M. Rignanese, G. Petretto, D. Waroquiers, S. M. Griffin, J. B. Neaton, D. C. Chrzan, M....

  55. [55]

    A. A. Emery and C. Wolverton, Scientific Data4, 170153 (2017)

  56. [56]

    Priya and N

    P. Priya and N. R. Aluru, npj Computational Materials 7, 90 (2021)

  57. [57]

    Woods-Robinson, D

    R. Woods-Robinson, D. Broberg, A. Faghaninia, A. Jain, S. S. Dwaraknath, and K. A. Persson, Chemistry of Ma- terials30, 8375 (2018)

  58. [58]

    Mueller, G

    T. Mueller, G. Hautier, A. Jain, and G. Ceder, Chemistry of Materials23, 3854 (2011)

  59. [59]

    Kirklin, J

    S. Kirklin, J. E. Saal, V. I. Hegde, and C. Wolverton, Acta Materialia102, 125 (2016)

  60. [60]

    Aykol, S

    M. Aykol, S. S. Dwaraknath, W. Sun, and K. A. Persson, Science Advances4, eaaq0148 (2018)

  61. [61]

    Warford, F

    T. Warford, F. L. Thiemann, and G. Cs´ anyi, arXiv preprint arXiv:2601.21056 (2026)

  62. [62]

    H. Yu, M. Giantomassi, G. Materzanini, J. Wang, and G. Rignanese, Materials Genome Engineering Advances 2, e58 (2024)

  63. [63]

    S. N. Pozdnyakov and M. Ceriotti, arXiv preprint arXiv:2305.19302 (2024)

  64. [64]

    F. Bigi, P. Pegolo, A. Mazitov, and M. Ceriotti, arXiv preprint arXiv:2601.16195 (2026)

  65. [65]

    Riebesell, R

    J. Riebesell, R. E. A. Goodall, P. Benner, Y. Chiang, B. Deng, G. Ceder, M. Asta, A. A. Lee, A. Jain, and K. A. Persson, Nature Machine Intelligence7, 836 (2025)

  66. [66]

    Barroso-Luque, M

    L. Barroso-Luque, M. Shuaibi, X. Fu, B. M. Wood, M. Dzamba, M. Gao, A. Rizvi, C. L. Zitnick, and Z. W. Ulissi, arXiv preprint arXiv:2410.12771 (2024)

  67. [67]

    G. I. Csonka, J. P. Perdew, A. Ruzsinszky, P. H. T. Philipsen, S. Leb` egue, J. Paier, O. A. Vydrov, and J. G. ´Angy´ an, Physical Review B79, 155107 (2009)

  68. [68]

    Chorna, D

    S. Chorna, D. Tisi, C. Malosso, W. B. How, M. Ce- riotti, and S. Chong, Advanced Intelligent Systems8, 10.1002/aisy.202501497 (2026)

  69. [69]

    Breiman, Machine Learning45, 5 (2001)

    L. Breiman, Machine Learning45, 5 (2001)

  70. [70]

    K. P. Murphy,Machine learning: a probabilistic perspec- tive, 4th ed., Adaptive computation and machine learn- ing series (MIT Press, Cambridge, Mass., 2013) chapter 14.4.3, pp. 492-493. 13

  71. [71]

    C. E. Rasmussen and C. K. I. Williams,Gaussian Pro- cesses for Machine Learning(The MIT Press, 2005)

  72. [72]

    L. Ward, A. Agrawal, A. Choudhary, and C. Wolverton, npj Computational Materials2, 16028 (2016)

  73. [73]

    L. Ward, A. Dunn, A. Faghaninia, N. E. R. Zimmermann, S. Bajaj, Q. Wang, J. Montoya, J. Chen, K. Bystrom, M. Dylla, K. Chard, M. Asta, K. A. Persson, G. J. Sny- der, I. Foster, and A. Jain, Computational Materials Sci- ence152, 60 (2018)

  74. [74]

    C. J. Bartel, A. W. Weimer, S. Lany, C. B. Musgrave, and A. M. Holder, npj Computational Materials5, 1 (2019)

  75. [75]

    Lejaeghere, G

    K. Lejaeghere, G. Bihlmayer, T. Bj¨ orkman, P. Blaha, S. Bl¨ ugel, V. Blum, D. Caliste, I. E. Castelli, S. J. Clark, A. Dal Corso, S. de Gironcoli, T. Deutsch, J. K. Dewhurst, I. Di Marco, C. Draxl, M. Du lak, O. Eriksson, J. A. Flores-Livas, K. F. Garrity, L. Gen- ovese, P. Giannozzi, M. Giantomassi, S. Goedecker, X. Gonze, O. Gr ˚ an¨ as, E. K. U. Gross...

  76. [76]

    Kozhevnikov, M

    A. Kozhevnikov, M. Taillefumier, and S. Pintarelli, SIR- IUS, https://github.com/electronic-structure/SIRIUS, accessed: 2024-06-12

  77. [77]

    Giannozzi, S

    P. Giannozzi, S. Baroni, N. Bonini, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, G. L. Chiarotti, M. Cococ- cioni, I. Dabo, A. D. Corso, S. d. Gironcoli, S. Fabris, G. Fratesi, R. Gebauer, U. Gerstmann, C. Gougoussis, A. Kokalj, M. Lazzeri, L. Martin-Samos, N. Marzari, F. Mauri, R. Mazzarello, S. Paolini, A. Pasquarello, L. Paulatto, C. Sbraccia, S. Sc...

  78. [78]

    Giannozzi, O

    P. Giannozzi, O. Andreussi, T. Brumme, O. Bunau, M. B. Nardelli, M. Calandra, R. Car, C. Cavazzoni, D. Ceresoli, M. Cococcioni, N. Colonna, I. Carnimeo, A. D. Corso, S. d. Gironcoli, P. Delugas, R. A. DiStasio, A. Ferretti, A. Floris, G. Fratesi, G. Fugallo, R. Gebauer, U. Gerstmann, F. Giustino, T. Gorni, J. Jia, M. Kawa- mura, H.-Y. Ko, A. Kokalj, E. K¨...

  79. [79]

    AiiDA-QuantumESPRESSO, https://github.com/aiidateam/aiida-quantumespresso, accessed: 2024-06-12, version 4.4.0

  80. [80]

    de Miranda Nascimento, F

    G. de Miranda Nascimento, F. J. dos Santos, M. Bercx, D. Grassano, G. Pizzi, and N. Marzari, npj Computa- tional Materials 10.1038/s41524-026-02097-8 (2026)

Showing first 80 references.