Pith. sign in

REVIEW 3 major objections 4 minor 3 cited by

The carbon cost of materials discovery: Can machine learning really accelerate the discovery of new photovoltaics?

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper establishes that machine-learning surrogates can replace the most expensive density-functional-theory steps in photovoltaic materials screening, with direct prediction of a material's maximum efficiency beating stepwise…

desk verdict A careful, honest benchmark with a real caveat: the accuracy rankings are relative to a scissor-corrected GGA reference, not HSE truth, but the carbon-cost framework and external validation make it worth review. read the letter →

arxiv 2507.13246 v1 pith:PJD5ZIHE submitted 2025-07-17 cond-mat.mtrl-sci cs.LG

classification cond-mat.mtrl-scics.LG
keywords machinelearningdensityfunctionaltheoryphotovoltaicsSLMEcarbonemissionsmaterialsscreeningsurrogatemodelsscissorcorrection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether machine learning can replace expensive density-functional-theory calculations in the search for new photovoltaic materials without losing accuracy, and it frames the answer in carbon emissions as well as error. By decomposing the standard SLME (spectroscopic limited maximum efficiency) workflow into components and swapping each for an ML surrogate, the authors find that directly predicting the final efficiency number is more tractable than predicting absorption spectra along the way. A hybrid approach that learns a 'scissor correction' to cheap GGA spectra lands on the accuracy–cost Pareto front with mean absolute error under one percentage point. The authors also show that their best ML model's error on an external dataset is comparable to the disagreement between two different DFT methods, suggesting ML can be as reliable as methodological choices within DFT. The point is that low-carbon, high-throughput screening is achievable now if the right quantities are predicted.

What carries the argument

The central object is the SLME workflow, a chain of components: fundamental band gap, direct dipole-allowed gap, absorption spectrum, absorptance, radiative-recombination offset, and final efficiency. The paper compares eight methods (I–VIII) that replace different components with either GGA/hybrid DFT calculations or ALIGNN graph-neural-network predictions. The load-bearing identity is the scissor correction $\alpha_{\mathrm{HSE}} \approx \alpha_{\mathrm{GGA}}(E-\Delta E)$ with $\Delta E = E_g^{\mathrm{HSE}}-E_g^{\mathrm{GGA}}$, which turns cheap GGA spectra into the reference 'truth' that all methods are scored against. That identity defines the test-set labels, so the entire accuracy ranking rests on it.

What would settle it

Compute true HSE or GW absorption spectra and SLMEs for a set of indirect-gap and dipole-forbidden absorbers, then compare those values to the scissor-corrected GGA reference used in this paper; if the gap between them is comparable to or larger than the 6.8 to 7.2 percentage-point differences the paper treats as interchangeable, the Pareto-front ranking of Methods I through VIII would not survive.

Watch

Extended reading notes

Core claim

The paper establishes that for screening photovoltaic absorbers by their spectroscopic limited maximum efficiency (SLME), the most accurate and lowest-carbon strategy is to train a graph neural network to predict the SLME directly from structure, rather than to predict intermediate optical or electronic properties. An intermediate method that keeps GGA-level absorption spectra but applies a learned scissor correction to the gap is nearly as accurate, with mean absolute error under 1.0 percentage points relative to the reference, at roughly two orders of magnitude less cost than hybrid-functional calculations. Against an external test set, the direct SLME model achieves 6.8 percentage points MAE, which is no larger than the 7.2 percentage points MAE between two independently used DFT functionals (TB-mBJ and delta-sol-corrected GGA). The authors therefore claim that ML surrogates trained on consistent high-fidelity data can match the variability of DFT methodology itself, and that the main bottleneck is the scarcity of consistent, high-fidelity SLME datasets rather than model architecture.

Load-bearing premise

The ranking of methods is only as trustworthy as the scissor-corrected GGA spectra used as the reference for true SLME; if those spectra misrepresent real optics for indirect-gap or dipole-forbidden absorbers, then all accuracy claims are measured against a benchmark whose own error was never bounded.

Editorial extensions

If this is right

  • Direct ML prediction of SLME should be the default screening tool for initial large-scale searches, reserving DFT for shortlisted candidates.
  • Method III, using GGA spectra plus a learned scissor correction, offers a near-HSE accuracy sweet spot at low carbon cost for users who need interpretability.
  • Expanding high-fidelity SLME datasets matters more than expanding low-quality datasets, while spectral prediction still needs both better architectures and more data.
  • Reporting carbon emissions alongside accuracy should become standard in computational materials discovery.
  • ML surrogate error on external data being comparable to DFT-functional spread means that benchmark comparisons against a single DFT method can overstate or understate model quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the scissor-corrected GGA reference is wrong for indirect-gap absorbers, then Methods I and III's apparent accuracy may partly reflect learning the reference rather than physical efficiency; a recalibration against GW-quality SLMEs could reshuffle the Pareto front.
  • The same carbon-accuracy Pareto methodology could be applied to other materials design targets, such as battery electrolytes or thermoelectrics, where a scalar figure of merit can be directly predicted.
  • Because a single ML inference costs about 1/2000 of a static GGA calculation, the dominant carbon cost of the recommended workflows becomes the training run; publishing trained model weights would let downstream users amortize that cost to near zero.
  • The 6.8 versus 7.2 percentage-point comparison suggests that model epistemic uncertainty, not just point predictions, should be reported so screening decisions can account for ML versus DFT disagreement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper constructs eight computational workflows for estimating the spectroscopic limited maximum efficiency (SLME) of photovoltaic materials, ranging from hybrid-functional DFT to direct machine-learning prediction of the SLME, and compares them on a fixed held-out test set in terms of both predictive accuracy (MAE in SLME and in rank) and estimated CO2 emissions. The central findings are that direct scalar SLME prediction (Method I) and a learned scissor correction on GGA spectra (Method III) dominate the accuracy-emissions front, that intermediate spectral prediction (Method II) is markedly less accurate, and that ML predictions are competitive with the spread between different DFT approximations on an external dataset. The study includes learning curves with a null baseline, three-subsample averaging, and an external validation set from Fabini et al., and it releases code and data.

Significance. If the conclusions withstand scrutiny, the paper provides a valuable quantitative template for jointly assessing predictive accuracy and environmental cost in computational discovery, and it gives concrete guidance on where ML surrogates are currently most useful in PV screening. The strengths are the careful experimental design: a fixed test set shared by all models, a null-hypothesis baseline in the learning curves, three-subsample averaging, explicit carbon accounting via CodeCarbon, and an external validation set (the Fabini delta-sol data) that protects the central comparison from the most obvious same-workflow circularity. The available code and data are a further positive point. However, the benchmark against which all accuracies are measured is itself an approximation whose error is not quantified, and this limits the strength of several absolute and comparative claims.

major comments (3)
  1. [II.C, Table I, Figures 3 and 4] All accuracy numbers in the paper, including the headline claim that Method III achieves 'MAEs under 1.0 percentage points' (Section III.D), are distances to a reference defined as a scissor-corrected GGA spectrum with GGA-level offsets (Method VIII). The authors explicitly state in Section II.C that 'the test dataset used GGA-level offsets, so this source of error was not examined in this work.' Since the scissor approximation is known to be least reliable for indirect-gap or dipole-forbidden absorbers, which are precisely the cases SLME was designed to treat, the absolute errors and the ranking of Methods I-VIII are measured against an unquantified benchmark. The external validation in Section III.E shows that the ML error (6.8 pp) is comparable to the inter-DFT spread between two approximate references (7.2 pp), but it does not bound the benchmark bias itself. The paper should quantify the reference error on at least a subset with true HSE-level spectra, or substantially soften the conclusions that depend on absolute accuracy.
  2. [III.E and Abstract] The abstract claims that 'ML models trained on DFT data can outperform DFT workflows using alternative exchange-correlation functionals in screening applications.' The evidence in Section III.E is that the ML-versus-delta-sol MAE is 6.8 percentage points, compared with 7.2 percentage points for the TB-mBJ-versus-delta-sol comparison. These numbers are very close, no uncertainty or significance test is reported, and both references are approximate DFT methods. The data support the conclusion that ML predictions fall within the inter-method spread of DFT approximations, not that ML outperforms a DFT workflow. This claim should be rephrased to match the evidence.
  3. [III.A and Figure 2] The statement that 'extrapolation at the current rate of model improvement with several tens of thousands of high-quality estimates of SLME a model with negligible errors is possible' is not justified by the presented data. The learning curves in Figure 2 are averaged over only three random subsamples, are shown without error bars, and are not fitted to a scaling law. The claim of a steepest gradient for direct SLME prediction is therefore an informal visual extrapolation rather than a quantitative result. This overstates the expected benefit of dataset expansion and should be removed or replaced with a more cautious statement.
minor comments (4)
  1. [III.E] The text refers to 'rank correlation plots (Figure 4)', but Figure 4 is the cost-accuracy Pareto plot; the rank correlation plots appear to be in the supplementary information. Please correct the figure reference.
  2. [II.B] The equation for the black-body spectrum is incorrect as printed: it shows phi_BB(E) = (2 e pi / (h^3 c^2)) integral A(E) E^2 / (exp(E/k_B T) - 1) dE, which contains an extra integral, an extra factor of the electron charge e, and an erroneous A(E) inside the integral. The standard definition is phi_BB(E) = (2 pi / (h^3 c^2)) E^2 / (exp(E/k_B T) - 1). Please correct this equation so that the methodology is unambiguous.
  3. [I, II.A] There are minor typographical errors: 'characterisaton' in the Introduction and 'theCodeCarbon' in Section II.A.
  4. [Figure 2] The learning curves would benefit from error bars or shaded confidence intervals on the three-subsample averages, since the gradient comparisons in the text rest on these curves.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the ML surrogates are trained and evaluated on defined DFT-workflow outputs, with independent external validation grounding the headline generalization claims.

full rationale

The derivation chain is self-contained and not circular. The ML models are trained to reproduce specific DFT-workflow outputs (band gaps, spectra, scissor corrections, and SLME values), and their reported errors are genuine held-out test-set errors against labels generated by the same canonical workflow described as Method VIII. This is standard supervised regression on a defined target, not a prediction that is equivalent to its input by construction. The paper explicitly flags the benchmark's own limitation: 'the test dataset used GGA-level offsets, so this source of error was not examined in this work' and 'these errors are relative to the fully hybrid approach, which is itself limited.' These are accuracy and external-validity caveats, not circularity. The headline generalization claims are additionally grounded by independent external data: the Fabini Δ-sol external set and the Yu-Zunger GW SLME set, with inter-DFT MAE (7.2 pp TB-mBJ vs Δ-sol) comparable to ML-vs-Δ-sol MAE (6.8 pp). The scissor-correction approximation is supported by an external citation (Yang et al.), not by a self-citation. No load-bearing step reduces to its own input, and no self-citation chain is used to force the conclusions.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The paper's conclusions rest on standard physics (detailed balance, SLME), on the adequacy of scissor-corrected GGA spectra as a stand-in for HSE (borrowed from Yang et al.), on the SLME as a sufficient screening metric, and on CodeCarbon's UK-grid emission estimates. The ML weights and the 100-bin spectral discretization are fitted or hand-chosen quantities, but they are components of the surrogate models rather than parameters of a derivation. The central numeric results are benchmark measurements, not derivations, so the ledger is short of invented entities but carries real domain assumptions about the reference workflow.

free parameters (4)
  • Spectral binning dimension = 100
    Absorption spectra are binned into 100 dimensions via numpy.interpolate() rather than a latent compression; chosen by hand and justified by Kaundinya et al. This discretization affects what the spectral prediction models output and therefore the error structure of Method II.
  • Training schedule = 300 epochs, batch size 64 (batch size 2 for learning curves)
    Fixed for all models 'for consistency'; the text notes no hyperparameter tuning was performed. Model capacity is therefore not optimized per property, which shapes all accuracy numbers.
  • SLME device thickness d = not stated in the provided text
    The absorptance A(E) = 1 - exp(-2d*alpha(E)) in Section IIB requires a thickness; without its value the reported efficiency numbers cannot be reproduced from the manuscript text alone.
  • ALIGNN model weights = learned from about 4.8k training materials
    All accuracy and ranking numbers in Figures 3 and 4 are produced by weights fitted to the W-R/Kim DFT labels; this is standard supervised fitting, but the headline performance figures are conditional on this fit and on the fixed hyperparameters.
assumptions (5)
  • standard math Detailed balance (Shockley-Queisser) physics governs maximum PV efficiency
    Section IIB uses the detailed balance expressions for J_SC and J_0, the AM1.5G spectrum, and the black-body spectrum; standard physics taken from refs. 26-27.
  • domain assumption Scissor-corrected GGA spectrum approximates the HSE spectrum: alpha_HSE ~ alpha_GGA(E - DeltaE)
    Section IIC. This is the reference 'truth' (Method VIII) for every accuracy ranking; supported by a citation to Yang et al., but its error is never quantified in this work.
  • domain assumption SLME with unit internal quantum efficiency and Lambert-Beer absorptance is an adequate figure of merit for PV screening
    Section IIB defines the FoM; Section IIIF acknowledges the Blank selection metric would be more accurate but lacks large datasets. The ranking conclusions are conditional on SLME being a meaningful target.
  • domain assumption CodeCarbon energy monitoring with UK grid emission factors yields meaningful relative carbon costs
    Section IIA: emissions are estimated from UK grid intensity via CodeCarbon; the paper mostly reports relative costs to mitigate this, but every relative number inherits the monitoring assumptions.
  • domain assumption Test labels from the W-R/Kim HSE-scissor workflow are a fair reference for ranking all methods
    Section III and Table I: all internal accuracy metrics are differences from Method VIII, and the test set uses GGA-level offsets, a source of error the authors state was not examined.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The carbon cost of materials discovery: Can machine learning really accelerate the discovery of new photovoltaics?." pith.science (2026). https://pith.science/paper/PJD5ZIHE

@misc{pith2026250713246,
  author       = {Pith},
  title        = {Pith review of: The carbon cost of materials discovery: Can machine learning really accelerate the discovery of new photovoltaics?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PJD5ZIHE}},
  note         = {Machine review of arXiv:2507.13246}
}
abstract

Computational screening has become a powerful complement to experimental efforts in the discovery of high-performance photovoltaic (PV) materials. Most workflows rely on density functional theory (DFT) to estimate electronic and optical properties relevant to solar energy conversion. Although more efficient than laboratory-based methods, DFT calculations still entail substantial computational and environmental costs. Machine learning (ML) models have recently gained attention as surrogates for DFT, offering drastic reductions in resource use with competitive predictive performance. In this study, we reproduce a canonical DFT-based workflow to estimate the maximum efficiency limit and progressively replace its components with ML surrogates. By quantifying the CO$_2$ emissions associated with each computational strategy, we evaluate the trade-offs between predictive efficacy and environmental cost. Our results reveal multiple hybrid ML/DFT strategies that optimize different points along the accuracy--emissions front. We find that direct prediction of scalar quantities, such as maximum efficiency, is significantly more tractable than using predicted absorption spectra as an intermediate step. Interestingly, ML models trained on DFT data can outperform DFT workflows using alternative exchange--correlation functionals in screening applications, highlighting the consistency and utility of data-driven approaches. We also assess strategies to improve ML-driven screening through expanded datasets and improved model architectures tailored to PV-relevant features. This work provides a quantitative framework for building low-emission, high-throughput discovery pipelines.

Figures

Figures reproduced from arXiv: 2507.13246 by the authors.

Figure 1
Figure 1. FIG. 1. Plot of the methods for estimating SLMEs outlined [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Learning curves for each property, looking at a) relative errors for the given property and b) the resultant error when [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Violin plot comparing the success of the methods out [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: FIG. 4. Pareto front for performance vs cost for the methods outlined in Table I, where performance is measured as a) MAE [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5. Violin plot comparing the successes of the SLME [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Discovery and recovery of crystalline materials with property-conditioned transformers

    cond-mat.mtrl-sci 2025-11 conditional novelty 6.0 of 10

    Conditioning the attention layers of a crystal-writing transformer on continuous property values enables XRD-based structure recovery and targeted generation of photovoltaic candidates.

  2. Learning disentangled latent representations facilitates discovery and design of functional materials

    cond-mat.mtrl-sci 2025-07 conditional novelty 6.0 of 10

    An unsupervised disentangling autoencoder learns a latent dimension in optical absorption spectra that correlates with photovoltaic efficiency, reflects the direct-to-indirect gap transition, and accelerates simulated...

  3. VASP Plugins: Linking the Vienna ab-initio Simulation Package with Python

    cond-mat.mtrl-sci 2026-07 accept novelty 5.5 of 10

    A C++/pybind11 shared-memory plugin layer exposes VASP SCF and ionic data as NumPy arrays so Python can modify structure, forces, local potential, and occupancies in place.

Reference graph

Works this paper leans on

72 extracted references · 60 canonical work pages · cited by 3 Pith papers

  1. [1]

    A. K. Cheetham, R. Seshadri, and F. Wudl, Nature Syn- thesis 1, 514 (2022)

  2. [2]

    Masson, E

    G. Masson, E. Bosch, A. Van Rechem, and M. de l’Epine, IEA Photovoltaic Power Systems Programme https://doi.org/10.69766/VHRF4040 (2024)

  3. [3]

    N. M. Haegel, H. Atwater, T. Barnes, C. Breyer, A. Bur- rell, Y.-M. Chiang, S. De Wolf, B. Dimmler, D. Feldman, S. Glunz, J. C. Goldschmidt, D. Hochschild, R. Inzunza, I. Kaizuka, B. Kroposki, S. Kurtz, S. Leu, R. Margo- lis, K. Matsubara, A. Metz, W. K. Metzger, M. Mor- jaria, S. Niki, S. Nowak, I. M. Peters, S. Philipps, T. Reindl, A. Richter, D. Rose, ...

  4. [4]

    J. C. Blakesley, R. S. Bonilla, M. Freitag, A. M. Ganose, N. Gasparini, P. Kaienburg, G. Koutsourakis, J. D. Ma- jor, J. Nelson, N. K. Noel, B. Roose, J. S. Yun, S. Aliwell, P. P. Altermatt, T. Ameri, V. Andrei, A. Armin, D. Bag- nis, J. Baker, H. Beath, M. Bellanger, P. Berrouard, J. Blumberger, S. A. Boden, H. Bronstein, M. J. Carnie, C. Case, F. A. Cas...

  5. [5]

    Battaglia, A

    C. Battaglia, A. Cuevas, and S. D. Wolf, Energy & En- vironmental Science9, 1552 (2016)

  6. [6]

    Sayed, A

    H. Sayed, A. M. Ahmed, A. Hajjiah, M. A. Abdelkawy, and A. H. Aly, Scientific Reports15, 16529 (2025)

  7. [7]

    Ramanujam and U

    J. Ramanujam and U. P. Singh, Energy & Environmental Science 10, 1306 (2017)

  8. [8]

    M. A. Scarpulla, B. McCandless, A. B. Phillips, Y. Yan, M. J. Heben, C. Wolden, G. Xiong, W. K. Metzger, D. Mao, D. Krasikov, I. Sankin, S. Grover, A. Munshi, W. Sampath, J. R. Sites, A. Bothwell, D. Albin, M. O. Reese, A. Romeo, M. Nardone, R. Klie, J. M. Walls, T. Fiducia, A. Abbas, and S. M. Hayes, Solar Energy Materials and Solar Cells255, 112289 (2023)

Show all 72 references
  1. [9]

    E. K. Solak and E. Irmak, RSC Advances 13, 12244 (2023)

  2. [10]

    O’Regan and M

    B. O’Regan and M. Grätzel, Nature353, 737 (1991)

  3. [11]

    M. A. Green, A. Ho-Baillie, and H. J. Snaith, Nature Photonics 8, 506 (2014)

  4. [12]

    Zakutayev, J

    A. Zakutayev, J. D. Major, X. Hao, A. Walsh, J. Tang, T. K. Todorov, L. H. Wong, and E. Saucedo, Journal of Physics: Energy3, 032003 (2021)

  5. [13]

    D. B. Needleman, J. R. Poindexter, R. C. Kurchin, I. M. Peters, G. Wilson, and T. Buonassisi, Energy & Environ- mental Science9, 2122 (2016)

  6. [14]

    S. Y. Yang, J. Seidel, S. J. Byrnes, P. Shafer, C.-H. Yang, M. D. Rossell, P. Yu, Y.-H. Chu, J. F. Scott, J. W. Ager, L. W. Martin, and R. Ramesh, Nature Nanotechnology 5, 143 (2010)

  7. [15]

    Wu and Y

    L. Wu and Y. Yang, Advanced Materials Interfaces9, 2201415 (2022)

  8. [16]

    A. W. Welch, L. L. Baranowski, H. Peng, H. Hempel, R. Eichberger, T. Unold, S. Lany, C. Wolden, and A. Za- kutayev, Advanced Energy Materials7, 1601935 (2017)

  9. [17]

    Yadav, R

    S. Yadav, R. K. Chauhan, and R. Mishra, Renewable Energy 255, 123810 (2025)

  10. [18]

    A. A. Ahmad, A. B. Migdadi, A. M. Alsaad, I. A. Qat- tan, Q. M. Al-Bataineh, and A. Telfah, Heliyon8, e08683 (2022)

  11. [19]

    Vidal, S

    J. Vidal, S. Lany, M. d’Avezac, A. Zunger, A. Zaku- tayev, J. Francis, and J. Tate, Applied Physics Letters 100, 032104 (2012)

  12. [20]

    A. M. Ganose, K. T. Butler, A. Walsh, and D. O. Scan- lon, Journal of Materials Chemistry A4, 2060 (2016)

  13. [21]

    X. Wang, S. R. Kavanagh, D. O. Scanlon, and A. Walsh, Joule 8, 2105 (2024)

  14. [22]

    Yang, W.-J

    J.-H. Yang, W.-J. Yin, J.-S. Park, J. Ma, and S.-H. Wei, Semiconductor Science and Technology31, 083002 (2016)

  15. [23]

    Z. Yuan, D. Dahliah, M. R. Hasan, G. Kassa, A. Pike, S.Quadir, R.Claes, C.Chandler, Y.Xiong, V.Kyveryga, P. Yox, G.-M. Rignanese, I. Dabo, A. Zakutayev, D. P. Fenning, O. G. Reid, S. Bauers, J. Liu, K. Kovnir, and G. Hautier, Joule8, 1412 (2024)

  16. [24]

    J. P. Perdew, K. Burke, and M. Ernzerhof, Physical Re- view Letters77, 3865 (1996)

  17. [25]

    Courty, V

    B. Courty, V. Schmidt, S. Luccioni, Goyal-Kamal, Mari- onCoutarel, B. Feld, J. Lecourt, LiamConnell, A. Saboni, Inimaz, supatomic, M. Léval, L. Blanche, A. Cru- veiller, ouminasara, F. Zhao, A. Joshi, A. Bogroff, H. d. Lavoreille, N. Laskaris, E. Abati, D. Blank, Z. Wang, A. C...

  18. [26]

    Shockley and H

    W. Shockley and H. J. Queisser, Journal of Applied Physics 32, 510 (1961)

  19. [27]

    Yu and A

    L. Yu and A. Zunger, Physical Review Letters 108, 068701 (2012)

  20. [28]

    J. Heyd, G. E. Scuseria, and M. Ernzerhof, The Journal of Chemical Physics118, 8207 (2003)

  21. [29]

    R. X. Yang, M. K. Horton, J. Munro, and K. A. Persson, High-throughput optical absorption spectra for inorganic semiconductors (2022), arXiv preprint arXiv:2209.02918

  22. [30]

    De Angelis, ACS Energy Letters8, 1270 (2023)

    F. De Angelis, ACS Energy Letters8, 1270 (2023)

  23. [31]

    Lany, Nature Computational Science3, 675 (2023)

    M.D.Witman, A.Goyal, T.Ogitsu, A.H.McDaniel,and S. Lany, Nature Computational Science3, 675 (2023)

  24. [32]

    Woods-Robinson, Y

    R. Woods-Robinson, Y. Xiong, J.-X. Shen, N. Winner, M. K. Horton, M. Asta, A. M. Ganose, G. Hautier, and K. A. Persson, Matter6, 3021 (2023)

  25. [33]

    T. W. Hesterberg, C. A. Lapin, and W. B. Bunn, Envi- ronmental Science & Technology42, 6437 (2008)

  26. [34]

    Banko and E

    M. Banko and E. Brill, inProceedings of the 39th An- nual Meeting on Association for Computational Linguis- tics, ACL ’01 (Association for Computational Linguis- tics, USA, 2001) pp. 26–33

  27. [35]

    Amodei, R

    D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos, E. Elsen, J. Engel, L. Fan, C. Fougner, T. Han, A. Hannun, B. Jun, P. LeGres- ley, L. Lin, S. Narang, A. Ng, S. Ozair, R. Prenger, J. Raiman, S. Satheesh, D. ...

  28. [36]

    Hestness, S

    J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou, DeepLearningScalingisPredictable, Empirically(2017)

  29. [37]

    C. Sun, A. Shrivastava, S. Singh, and A. Gupta, Revisit- ing Unreasonable Effectiveness of Data in Deep Learning Era (2017), arXiv preprint arXiv.1707.02968

  30. [38]

    Bailly, C

    A. Bailly, C. Blanc, É. Francis, T. Guillotin, F. Jamal, B.Wakim,andP.Roy,ComputerMethodsandPrograms in Biomedicine213, 106504 (2022)

  31. [39]

    Alampara, M

    N. Alampara, M. Schilling-Wilhelmi, and K. M. Jablonka, Lessons from the trenches on evaluating machine-learning systems in materials science (2025), arXiv preprint:arXiv.2503.10837

  32. [40]

    A. M. Ganose, H. Sahasrabuddhe, M. Asta, K. Beck, T. Biswas, A. Bonkowski, J. Bustamante, X. Chen, Y. Chiang, D. C. Chrzan, J. Clary, O. A. Cohen, C. Ertu- ral, M. C. Gallant, J. George, S. Gerits, R. E. A. Goodall, R. D. Guha, G. Hautier, M. Horton, T. J. Inizan, A. D. Kaplan...

  33. [41]

    Choudhary and B

    K. Choudhary and B. DeCost, npj Computational Ma- terials 7, 1 (2021)

  34. [42]

    Choudhary, Q

    K. Choudhary, Q. Zhang, A. C. E. Reid, S. Chowdhury, N. Van Nguyen, Z. Trautt, M. W. Newrock, F. Y. Congo, and F. Tavazza, Scientific Data5, 180082 (2018)

  35. [43]

    D. H. Fabini, M. Koerner, and R. Seshadri, Chemistry of Materials 31, 1561 (2019)

  36. [44]

    M. K. Y. Chan and G. Ceder, Physical Review Letters 105, 196403 (2010)

  37. [45]

    Tran and P

    F. Tran and P. Blaha, Physical Review Letters 102, 226401 (2009)

  38. [46]

    Geiger and T

    M. Geiger and T. Smidt, e3nn: Euclidean Neural Net- works (2022)

  39. [47]

    Geiger, T

    M. Geiger, T. Smidt, A. M, B. K. Miller, W. Boomsma, B. Dice, K. Lapchevskyi, M. Weiler, M. Tyszkiewicz, S. Batzner, D. Madisetti, M. Uhrin, J. Frellsen, N. Jung, S. Sanborn, M. Wen, J. Rackers, M. Rød, and M. Bailey, Euclidean neural networks: e3nn (2022)

  40. [48]

    Thomas, T

    N. Thomas, T. Smidt, S. Kearnes, L. Yang, L. Li, K. Kohlhoff, and P. Riley, Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds (2018), _eprint: 1802.08219

  41. [49]

    Weiler, M

    M. Weiler, M. Geiger, M. Welling, W. Boomsma, and T. Cohen, 3D Steerable CNNs: Learning Rotationally EquivariantFeaturesinVolumetricData(2018),_eprint: 1807.02547

  42. [50]

    Kondor, Z

    R. Kondor, Z. Lin, and S. Trivedi, Clebsch-Gordan Nets: a Fully Fourier Space Spherical Convolutional Neural Network (2018), _eprint: 1806.09231

  43. [51]

    Batatia, S

    I. Batatia, S. Batzner, D. P. Kovács, A. Musaelian, G. N. C. Simm, R. Drautz, C. Ortner, B. Kozinsky, and G. Csányi, Nature Machine Intelligence7, 56 (2025)

  44. [52]

    Batzner, A

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, Nature Communications13, 2453 (2022)

  45. [53]

    Musaelian, S

    A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Nature Commu- nications 14, 579 (2023)

  46. [54]

    R. Ruff, P. Reiser, J. Stühmer, and P. Friederich, Digital Discovery 3, 594 (2024)

  47. [55]

    Grunert, M

    M. Grunert, M. Großmann, and E. Runge, Physical Re- view Materials8, L122201 (2024)

  48. [56]

    N. T. Hung, R. Okabe, A. Chotrattanapituk, and M. Li, Advanced Materials36, 2409175 (2024)

  49. [57]

    Sbailò, Á

    L. Sbailò, Á. Fekete, L. M. Ghiringhelli, and M. Scheffler, npj Computational Materials8, 250 (2022)

  50. [58]

    P.Ong, G.Hautier, W.Chen, W.D.Richards, S

    A.Jain, S. P.Ong, G.Hautier, W.Chen, W.D.Richards, S. Dacek, S. Cholia, D. Gunter, D. Skinner, G. Ceder, and K. A. Persson, APL Materials1, 011002 (2013)

  51. [59]

    M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, J. Bouw- man, A. J. Brookes, T. Clark, M. Crosas, I. Dillo, O. Dumon, S. Edmunds, C. T. Evelo, R. Finkers, A.Gonzalez-Beltran, A.J....

  52. [60]

    C. Fare, P. Fenner, M. Benatan, A. Varsi, and E. O. Pyzer-Knapp, npj Computational Materials 8, 257 (2022)

  53. [61]

    Hoffmann, J

    N. Hoffmann, J. Schmidt, S. Botti, and M. A. L. Mar- ques, Digital Discovery2, 1368 (2023). 12

  54. [62]

    H. Kaur, F. D. Pia, I. Batatia, X. R. Advincula, B. X. Shi, J. Lan, G. Csányi, A. Michaelides, and V. Kapil, Faraday Discussions256, 120 (2025)

  55. [63]

    Mitchell, S

    M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasser- man, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru (2019) pp. 220–229

  56. [64]

    Blank, T

    B. Blank, T. Kirchartz, S. Lany, and U. Rau, Physical Review Applied8, 024032 (2017)

  57. [65]

    D. P. Kingma and M. Welling, Auto-Encoding Varia- tional Bayes (2022), arXiv preprint: arXiv.1312.6114

  58. [66]

    P. R. Kaundinya, K. Choudhary, and S. R. Kalidindi, JOM 74, 1395 (2022)

  59. [67]

    S. Kim, M. Lee, C. Hong, Y. Yoon, H. An, D. Lee, W. Jeong, D. Yoo, Y. Kang, Y. Youn, and S. Han, Sci- entific Data7, 387 (2020)

  60. [68]

    Kresse and J

    G. Kresse and J. Hafner, Journal of Physics: Condensed Matter 6, 8245 (1994)

  61. [69]

    Kresse and D

    G. Kresse and D. Joubert, Physical Review B59, 1758 (1999)

  62. [70]

    Kresse and J

    G. Kresse and J. Hafner, Physical Review B 47, 558 (1993)

  63. [71]

    Kresse and J

    G. Kresse and J. Furthmüller, Computational Materials Science 6, 15 (1996)

  64. [72]

    Kresse and J

    G. Kresse and J. Furthmüller, Physical Review B54, 11169 (1996)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.