Pith. sign in

REVIEW 3 major objections 5 minor 27 references

Predicting large-supercell defect formation energies from machine-learning charge density models trained on small supercells

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Machine learning on charge density from 96 small supercells predicts defect formation energies in 360-atom supercells with mean absolute error below 0.05 eV.

desk verdict Useful mixed-size training idea for MLCD defect prediction, but the headline MAE and the cost claim both overstate what the body actually shows. read the letter →

arxiv 2608.07997 v1 pith:I2OF5R6E submitted 2026-08-08 cond-mat.mtrl-sci

classification cond-mat.mtrl-sci
keywords defectformationenergymachinelearningchargedensitygalliumnitridesupercellsizeconvergencemixed-sizetrainingequivariantneuralnetworkNSCFcalculationcross-sizetransfer
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a machine-learning model trained on the charge density of small defect supercells can predict defect formation energies in target supercells much larger than anything in the training set, provided the training cells span a range of sizes. For four intrinsic defects in GaN, a model trained on only 96 structures containing 16–96 atoms reaches a mean absolute formation-energy error below 0.05 eV in 360-atom supercells. The authors argue this works because real-space charge density is a local, transferable quantity, whereas direct energy–force potentials trained on the same data err by more than 1 eV. If right, the result would cut the data-generation cost of first-principles defect calculations by roughly two-thirds relative to single-size training, and it suggests a general recipe for extrapolating supercell-size-dependent properties from mixed-size datasets.

What carries the argument

The central object is the real-space charge density, treated as a transferable local descriptor. The mechanism is a three-part pipeline: an equivariant atom-probe network (EAC-Net) is trained on perturbed small-supercell DFT charge densities; the trained model predicts the charge density of the large target supercell; a non-self-consistent DFT calculation using that density yields total energies and derived formation energies. The data design is governed by a criterion on the defect influence radius, measured by differential charge density and structural similarity, and by a three-region partition of each supercell into defect, transition, and bulk-like environments. This combination lets small cells supply configurational diversity while a few larger cells act as anchors that constrain the extrapolation along the supercell-size axis.

What would settle it

Run the same 96-structure mixed-size MLCD protocol on a GaN defect whose influence radius clearly exceeds 5.62 Å (for example a charged vacancy or an extended interstitial), with no training supercell larger than 96 atoms, and compare against 360-atom DFT formation energies: if the mean absolute error stays below 0.05 eV the cross-size claim survives, and it fails if the error rises well above that threshold.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that machine-learning charge density (MLCD) can serve as a cross-size transfer route for defect calculations. The model, trained on DFT charge densities of small supercells, predicts the charge density of a 360-atom supercell, from which total energies and formation energies are obtained by a non-self-consistent-field DFT step. A mixed-size training set containing 16-, 24-, 32-, 64-, and 96-atom supercells reaches the 0.05 eV target accuracy with 96 structures, whereas training on a single 96-atom size requires about 240 structures and training on 16–32-atom cells alone saturates around 0.1 eV even with 720 structures. The authors identify the controlling length scale: half of the training-cell lattice length must be comparable to or larger than the defect influence radius, obtained from differential charge-density analysis. The same predicted densities also reproduce band structures, defect levels, and spin polarization in the large cells.

Load-bearing premise

The method's data efficiency rests on knowing the defect influence radius from DFT charge densities in the 360-atom target supercells; if that large-cell information is unavailable, the minimum training size cannot be set by the proposed criterion.

Editorial extensions

If this is right

  • With 96 mixed-size structures, MLCD reaches formation-energy errors below 0.05 eV for four intrinsic defects in 360-atom GaN supercells, cutting the dataset size by about two-thirds and the total DFT preparation cost by a factor of 6.2 relative to single-size 96-atom training.
  • The half-lattice-length criterion gives a quantitative rule for when a training supercell size is adequate: half of its lattice vector must be at least the defect influence radius.
  • MLIPs trained on the same datasets err by more than 1 eV, indicating that direct energy-force fitting lacks the local transferability that charge-density learning provides.
  • MLCD-predicted densities, through NSCF calculations, reproduce band structures, defect levels, and spin polarization of the large supercells, so formation energies are not the only accessible property.
  • Mixed-size training is framed as a continuation problem, where intermediate-size supercells anchor the size-dependent property curve and reduce extrapolation uncertainty.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The continuation-by-anchors view suggests a practical shortcut: estimate the defect influence radius from the size-dependent curve itself, using small- and medium-cell data, which would remove the current need for target-scale DFT charge densities to set the training-size criterion.
  • The strategy may generalize beyond point defects to other size-convergent properties such as charged-defect transition levels, alloy disorder, or interface segregation energies, where mixed-size anchors could play the same role.
  • Because the study covers four neutral intrinsic defects in GaN, the robustness of the claim for charged defects or defects with long-range strain fields is untested; those cases are natural stress tests for the 0.05 eV target.
  • The paper's explanation ties cross-size transfer to the spatially resolved local output of MLCD, implying that any model with a similar local-output design should show comparable robustness, not only the specific EAC-Net architecture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes a machine-learning charge-density (MLCD) workflow, based on the EAC-Net model, that is trained on defect-containing GaN supercells of 16–96 atoms and then used to predict charge densities for 360-atom defect supercells; total energies and defect formation energies are obtained from non-self-consistent-field (NSCF) DFT calculations using the predicted densities. A mixed-size training strategy is introduced, with the mix of supercell sizes selected using defect influence radii and structural-similarity measures. The authors claim that 96 configurations suffice to reach a mean formation-energy error below 0.05 eV for four intrinsic defects, that this is a 6.2× cost reduction relative to using only 96-atom training cells, and that an Allegro MLIP trained on the same data is far less accurate.

Significance. If the claims hold, the work is a useful demonstration that charge-density learning can transfer across supercell sizes more robustly than direct energy-force fitting, and the explicit length-scale criterion (half the training-cell lattice length should cover the defect influence radius) is a physically motivated and potentially reusable design heuristic. The paper benefits from a transparent dataset budget table, a comparison against an MLIP baseline, and repeated-run statistics for at least one model. However, the central practical claim of data efficiency is weakened by the fact that the training-set design relies on 360-atom DFT charge densities from the target supercells, which are not included in the cost accounting, and the abstract overstates the per-defect accuracy relative to the body of the paper.

major comments (3)
  1. [Abstract; Section 2.3, Fig. 3(e)(f)] The abstract claims 'defect-wise mean absolute error below 0.05 eV', but the body states in Fig. 3(e) that γ(16–96) keeps the MAE of every defect within 0.1 eV, and Fig. 3(f) reports the average formation-energy error dropping below 0.05 eV. These are different statements, and the abstract overstates the result. Please revise the abstract to say that the average error is below 0.05 eV and that per-defect errors remain within 0.1 eV.
  2. [Section 2.3, Fig. 3(a)(b); Fig. 4(g)] The data-efficiency claim omits a substantial part of the required DFT cost. The defect influence radii (V_N ≈ 3.25 Å, Ga_N ≈ 5.62 Å) are extracted from differential charge densities computed in the 360-atom supercells (Fig. 3(a)), and the structural-similarity criterion in Eqs. (3)–(6) uses the 360-atom structure as the reference (Fig. 3(b)). Since a self-consistent DFT calculation yields both the charge density and the total energy, applying the design protocol as described requires essentially the target-scale DFT calculations that the method aims to avoid. The 'only 96 supercells' claim in the abstract and the 6.2× cost reduction in Fig. 4(g) therefore need to be re-examined; either the 360-atom design calculations should be included in the cost, or the authors should show that the influence radii and similarity criterion can be obtained from smaller-supercell data or prior knowledge.
  3. [Section 2.3, Fig. 3(f); Section 2.1, Fig. 2(c)] The headline result that γ(16–96) reaches ΔE_form < 0.05 eV with 96 structures is reported as a single point estimate. The authors themselves show in Fig. 2(c) that the random grid-point sampling used for training introduces run-to-run scatter across 12 independent runs for the β(96) model. Equivalent multiple-seed statistics should be reported for γ(16–96), because the target accuracy of 0.05 eV is comparable to the stochastic spread shown for the other models, and a single run is not sufficient to establish that the target is reached.
minor comments (5)
  1. [References, Ref. [18]] The DOI '10.1103/h66h-y5k6' for Ref. [18] appears malformed or is a placeholder; please verify and provide the correct DOI.
  2. [Fig. 4(g)] The y-axis label 'Time of preparation' is vague; please specify whether this is DFT data-generation cost, model training cost, or total wall-clock time, and state the units.
  3. [Section 5.3] The phrase 'a×N_k ≈ 30–40 Å' is dimensionally unusual; if this means the k-point density is approximately 30–40 Å, the notation should be clarified, for example by writing the k-point spacing or the number of k-points along each direction.
  4. [Section 2.4] The claim that the single-particle energy error near the VBM is 7×10^(−5) eV for bulk is not tied to a specific figure or table; please indicate where this number is shown or add a supplementary figure.
  5. [Data availability] The data availability statement says data are available 'upon reasonable request'; given the paper's emphasis on data efficiency and reproducibility, providing the training and evaluation datasets in a repository would strengthen the manuscript.

Circularity Check

1 steps flagged · score 3.0 of 10

Data-efficiency claim rests on 360-atom DFT charge densities used to set the mixed-size training schedule; those target-scale calculations are omitted from the 96-structure and 6.2x cost accounting.

  1. other [Section 2.3, Fig. 3(a)-(b), Eqs. (5)-(6); cost claim in Discussion and Abstract]
    "Figure 3(a) shows DFT differential charge densities of VN and GaN relative to bulk in 360-atom supercells. ... We take the 360-atom structure as the reference, where ε_R,A = 1 denotes identical structures. ... This leads to a key criterion: half of the training-cell lattice length, ā/2, must be comparable to or larger than the defect influence radius. ... At fixed accuracy, this mixed-size design reduces the computational cost by a factor of 6.2 relative to training with a single 96-atom supercell size."

    The influence radii (3.25 Å for VN, 5.62 Å for GaN) and structural-similarity thresholds that decide which supercell sizes are included in γ(16–96) are extracted from SCF DFT charge densities and relaxed geometries of the 360-atom target supercells. In a plane-wave DFT code, one SCF run yields both the charge density and the total energy, so these design inputs are produced by the same large-supercell DFT calculation whose formation energy the workflow claims to avoid. The reported 'only 96 supercells' dataset count and the 6.2x cost reduction exclude these target-scale SCF calculations.

full rationale

The core ML accuracy claim is not circular: the MLCD model is trained on small-supercell charge densities (16–96 atoms) and the 360-atom formation energies are obtained from NSCF calculations on MLCD-predicted densities, not used as training targets. The error against 360-atom DFT is a genuine external benchmark. No load-bearing self-citation was found; EAC-Net [15] is third-party code and the one author-overlap citation [10] is incidental background. The circularity found is limited to the data-efficiency/cost claim: the mixed-size training schedule is selected using differential charge densities from the 360-atom target supercells (Fig. 3a) and a structural-similarity measure referenced to the 360-atom structure (Fig. 3b, Eqs. 5–6). Since an SCF DFT calculation on the target supercell yields both the charge density used for this design and the total energy needed for formation energies, the 'only 96 supercells' and '6.2x reduction' statements understate the large-supercell DFT effort actually required by the method as described. This is a partial, cost-accounting circularity, not a collapse of the predictive model. Score 3 reflects one mild but identifiable use of target-scale information to shape the input design.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The method introduces no new physical entities. The central claim depends on several chosen hyperparameters and, more importantly, on influence radii extracted from the 360-atom target calculations, which couples the training-set design to the very large-scale calculations the method aims to avoid.

free parameters (3)
  • Defect influence radii = 3.25 Angstrom (V_N); 5.62 Angstrom (Ga_N)
    Derived from differential charge densities in 360-atom target supercells (Fig. 3a) and used to select which supercell sizes to include in training. This uses target-scale data to design the training set.
  • EAC-Net atomic and probe cutoff radius = 6.0 Angstrom
    Chosen hyperparameter; the transferability claim depends on the charge density being local within this cutoff.
  • Perturbation displacement range R_disp = 0 to 0.5 Angstrom
    Data augmentation amplitude chosen for generating training configurations; affects coverage of the potential-energy surface.
assumptions (3)
  • domain assumption Charge density is a local function of the atomic environment within the 6.0 Angstrom cutoff.
    EAC-Net approximates density from local atomic and probe features; the cross-size transfer result depends on this locality.
  • domain assumption Total energy from an NSCF calculation at a predicted charge density accurately reproduces the self-consistent total energy for defect supercells.
    The workflow computes formation energies from NSCF energies using the MLCD density; the error analysis assumes this approximation is valid.
  • ad hoc to paper Structural similarity (RDF/ADF overlap) is a valid proxy for charge-density transferability across supercell sizes.
    The paper uses epsilon_R,A to conclude when small-cell local structures match the 360-atom reference; this metric is introduced here without independent validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Predicting large-supercell defect formation energies from machine-learning charge density models trained on small supercells." pith.science (2026). https://pith.science/paper/I2OF5R6E

@misc{pith2026260807997,
  author       = {Pith},
  title        = {Pith review of: Predicting large-supercell defect formation energies from machine-learning charge density models trained on small supercells},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I2OF5R6E}},
  note         = {Machine review of arXiv:2608.07997}
}
read the original abstract

First-principles defect calculations are often limited by the cost of the large supercells required to suppress image interactions. Machine-learning interatomic potentials (MLIPs) provide another alternative, but training defect MLIPs typically requires thousands of structures and weeks of data generation. Since charge density is the key to density-functional-theory (DFT), we propose a machine-learning charge density (MLCD) route for predicting defect formation energies with higher data efficiency. We optimize the training set by integrating small supercells of varying sizes for better extrapolation, allocating their proportions based on spatial charge-density analysis. With only 96 supercells containing 16--96 atoms as the dataset, MLCD accurately predicts the formation energies of four intrinsic defects in 360-atom supercells, with defect-wise mean absolute error below 0.05 eV. In contrast, MLIPs trained on the same dataset can err by more than 1 eV. These results show that charge-density learning enables more robust cross-size transfer than direct energy-force fitting and that mixed-size data design can substantially reduce the cost of defect prediction.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 11 canonical work pages

  1. [1]

    Freysoldt, C.et al.First-principles calculations for point defects in solids.Rev. Mod. Phys.86, 253–305 (2014). URL https://link.aps.org/doi/10.1103/RevModPhys.86.253

  2. [2]

    Alkauskas, A., McCluskey, M. D. & Van de Walle, C. G. Tutorial: defects in semiconductors— combining experiment and theory.J. Appl. Phys.119, 181101 (2016). URL https://doi.org/10. 1063/1.4948245

  3. [3]

    Overcoming the doping bottleneck in semiconductors.Comput

    Wei, S.-H. Overcoming the doping bottleneck in semiconductors.Comput. Mater. Sci.30, 337–348 (2004). URL https://www.sciencedirect.com/science/article/pii/S092702560400117X

  4. [4]

    & Zhang, S

    Wei, S.-H. & Zhang, S. B. Chemical trends of defect formation and doping limit in II-VI semicon- ductors: the case of CdTe.Phys. Rev. B66, 155211 (2002). URL https://link.aps.org/doi/10.1103/ PhysRevB.66.155211

  5. [5]

    Van de Walle, C. G. & Neugebauer, J. First-principles calculations for defects and impurities: applications to III-nitrides.J. Appl. Phys.95, 3851–3879 (2004). URL https://doi.org/10.1063/1. 1682673

  6. [6]

    & Van de Walle, C

    Freysoldt, C., Neugebauer, J. & Van de Walle, C. G. Fully ab initio finite-size corrections for charged- defect supercell calculations.Phys. Rev. Lett.102, 016402 (2009). URL https://link.aps.org/doi/ 10.1103/PhysRevLett.102.016402

  7. [7]

    & Furthm¨ uller, J

    Kresse, G. & Furthm¨ uller, J. Efficient iterative schemes for ab initio total-energy calculations using a plane-wave basis set.Phys. Rev. B54, 11169–11186 (1996). URL https://link.aps.org/doi/10. 1103/PhysRevB.54.11169

  8. [8]

    & Ong, S

    Chen, C. & Ong, S. P. A universal graph deep learning interatomic potential for the periodic table. Nat. Comput. Sci.2, 718–728 (2022). URL https://doi.org/10.1038/s43588-022-00349-3

Show all 27 references
  1. [9]

    Deng, B.et al.CHGNet as a pretrained universal neural network potential for charge-informed atomistic modelling.Nat. Mach. Intell.5, 1031–1041 (2023). URL https://doi.org/10.1038/ s42256-023-00716-3

  2. [10]

    & Kavanagh, S

    Mannodi-Kanakkithodi, A., Huang, M., Gorai, P. & Kavanagh, S. R. Accelerating point-defect simulations using data-driven and machine learning approaches.MRS Bull.51, 600–614 (2026). URL https://doi.org/10.1557/s43577-026-01103-0

  3. [11]

    Commun.13, 2453 (2022)

    Batzner, S.et al.E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials.Nat. Commun.13, 2453 (2022). URL https://doi.org/10.1038/s41467-022-29939-5

  4. [12]

    P., Simm, G

    Batatia, I., Kov´ acs, D. P., Simm, G. N. C., Ortner, C. & Cs´ anyi, G. MACE: higher order equivariant message passing neural networks for fast and accurate force fields (2023). URL https://arxiv.org/ abs/2206.07697. arXiv:2206.07697. 9

  5. [13]

    & Corminboeuf, C

    Fabrizio, A., Grisafi, A., Meyer, B., Ceriotti, M. & Corminboeuf, C. Electron density learning of non-covalent systems.Chem. Sci.10, 9424–9432 (2019). URL https://doi.org/10.1039/c9sc02696g

  6. [14]

    Koker, T., Quigley, K., Taw, E., Tibbetts, K. & Li, L. Higher-order equivariant neural networks for charge density prediction in materials.npj Comput. Mater.10, 161 (2024). URL https://doi.org/ 10.1038/s41524-024-01343-1

  7. [15]

    & Zhong, Z

    Qin, X., Lv, T. & Zhong, Z. EAC-Net: predicting real-space charge density via equivariant atomic contributions.J. Chem. Theory Comput.22, 4813–4821 (2026). URL https://doi.org/10.1021/acs. jctc.6c00283

  8. [16]

    Li, H.et al.Deep-learning density functional theory Hamiltonian for efficient ab initio electronic- structure calculation.Nat. Comput. Sci.2, 367–377 (2022). URL https://doi.org/10.1038/ s43588-022-00265-6

  9. [17]

    & Xiang, H

    Zhong, Y., Yu, H., Su, M., Gong, X. & Xiang, H. Transferable equivariant graph neural networks for the Hamiltonians of molecules and solids.npj Comput. Mater.9, 182 (2023). URL https: //doi.org/10.1038/s41524-023-01130-4

  10. [18]

    & Kumagai, Y

    Kiyohara, S., Shibui, C., Bae, S. & Kumagai, Y. Machine-learning prediction of charged-defect formation energies from crystal structures.Phys. Rev. Lett.135, 246101 (2025). URL https://link. aps.org/doi/10.1103/h66h-y5k6

  11. [19]

    Dutta, S.et al.Characterizing defect dynamics in silicon carbide using symmetry-adapted collective variables and machine learning interatomic potentials.J. Chem. Theory Comput.22, 4728–4741 (2026). URL https://doi.org/10.1021/acs.jctc.6c00063

  12. [20]

    Shimizu, K.et al.Using neural network potentials to study defect formation and phonon properties of nitrogen vacancies with multiple charge states in GaN.Phys. Rev. B106, 054108 (2022). URL https://link.aps.org/doi/10.1103/PhysRevB.106.054108

  13. [21]

    & Erhart, P

    Linder¨ alv, C., ¨Osterbacka, N., Wiktor, J. & Erhart, P. Optical line shapes of color centers in solids from classical autocorrelation functions.npj Comput. Mater.11, 101 (2025). URL https: //doi.org/10.1038/s41524-025-01565-x

  14. [22]

    & Kohn, W

    Hohenberg, P. & Kohn, W. Inhomogeneous electron gas.Phys. Rev.136, B864–B871 (1964). URL https://link.aps.org/doi/10.1103/PhysRev.136.B864

  15. [23]

    & Sham, L

    Kohn, W. & Sham, L. J. Self-consistent equations including exchange and correlation effects.Phys. Rev.140, A1133–A1138 (1965). URL https://link.aps.org/doi/10.1103/PhysRev.140.A1133

  16. [24]

    & Furthm¨ uller, J

    Kresse, G. & Furthm¨ uller, J. Efficiency of ab-initio total energy calculations for metals and semi- conductors using a plane-wave basis set.Comput. Mater. Sci.6, 15–50 (1996). URL https: //www.sciencedirect.com/science/article/pii/0927025696000080

  17. [25]

    & Joubert, D

    Kresse, G. & Joubert, D. From ultrasoft pseudopotentials to the projector augmented-wave method. Phys. Rev. B59, 1758–1775 (1999). URL https://link.aps.org/doi/10.1103/PhysRevB.59.1758

  18. [26]

    P., Burke, K

    Perdew, J. P., Burke, K. & Ernzerhof, M. Generalized gradient approximation made simple.Phys. Rev. Lett.77, 3865–3868 (1996). URL https://link.aps.org/doi/10.1103/PhysRevLett.77.3865

  19. [27]

    Bl¨ ochl, P. E. Projector augmented-wave method.Phys. Rev. B50, 17953–17979 (1994). URL https://link.aps.org/doi/10.1103/PhysRevB.50.17953. 10 T able 1Configuration datasets used for MLCD model training. Model name Defect supercell Natom Configs. per Natom per defect Total def...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.