{"id":"1f2d6ac6-4eab-474b-867d-4cdf97ba71f4","arxiv_id":"2608.07997","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A charge-density machine-learning model trained on 96 small mixed-size supercells predicts large-supercell GaN defect formation energies with about 0.05 eV average error, outperforming energy-force MLIPs on the same data.","lead":"A machine-learning model that predicts electron charge density from atomic structure can estimate defect formation energies in large supercells using only small-supercell training data, with reported errors around 0.05 eV. The approach could cut the cost of semiconductor defect calculations by a factor of six if the results hold.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed data-efficiency gain rests on 360-atom DFT charge densities used to choose the mixed-size training schedule; those target-scale calculations are omitted from the cost accounting.","rationale":"Agree with the reader: the weakest point is the use of 360-atom DFT charge densities to set the training-size design. This is load-bearing because the paper's central contribution is data efficiency, not just predictive accuracy. Without the target-scale influence radii, a user would not know that 96-atom cells are needed for Ga_N and N_Ga while 32-atom cells suffice for V_N, and the mixed-size strategy shown in Fig. 3(d)-(f) would not be uniquely determined. Costing only the 96 small supercells while ignoring the large-cell SCF calculations used in Fig. 3(a)-(b) makes the 6.2x claim incomplete. The proposed blind-retraining test would settle whether the target-scale information is actually required. I also note the reader's secondary point about the abstract's 'defect-wise MAE < 0.05 eV' versus the body's per-defect values within 0.1 eV; this is a reporting inconsistency, but the influence-radius issue is the more fundamental reason to keep the verdict conditional rather than accept.","tokens_in":11008,"tokens_out":6927,"duration_ms":71204,"concrete_test":"Retrain γ(16–96) with a blinded size-selection protocol: choose the training sizes using only supercells up to 96 atoms, e.g., estimate influence radii from 96-atom differential charge densities or from convergence of small-cell formation energies, and do not use any 360-atom DFT geometry, density, RDF/ADF, or energy. If the blinded model still reaches 360-atom average |ΔE_form| < 0.05 eV, the target-scale design information is not load-bearing. If it does not, the paper should restate the cost claim to include the 360-atom SCF calculations required to determine the influence radii.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The mixed-size schedule that produces the headline 96-structure result is chosen using information from the target supercells. Section 2.3 obtains the defect influence radii (V_N ≈ 3.25 Å, Ga_N ≈ 5.62 Å) from DFT differential charge densities in the 360-atom supercells (Fig. 3a), and Eq. (3)–(6) defines the structural-similarity criterion against the 360-atom reference (Fig. 3b). The rule that half the training-cell lattice constant must reach the defect influence radius then decides which sizes are included in γ(16–96). This means the design protocol already requires SCF-level DFT charge densities and relaxed geometries for the large supercell. Because a single SCF calculation yields both the charge density and the total energy, obtaining these design inputs is essentially performing the target-scale DFT calculation the method aims to avoid. The abstract's 'only 96 supercells containing 16–96 atoms' and the 6.2x cost reduction in Fig. 4(g) therefore undercount the large-supercell DFT effort needed to apply the method as described. The accuracy of the MLCD/NSCF pipeline is not at issue; the practical data-efficiency claim is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a machine-learning charge-density (MLCD) workflow, based on the EAC-Net model, that is trained on defect-containing GaN supercells of 16–96 atoms and then used to predict charge densities for 360-atom defect supercells; total energies and defect formation energies are obtained from non-self-consistent-field (NSCF) DFT calculations using the predicted densities. A mixed-size training strategy is introduced, with the mix of supercell sizes selected using defect influence radii and structural-similarity measures. The authors claim that 96 configurations suffice to reach a mean formation-energy error below 0.05 eV for four intrinsic defects, that this is a 6.2× cost reduction relative to using only 96-atom training cells, and that an Allegro MLIP trained on the same data is far less accurate.","tokens_in":11279,"tokens_out":4676,"duration_ms":51862,"significance":"If the claims hold, the work is a useful demonstration that charge-density learning can transfer across supercell sizes more robustly than direct energy-force fitting, and the explicit length-scale criterion (half the training-cell lattice length should cover the defect influence radius) is a physically motivated and potentially reusable design heuristic. The paper benefits from a transparent dataset budget table, a comparison against an MLIP baseline, and repeated-run statistics for at least one model. However, the central practical claim of data efficiency is weakened by the fact that the training-set design relies on 360-atom DFT charge densities from the target supercells, which are not included in the cost accounting, and the abstract overstates the per-defect accuracy relative to the body of the paper.","major_comments":[{"comment":"The abstract claims 'defect-wise mean absolute error below 0.05 eV', but the body states in Fig. 3(e) that γ(16–96) keeps the MAE of every defect within 0.1 eV, and Fig. 3(f) reports the average formation-energy error dropping below 0.05 eV. These are different statements, and the abstract overstates the result. Please revise the abstract to say that the average error is below 0.05 eV and that per-defect errors remain within 0.1 eV.","section":"Abstract; Section 2.3, Fig. 3(e)(f)"},{"comment":"The data-efficiency claim omits a substantial part of the required DFT cost. The defect influence radii (V_N ≈ 3.25 Å, Ga_N ≈ 5.62 Å) are extracted from differential charge densities computed in the 360-atom supercells (Fig. 3(a)), and the structural-similarity criterion in Eqs. (3)–(6) uses the 360-atom structure as the reference (Fig. 3(b)). Since a self-consistent DFT calculation yields both the charge density and the total energy, applying the design protocol as described requires essentially the target-scale DFT calculations that the method aims to avoid. The 'only 96 supercells' claim in the abstract and the 6.2× cost reduction in Fig. 4(g) therefore need to be re-examined; either the 360-atom design calculations should be included in the cost, or the authors should show that the influence radii and similarity criterion can be obtained from smaller-supercell data or prior knowledge.","section":"Section 2.3, Fig. 3(a)(b); Fig. 4(g)"},{"comment":"The headline result that γ(16–96) reaches ΔE_form < 0.05 eV with 96 structures is reported as a single point estimate. The authors themselves show in Fig. 2(c) that the random grid-point sampling used for training introduces run-to-run scatter across 12 independent runs for the β(96) model. Equivalent multiple-seed statistics should be reported for γ(16–96), because the target accuracy of 0.05 eV is comparable to the stochastic spread shown for the other models, and a single run is not sufficient to establish that the target is reached.","section":"Section 2.3, Fig. 3(f); Section 2.1, Fig. 2(c)"}],"minor_comments":[{"comment":"The DOI '10.1103/h66h-y5k6' for Ref. [18] appears malformed or is a placeholder; please verify and provide the correct DOI.","section":"References, Ref. [18]"},{"comment":"The y-axis label 'Time of preparation' is vague; please specify whether this is DFT data-generation cost, model training cost, or total wall-clock time, and state the units.","section":"Fig. 4(g)"},{"comment":"The phrase 'a×N_k ≈ 30–40 Å' is dimensionally unusual; if this means the k-point density is approximately 30–40 Å, the notation should be clarified, for example by writing the k-point spacing or the number of k-points along each direction.","section":"Section 5.3"},{"comment":"The claim that the single-particle energy error near the VBM is 7×10^(−5) eV for bulk is not tied to a specific figure or table; please indicate where this number is shown or add a supplementary figure.","section":"Section 2.4"},{"comment":"The data availability statement says data are available 'upon reasonable request'; given the paper's emphasis on data efficiency and reproducibility, providing the training and evaluation datasets in a repository would strengthen the manuscript.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The main technical advance is plausible and the paper is within the journal's scope, but the practical data-efficiency claim should be re-evaluated after the cost-accounting issue is addressed. In particular, the authors should be asked to either include the 360-atom DFT calculations used to determine the influence radii and similarity criterion in their reported cost, or demonstrate that the training-set design can be performed without target-scale information. The abstract's error metric should also be corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper has a useful finding: charge-density learning transfers across supercell sizes better than energy/force MLIPs, and mixing small and medium supercells in the training set helps a lot. The GaN demonstration is careful in places—12 repeated runs, per-defect breakdowns, cross-checks against DFT. Second, the headline is overclaimed. The abstract says defect-wise MAE below 0.05 eV; the body says the average formation-energy error reaches under 0.05 eV while per-defect MAE stays within 0.1 eV. That is a real discrepancy and should be fixed.\n\nThe bigger problem is the cost accounting. The influence radii that decide which supercell sizes go into the training set come from differential charge densities computed in the 360-atom target supercells (Fig. 3a). So the \"only 96 supercells with 16–96 atoms\" story undercounts the actual DFT effort: you need at least a few 360-atom SCF calculations before you can design the mixed-size set. The stress-test note is correct. To be fair, one can imagine estimating a radius from smaller cells or physical intuition, but that is not what the paper does, and the data-efficiency claim as written depends on that omitted work.\n\nWhat I like: the mixed-size training as a continuation argument, and the concrete comparison with Allegro on the same data. The spin-polarization transfer for Ga_N is a nice, nontrivial result. The paper is clearly written, and the methods are described well enough to reproduce in principle. The citation list is standard and not self-inflating.\n\nWhat worries me beyond the above: data and trained models are not public (data \"available upon request\"), which is weak for a machine-learning paper. There are no error bars on the headline 0.05 eV number, only distributions. Charged defects are not addressed, and the method is demonstrated on only one material. These are limitations, not fatal flaws.\n\nAll in all, the core idea is worth refereeing. The paper deserves a serious review, but only if the authors are pushed to correct the abstract, open a defensible accounting of the large-supercell DFT used in the design, and ideally release data. As is, I would not trust the \"6.2x cost reduction\" number without those fixes.","headline":"Useful mixed-size training idea for MLCD defect prediction, but the headline MAE and the cost claim both overstate what the body actually shows.","tokens_in":11760,"tokens_out":3052,"would_cite":true,"duration_ms":33960,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine learning on charge density from 96 small supercells predicts defect formation energies in 360-atom supercells with mean absolute error below 0.05 eV.","keywords":["defect formation energy","machine learning charge density","gallium nitride","supercell size convergence","mixed-size training","equivariant neural network","NSCF calculation","cross-size transfer"],"falsifier":"Run the same 96-structure mixed-size MLCD protocol on a GaN defect whose influence radius clearly exceeds 5.62 Å (for example a charged vacancy or an extended interstitial), with no training supercell larger than 96 atoms, and compare against 360-atom DFT formation energies: if the mean absolute error stays below 0.05 eV the cross-size claim survives, and it fails if the error rises well above that threshold.","tokens_in":10803,"feed_emoji":"⚛️","tokens_out":7315,"duration_ms":69402,"temperature":0.7,"pith_summary":"The paper claims that a machine-learning model trained on the charge density of small defect supercells can predict defect formation energies in target supercells much larger than anything in the training set, provided the training cells span a range of sizes. For four intrinsic defects in GaN, a model trained on only 96 structures containing 16–96 atoms reaches a mean absolute formation-energy error below 0.05 eV in 360-atom supercells. The authors argue this works because real-space charge density is a local, transferable quantity, whereas direct energy–force potentials trained on the same data err by more than 1 eV. If right, the result would cut the data-generation cost of first-principles defect calculations by roughly two-thirds relative to single-size training, and it suggests a general recipe for extrapolating supercell-size-dependent properties from mixed-size datasets.","feed_headline":"96 small-cell runs predict 360-atom defect energies to 0.05 eV","feed_subtitle":"Mixed-size supercell training beats energy-force potentials and needs only 96 structures in GaN.","key_machinery":"The central object is the real-space charge density, treated as a transferable local descriptor. The mechanism is a three-part pipeline: an equivariant atom-probe network (EAC-Net) is trained on perturbed small-supercell DFT charge densities; the trained model predicts the charge density of the large target supercell; a non-self-consistent DFT calculation using that density yields total energies and derived formation energies. The data design is governed by a criterion on the defect influence radius, measured by differential charge density and structural similarity, and by a three-region partition of each supercell into defect, transition, and bulk-like environments. This combination lets small cells supply configurational diversity while a few larger cells act as anchors that constrain the extrapolation along the supercell-size axis.","core_discovery":"On its own terms, the paper establishes that machine-learning charge density (MLCD) can serve as a cross-size transfer route for defect calculations. The model, trained on DFT charge densities of small supercells, predicts the charge density of a 360-atom supercell, from which total energies and formation energies are obtained by a non-self-consistent-field DFT step. A mixed-size training set containing 16-, 24-, 32-, 64-, and 96-atom supercells reaches the 0.05 eV target accuracy with 96 structures, whereas training on a single 96-atom size requires about 240 structures and training on 16–32-atom cells alone saturates around 0.1 eV even with 720 structures. The authors identify the controlling length scale: half of the training-cell lattice length must be comparable to or larger than the defect influence radius, obtained from differential charge-density analysis. The same predicted densities also reproduce band structures, defect levels, and spin polarization in the large cells.","pith_inferences":["The continuation-by-anchors view suggests a practical shortcut: estimate the defect influence radius from the size-dependent curve itself, using small- and medium-cell data, which would remove the current need for target-scale DFT charge densities to set the training-size criterion.","The strategy may generalize beyond point defects to other size-convergent properties such as charged-defect transition levels, alloy disorder, or interface segregation energies, where mixed-size anchors could play the same role.","Because the study covers four neutral intrinsic defects in GaN, the robustness of the claim for charged defects or defects with long-range strain fields is untested; those cases are natural stress tests for the 0.05 eV target.","The paper's explanation ties cross-size transfer to the spatially resolved local output of MLCD, implying that any model with a similar local-output design should show comparable robustness, not only the specific EAC-Net architecture."],"forward_implications":["With 96 mixed-size structures, MLCD reaches formation-energy errors below 0.05 eV for four intrinsic defects in 360-atom GaN supercells, cutting the dataset size by about two-thirds and the total DFT preparation cost by a factor of 6.2 relative to single-size 96-atom training.","The half-lattice-length criterion gives a quantitative rule for when a training supercell size is adequate: half of its lattice vector must be at least the defect influence radius.","MLIPs trained on the same datasets err by more than 1 eV, indicating that direct energy-force fitting lacks the local transferability that charge-density learning provides.","MLCD-predicted densities, through NSCF calculations, reproduce band structures, defect levels, and spin polarization of the large supercells, so formation energies are not the only accessible property.","Mixed-size training is framed as a continuation problem, where intermediate-size supercells anchor the size-dependent property curve and reduce extrapolation uncertainty."],"supporting_citations":[{"why":"Supplies the EAC-Net model used to predict charge densities from atomic structure.","marker":"[15]"},{"why":"Establishes that the charge density determines the total energy, the theoretical basis for the MLCD–NSCF route.","marker":"[22]"},{"why":"Provides the Kohn-Sham framework that the NSCF energy evaluation step uses.","marker":"[23]"},{"why":"Documents why large supercells are needed to suppress defect image interactions, the problem the paper addresses.","marker":"[1]"},{"why":"Gives the O(N^3) scaling that motivates cheaper cross-size defect prediction.","marker":"[7]"},{"why":"Provides a 9100-structure MLIP GaN dataset, a baseline for the data cost of defect MLIPs.","marker":"[20]"},{"why":"Shows an MLIP defect study using 1341 large supercells, another baseline for dataset cost.","marker":"[21]"},{"why":"Provides a prior method for learning charge densities, contextualizing the MLCD approach.","marker":"[14]"}],"fun_headline_variants":["Charge-density learning leapfrogs force-fitting for defect energies","96-cell MLCD beats MLIPs by 20x in defect prediction","Small-cell charge density nails 360-atom defect energies","Mixed-size training: 96 cells predict defects to 0.05 eV","ML charge density transfers across sizes for defect energies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's data efficiency rests on knowing the defect influence radius from DFT charge densities in the 360-atom target supercells; if that large-cell information is unavailable, the minimum training size cannot be set by the proposed criterion.","fun_headline_variants_meta":{"raw":{"variants":["Charge-density learning leapfrogs force-fitting for defect energies","96-cell MLCD beats MLIPs by 20x in defect prediction","Small-cell charge density nails 360-atom defect energies","Mixed-size training: 96 cells predict defects to 0.05 eV","ML charge density transfers across sizes for defect energies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000618,"raw_usage":{"total_tokens":2869,"prompt_tokens":947,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":1834}},"tokens_in":563,"tokens_out":1922,"duration_ms":15500,"temperature":1.0,"reasoning_tokens":1834,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:33:58.988623+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 96-structure mixed-size MLCD protocol on a GaN defect whose influence radius clearly exceeds 5.62 Å (for example a charged vacancy or an extended interstitial), with no training supercell larger than 96 atoms, and compare against 360-atom DFT formation energies: if the mean absolute error stays below 0.05 eV the cross-size claim survives, and it fails if the error rises well above that threshold.","supporting_citations":[{"cited_title":"& Furthm¨ uller, J","cited_arxiv_id":null,"evidence_quote":"Gives the O(N^3) scaling that motivates cheaper cross-size defect prediction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides a 9100-structure MLIP GaN dataset, a baseline for the data cost of defect MLIPs."},{"cited_title":"& Erhart, P","cited_arxiv_id":null,"evidence_quote":"Shows an MLIP defect study using 1341 large supercells, another baseline for dataset cost."}],"review_version":1}