{"id":"3120b75b-ff12-403e-944b-13d168bc7dce","arxiv_id":"2412.19330","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A tiered geometry, electrostatics, and machine-learning screen predicts thousands of new split-vacancy defects in the Materials Project.","lead":"This computational study searches for split vacancies, a rearranged defect shape where one missing atom turns into a two-vacancy plus interstitial complex. It predicts these defects are far more common than the handful of known cases, appearing in roughly 10 percent of cation vacancies across the Materials Project database.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline 10% prevalence is an ML-only census validated on oxides and nitrides with ~40% precision; without a stratified DFT check across the full Materials Project chemistry space, the central discovery claim is not yet established.","rationale":"I read the paper as a methods-plus-discovery claim. The tiered screening workflow is plausible, and the DFT validation on known split vacancies and the 444-oxide set is meaningful independent support. However, the most load-bearing unverified step is the transfer of MACE-mp from the oxide/nitride validation sets to the full MP database. The numerical headline — 10% of cation vacancies — is not directly DFT-confirmed; the authors' own 40% precision lower bound would reduce the split-vacancy count from 29,000 to roughly 12,000 (~4% prevalence), yet the Discussion quotes 10% without correction. Because the paper's new scientific conclusion is prevalence at scale, this needs a stratified DFT confirmation before the claim is accepted as stated. The reader's conditions already include a random DFT sample for MP predictions, so I agree with the conditional verdict, but I would elevate the ML-transfer and false-positive issue over the electrostatic prescreen as the primary risk: the electrostatic stage determines recall, while the ML stage determines the reported counts and prevalence. No ad hominem is intended; the concern is about extrapolation evidence, not author conduct.","tokens_in":39600,"tokens_out":7133,"duration_ms":68769,"concrete_test":"DFT-relax a stratified random sample of ~300 ML-predicted split vacancies from the MP screen, oversampling the high-prevalence chemistries in Fig. 7c,d (halides, Hg, C, chalcogenides, low-oxidation-state compounds) using the same PBE/MPRelaxSet supercell settings as the nitride validation, and compare each relaxed geometry and energy against the corresponding simple-vacancy relaxation. Compute per-stratum precision; if the overall precision is statistically consistent with 39-44%, the 10% claim gains support, while a drop below ~25% requires revising the Discussion prevalence and headline counts to lower-bound-corrected values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — ~29,000 (10%) of cation vacancies in the Materials Project are split vacancies, and that the workflow 'identifies' them — depends entirely on MACE-mp relaxations for the full MP census (Results, Fig. 7). The paper's own validation is limited to two non-representative subsets: 444 oxides (44% precision for lower-energy split vacancies) and stable nitrides (39% precision for split-classified relaxations). MACE-mp is trained mostly on oxide PBE data, yet Fig. 7c,d highlight the highest predicted prevalence in halides, Hg, C, chalcogenides and coinage metals — exactly the chemistries where formal-charge electrostatics and a PBE-trained potential are least tested. The paper uses the ~40% precision only to convert the 55,000 ML-predicted lower-energy structures into a lower bound (~22,000 true), which would imply ~12,000 true split vacancies (~4% prevalence), but the Discussion states 'around 10% of cation vacancies in all inorganic solids' without applying this correction. If precision in untested chemistries is below 40%, the headline counts and prevalence are inflated. The electrostatic prescreen (Alg. 1) is a false-negative risk, but the ML false-positive/transfer risk is the more direct threat to the quantitative headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a multi-stage screening workflow for identifying split-vacancy defects: enumeration of candidate V_X–X_i–V_X complexes, an Ewald-based formal-charge electrostatic prescreen with a 110% energy cutoff, relaxation of survivors with the MACE-mp foundation model, and final DFT confirmation of selected candidates. The workflow is first validated on known split-vacancy cases and on a 444-oxide DFT screen, which finds 93 lower-energy cation vacancies. The method is then applied to the full Materials Project database, yielding ML predictions of roughly 55,000 lower-energy cation vacancies and 29,000 ML-classified split vacancies. The paper argues that split vacancies are common (around 10% of cation vacancies) and that foundation ML potentials can accelerate such defect searches, with explicit caveats about their limited domain of validity.","tokens_in":39895,"tokens_out":6927,"duration_ms":64365,"significance":"If the central claims hold, the paper is significant: it offers a practical computational pipeline for finding a class of defect reconstructions that local structure-searching methods routinely miss, provides a concrete DFT-validated set of 93 lower-energy cation vacancies in oxides, and demonstrates that a foundation ML potential, combined with electrostatic screening, can search the full Materials Project space. The known-case benchmark (PBEsol vs PBE0 relative energies with R2=0.998) is a strong piece of evidence that semi-local DFT is adequate for these fully-ionized defects, and the open code and database (modulo the placeholder DOI) are valuable community resources. The main weakness is that the quantitative census results are ML predictions with only partial validation, and the manuscript's own precision-correction is not carried through to its headline prevalence statement.","major_comments":[{"comment":"The headline prevalence statement is inconsistent with the paper's own precision correction. The Results report that the ML model predicts 29,000 (10%) of cation vacancies in the Materials Project to be split vacancies, and the Discussion repeats 'around 10% of cation vacancies in all inorganic solids'. However, the preceding text states that applying the ~40% precision from the oxide and nitride validations yields ~12,000 true split vacancies, which is approximately 4% of cation vacancies, not 10%. Please revise the Discussion to quote the precision-corrected figure or to label the 10% explicitly as an uncorrected ML prediction (an upper bound). As written, the conclusion overstates the established prevalence.","section":"Discussion & Conclusions"},{"comment":"The census numbers (55,000 and 29,000) are based entirely on MACE-mp relaxations, with DFT confirmation only for an oxide subset (44% precision for lower-energy split vacancies) and a nitride subset (39% precision for split-classified relaxations). Both validation sets are compositionally limited, and Fig. 7c,d show the highest predicted prevalences in halides, carbon, chalcogenides, mercury, and coinage-metal compounds - precisely the chemistries where formal-charge electrostatics and the PBE-trained MACE-mp model are least tested. The ~40% precision is applied as a uniform correction factor, but no evidence is given that it transfers to these untested chemistry classes. Please provide a stratified DFT check across the chemistries that dominate the predicted census, or explicitly state that the 29,000 and 10% figures remain unconfirmed predictions.","section":"Machine Learning Acceleration (Fig. 7)"},{"comment":"The recall of the electrostatic prescreen is not quantified. The 110% cutoff is an ad hoc tuning parameter, and the paper shows only that the known split-vacancy cases fall below it; no DFT calculations are performed on candidates that fail the cutoff. Since the abstract claims the approach 'allows the screening of all solid-state compounds', the possibility of false negatives - especially in chemistries where strain, pair repulsion, or covalent effects dominate - should be addressed. I recommend relaxing a random sample of excluded candidates for a few diverse host compounds to estimate the prescreen's sensitivity, or clearly stating that the true recall is unknown and that the pipeline is a heuristic search rather than a complete enumeration.","section":"Algorithm 1 and 'Screening Split Cation Vacancies in Oxides'"}],"minor_comments":[{"comment":"The database and code DOI is given as a placeholder ('https://doi.org/10.5281/zenodo.XXXX'); please provide the actual DOI before publication.","section":"Data availability"},{"comment":"The caption refers to 'the full DFT calculated dataset (~1000 compounds)', while the oxide screen described in the text covers the first 444 compounds; please clarify what structures are included in this figure.","section":"Fig. 4b caption"},{"comment":"The metric definitions are crowded into the caption; consider moving the definitions of TPR/FPR/TNR/FNR and the prevalence note into the main text or a separate methods paragraph for readability.","section":"Table 2"},{"comment":"The statement that ShakeNBreak fails to identify the split-vacancy ground state for V_Ga in Ga2O3 is attributed to unpublished work; please add a citation or a public preprint if available.","section":"Introduction"},{"comment":"The '+0.35 eV' energy window in the 'exhaustive' ML criterion and the 10% electrostatic cutoff are both presented without sensitivity analysis; a brief justification of these numerical choices would help readers gauge how robust the screening is to their variation.","section":"Algorithm 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically solid in its core demonstration, but the census-level claims need to be reconciled with the reported validation precision. The author is also the main developer of the doped package that is central to the workflow, so there is a natural alignment of interests, though the manuscript is transparent about the tools used. I would suggest the editor ask for the Discussion to be rewritten around the precision-corrected prevalence and a clearer statement of which numbers are confirmed vs predicted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one. The core result is that split cation vacancies are not a freak occurrence: in 444 oxides, full DFT relaxations starting from electrostatically-filtered candidate geometries find 93 cation vacancies (about 10% of the possible ones) that sink 0.05–3 eV below the standard relaxed vacancy, mean 0.81 eV. Those energy differences change defect concentrations by orders of magnitude. That's real, reproducible evidence, and it's the strongest part of the paper.\n\nThe new thing is the tiered screening itself: generate V_X–X_i–V_X triples geometrically, prune with formal-charge Ewald electrostatics, relax the small remaining set with MACE-mp, then confirm the survivors with DFT. The known-case benchmark is excellent (PBEsol vs PBE0 relative energies, R^2=0.998) and the oxide DFT screen is a legitimate dataset. The pipeline is clearly described and the author is honest about the failure of local search methods like ShakeNBreak here.\n\nThe soft spots are all in the extrapolation from the DFT-tested sets to the full Materials Project census. The ~55,000 / ~29,000 / 10% numbers in Results are MACE-mp predictions, not DFT-confirmed. The paper's own validation on oxides and nitrides gives only ~39–44% precision for split-classified predictions, and those are chemistries MACE-mp knows reasonably well. The predicted hotspots—halides, Hg, chalcogenides, coinage metals—are exactly where neither formal-charge electrostatics nor a PBE-trained potential has been tested. The paper does compute a ~4% lower bound by applying the 40% precision, but then the Discussion states \"around 10%\" without the caveat, and the abstract says \"identifying thousands\" when most are ML predictions. That needs fixing. Also, the code and data repository is promised but not actually linked (placeholder Zenodo DOI), which blocks anyone from checking the database or reproducing the numbers.\n\nThe electrostatic prescreen could in principle miss covalent-stabilized split vacancies, but the paper acknowledges that limitation and it's not a load-bearing flaw. The lack of a stratified DFT sample across the untested chemistries is the real gap.\n\nVerdict: send to peer review. The methodology and the oxide/nitride test results deserve referee time, and the author should be pushed to release the repository and run a stratified DFT check before the headline numbers are quoted. This paper belongs in the defect-chemistry conversation.","headline":"A smart screening pipeline that makes a strong case that split vacancies are common, but the headline prevalence number rides on ML extrapolation that the paper hasn't yet validated in the chemistries where it matters most.","tokens_in":40427,"tokens_out":2702,"would_cite":true,"duration_ms":25563,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["61.72.jj","71.55.-i"],"model":"deepseek-v4-flash","headline":"A tiered screen combining geometric enumeration, formal-charge electrostatics, and a pretrained machine-learned potential can identify split-vacancy defects across essentially all known inorganic solids, and such defects are common…","keywords":["split vacancies","point defects","defect metastability","machine-learned interatomic potentials","electrostatic screening","high-throughput screening","cation vacancies","foundation models"],"falsifier":"Take a set of compounds with strongly covalent or low-symmetry bonding (e.g., small-gap semiconductors, transition-metal oxides with partially filled d shells, or layered van der Waals solids), enumerate all candidate split vacancies for a sample of cation vacancies, relax every candidate with DFT (no electrostatic pre-screen), and count how many low-energy split vacancies (ΔE < −0.025 eV) have initial electrostatic energies above the 110% cut-off; if a material-dependent fraction falls above the cut-off, the pre-screen's recall is not universal.","tokens_in":39337,"feed_emoji":"🔍","tokens_out":8476,"duration_ms":70363,"temperature":0.7,"pith_summary":"Split vacancies are defects in which removing one atom triggers a nearby atom to move into an interstitial position, leaving two vacancies flanking one interstitial. This paper argues that these reconstructions are common rather than exotic: in a large high-throughput test, roughly 10% of cation vacancies relax to a split vacancy that is lower in energy than the simple vacancy, with energy lowerings averaging about 0.8 eV and exceeding 2 eV in some cases. Because standard defect simulations usually relax from the unperturbed vacancy geometry, they systematically miss these lower-energy states. The paper shows that a cheap pre-screen based on formal-charge electrostatics, followed by relaxation with a general machine-learned interatomic potential, can find split vacancies across essentially all known inorganic solids, yielding thousands of predicted low-energy split-vacancy configurations.","feed_headline":"1 in 10 cation vacancies hides a split vacancy","feed_subtitle":"A tiered screen using electrostatics and machine learning finds thousands of missed low-energy defect states.","key_machinery":"The load-bearing object is the split vacancy itself, defined as the stoichiometry-conserving complex $[V_X + X_i + V_X]$ in which one host atom leaves its lattice site to sit between two empty sites. The workflow that carries the argument is a tiered screening algorithm: geometric enumeration of candidate complexes (with vacancy-interstitial distances under 5 Å), an electrostatic pre-screen using Ewald sums with formal ionic charges and a 110% energy cutoff relative to the simple vacancy, a relaxation pass with a foundation machine-learned interatomic potential (retaining geometries that stay split and lie within 0.35 eV of the simple vacancy), and final density functional theory evaluation. The paper's key empirical claim is that the formal-charge electrostatic energy, despite ignoring screening, strain, and covalency, ranks the candidate geometries well enough that the true low-energy split vacancies sit in the low-energy tail of the electrostatic distribution.","core_discovery":"The central claim is that low-energy split vacancies—stoichiometry-conserving complexes of the form $[V_X + X_i + V_X]$—are a widespread feature of cation vacancies in inorganic compounds, and that they can be identified systematically by a tiered workflow: enumerate all symmetry-inequivalent vacancy-interstitial-vacancy combinations within 5 Å, compute their formal-charge Ewald energies, keep only those within 110% of the simple vacancy's electrostatic energy, relax the survivors with a pretrained machine-learned interatomic potential, and finally confirm with density functional theory. Applied to a set of stable insulating metal oxides, the workflow finds 93 cation vacancies whose split (or split-like) geometry lies 0.05–3 eV below the best simple vacancy; extrapolated to a database of roughly 150,000 known and predicted compounds, the ML stage predicts about 29,000 cation vacancies classified as split vacancies, with density functional theory spot checks on oxides and nitrides confirming 40–60% of the predictions. If these numbers hold, the standard practice of relaxing defects from the ideal vacancy geometry is missing the true ground state for a substantial fraction of all cation vacancies.","pith_inferences":["If the 10% prevalence holds, experimental probes that assume simple monovacancy pictures—positron annihilation, EPR, deep-level spectroscopy—may need re-analysis across broader materials classes.","The same geometry+electrostatics+ML pipeline could be applied to other stoichiometry-conserving defect complexes (divacancies, antisite pairs, DX-like off-centre substitutions) and to colour-centre discovery for quantum technologies.","Given the reported 40–60% DFT validation accuracy, the ML screen over-predicts; the true number of split vacancies in the full database is likely in the thousands rather than 29,000, though still a large fraction of the initial estimate.","A testable extension is to benchmark the 110% electrostatic cut-off against a chemically diverse set of compounds to measure its recall, and to use a tighter cut-off where covalent bonding dominates."],"forward_implications":["Standard defect relaxations that start from the ideal vacancy and relax locally will miss the true ground state for roughly 10% of cation vacancies, so defect concentrations and properties computed from simple vacancies are systematically wrong for those materials.","The mean energy lowering of about 0.8 eV changes equilibrium defect populations by orders of magnitude: roughly a factor of 10^3 at 1000 K and 10^10 at 300 K for a single defect.","The ML-accelerated screen achieves a discovery acceleration factor of about 120 relative to random candidate selection, making whole-database defect structure searches feasible in about a GPU-day.","The predicted split-vacancy database is integrated into the defect-generation toolkit, so researchers are automatically alerted when their host compound has a likely split vacancy and at what confidence.","The method's success is specific to fully ionized (formal-charge) defects; it does not address metastabilities driven by charge localization, which require different handling."],"supporting_citations":[{"why":"supplies the prototypical split-vacancy case in β-Ga2O3 that motivates the search","marker":"[25]"},{"why":"catalogues known split vacancies in semiconducting oxides and frames the known-case set","marker":"[26]"},{"why":"provides the oxide test set and the vacancy classification algorithm used for validation","marker":"[43]"},{"why":"supplies the large database of computed materials that is screened at scale","marker":"[55]"},{"why":"supplies the pretrained machine-learned interatomic potential used for ML relaxation","marker":"[58]"},{"why":"is the local structure-searching baseline that fails to find split vacancies","marker":"[39]"},{"why":"implements the geometric enumeration, symmetry analysis, and classification routines","marker":"[45]"},{"why":"documents split vacancies in sapphire and validates semi-local DFT for fully ionized cases","marker":"[10]"}],"fun_headline_variants":["Thousands of split vacancies evade standard defect screens","AI and electrostatics reveal hidden split vacancy states","1 in 10 cation vacancies may be split vacancies","Split vacancy defects are far more common than thought","New screen finds thousands of low-energy split vacancies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The screen assumes that a formal-charge Ewald energy within 110% of the simple vacancy's electrostatic energy is a necessary condition for a split vacancy to be low in energy; if a material's bonding is dominated by covalency, strain, or charge localization, low-energy split vacancies may sit above that cut-off and never reach the machine-learning or DFT stages.","fun_headline_variants_meta":{"raw":{"variants":["Thousands of split vacancies evade standard defect screens","AI and electrostatics reveal hidden split vacancy states","1 in 10 cation vacancies may be split vacancies","Split vacancy defects are far more common than thought","New screen finds thousands of low-energy split vacancies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000969,"raw_usage":{"total_tokens":4151,"prompt_tokens":1003,"completion_tokens":3148,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":3076}},"tokens_in":619,"tokens_out":3148,"duration_ms":19116,"temperature":1.0,"reasoning_tokens":3076,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:41:45.307776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of compounds with strongly covalent or low-symmetry bonding (e.g., small-gap semiconductors, transition-metal oxides with partially filled d shells, or layered van der Waals solids), enumerate all candidate split vacancies for a sample of cation vacancies, relax every candidate with DFT (no electrostatic pre-screen), and count how many low-energy split vacancies (ΔE < −0.025 eV) have initial electrostatic energies above the 110% cut-off; if a material-dependent fraction falls above the cut-off, the pre-screen's recall is not universal.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"catalogues known split vacancies in semiconducting oxides and frames the known-case set"},{"cited_title":"Kumagai, N","cited_arxiv_id":null,"evidence_quote":"provides the oxide test set and the vacancy classification algorithm used for validation"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the large database of computed materials that is screened at scale"},{"cited_title":"Mosquera-Lois, S","cited_arxiv_id":null,"evidence_quote":"is the local structure-searching baseline that fails to find split vacancies"},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"implements the geometric enumeration, symmetry analysis, and classification routines"}],"review_version":1}