{"id":"ac399ba0-d19e-4726-9044-5e7fba63138f","arxiv_id":"2607.29510","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Graph neural networks trained on ordered perovskites transfer well to high-entropy perovskites for formation energy but not HOMO-LUMO gap; a small amount of HEPO data restores gap accuracy.","lead":"This paper tests whether AI models trained on simple ordered perovskite crystals can predict properties of disordered high-entropy perovskite oxides, and finds it works well for formation energy but not for electronic gaps. Adding a small set of disordered samples fixes the gap predictions, giving a practical recipe for screening these materials.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unpublished, self-generated SQS-HEPO labels underpin the core transfer comparison; independent validity of these labels is unverified, and a dataset artifact could produce the observed property-dependent transfer.","rationale":"The reader's weakest assumption correctly identifies the dependency on unpublished SQS-based HEPO data. The paper's transfer conclusions are only as sound as this dataset; if the SQS models are not representative or the DFT references are inconsistent, the entire comparison is in question. The lack of public data/code and the absence of any external validation (experimental or independent computational) leaves the core result unfalsifiable from the manuscript alone. However, the paper is honest about the provenance and the 0.0-fraction caveat, so the issue is a validation gap rather than a demonstrated error. The proposed independent SQS+DFT reproduction would settle whether the dataset supports the claims. Since the reader's conditional verdict already flags this, no change in verdict is needed.","tokens_in":14520,"tokens_out":7959,"duration_ms":78389,"concrete_test":"Generate independent SQS models for one equimolar HEPO composition (e.g., Ca(Ti0.2Zr0.2Hf0.2Sn0.2Ge0.2)O3) using a different SQS implementation (e.g., ATAT/mcsqs with different seeds/order) and compute PBE formation energy and HOMO-LUMO gap with a different DFT code (e.g., VASP, same pseudopotential class). Compare against the authors' reported values for the same composition. If the mean absolute difference exceeds ~5 meV/atom for Ef or ~0.1 eV for Eg, or if the ranking of formation vs gap transfer changes, the central conclusion is an artifact of the unpublished SQS construction. Also compute a per-A-site mean baseline on the HEPO test set; if its MAE is comparable to or lower than the reported 1.98 meV/atom, the 'effective transfer' claim is vacuous.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims — that formation-energy transfer works and HOMO-LUMO gap transfer fails — are quantified entirely on a 1810-structure SQS-based HEPO set whose DFT labels (Ef and Eg) come from the authors' own unpublished prior works (refs 33–34). No external dataset, no experimental validation, and no release of SQS structures or computed properties is provided. If the SQS models do not faithfully represent the configurational disorder of real HEPOs, or if the DFT reference phases/settings differ between the ordered and disordered domains, the comparison collapses. The paper also does not report any baseline (e.g., predicting per-family means) to contextualize the 1.98 meV/atom HEPO MAE; given the tiny formation-energy variance for Sr- and Ba-based families (Table 1: σ = 4.10 and 2.44 meV/atom), a non-trivial portion of the apparent transfer success may reduce to predicting the A-site average rather than a learned structure–property relationship.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper benchmarks four graph neural networks (CGCNN, GATGNN, ALIGNN, M3GNet) for predicting formation energy (Ef) and HOMO–LUMO gap (Eg) on a curated DFT dataset of ordered ABO3/A2BB'O6 perovskites (10,898 structures) and SQS-modeled high-entropy perovskite oxides (1,810 structures). The authors evaluate each model on held-out ordered perovskites, then blind-transfer the ordered-trained models to HEPO test structures, and finally fine-tune with fractions of HEPO training data. The central claims are: (i) ALIGNN is the most accurate model on the ordered domain; (ii) formation-energy prediction transfers effectively to HEPOs (blind MAE 1.98 meV/atom, close to the ordered test MAE of 2.47 meV/atom); (iii) HOMO–LUMO-gap prediction transfers poorly (blind HEPO MAE 0.28 eV vs 0.03 eV on ordered); and (iv) adding a small HEPO training fraction substantially improves gap prediction. UMAP embedding analysis is used to argue that angle-aware representations, especially ALIGNN, better organize chemistry and octahedral-tilting information.","tokens_in":14701,"tokens_out":4766,"duration_ms":50849,"significance":"If the empirical findings are robust, this is a useful practical contribution to ML-based screening of high-entropy perovskite oxides. The controlled design — fixed A/B elemental pools between ordered and disordered domains, composition-stratified splits, and a clearly described blind-transfer protocol — is a strength. The data-efficiency result, namely that a small HEPO-specific calibration set can close most of the gap-transfer error, is actionable. However, the central comparison currently depends on an unpublished SQS dataset with no released structures or external validation, and the paper lacks uncertainty quantification and trivial baselines. These gaps need to be addressed before the practical conclusions can be considered quantitatively established.","major_comments":[{"comment":"No uncertainty quantification is provided. Claims such as 'ALIGNN gives the best overall performance' rest on MAE differences as small as 0.01 eV for Eg (0.03 vs 0.04 eV) and, in the fine-tuning curve of Fig. 9(c), differences of 0.01 eV between fractions (e.g., 0.06 vs 0.05 eV). With a single training run per model and no seed variation, these differences are within typical run-to-run noise. Please report mean±std over multiple random seeds and, where relevant, paired statistical tests.","section":"Training models on ordered domain / Fig. 3"},{"comment":"Trivial baselines are missing. The HEPO formation-energy test distributions are very narrow for Sr and Ba (Table 1: σ = 4.10 and 2.44 meV/atom). For a normal distribution, a family-mean predictor gives MAE ≈ 0.798σ, i.e., ≈3.3 and ≈1.9 meV/atom for Sr and Ba, respectively. The reported Ba HEPO MAE of 2.24 meV/atom (Fig. 5c) is actually worse than this baseline, and the overall blind MAE of 1.98 meV/atom is only modestly better than a composition-resolved mean baseline. Without such baselines on the HEPO test set, 'formation-energy transfer works' is not quantitatively established. Please add trivial baselines (e.g., composition family mean, composition plus A-site mean) and discuss the results relative to them.","section":"Transferability from ordered perovskites to HEPOs / Table 1, Fig. 5, Fig. 9"},{"comment":"The entire transfer benchmark rests on the 1810 SQS-generated HEPO structures and their DFT labels, which are taken from the authors' own 'to be submitted' works (refs 33–34). No SQS structures, computed properties, or pseudopotential/reference-phase consistency details are released, and no independent external validation is provided. If the SQS models do not faithfully represent real HEPO disorder, or if the DFT reference phases/settings differ between the ordered and disordered domains, the central comparison could be an artifact. Please release the dataset or provide an independent reproduction, and explicitly document the consistency of DFT settings and reference phases across the ordered and HEPO domains.","section":"Dataset construction / Methods / refs [33,34]"},{"comment":"The 0.0 HEPO-fraction point is not a blind-transfer point: although no HEPO structures are in the training set, the HEPO validation set is used for model selection. The paper clearly acknowledges this, but the figure and surrounding text should make the distinction visually and verbally explicit. As presented, a reader may misinterpret the 0.14 eV value at fraction 0.0 as a blind-transfer error, when the true blind error is 0.28 eV. This does not invalidate the fine-tuning trend, but the labeling is important for the paper's central message.","section":"Fig. 9(c), 'Transferability from ordered perovskites to HEPOs'"}],"minor_comments":[{"comment":"Typo: 'using a a 80:10:10 split' should read 'using an 80:10:10 split'.","section":"Dataset construction"},{"comment":"Inconsistent capitalization: 'M3GNET' appears where 'M3GNet' is used elsewhere.","section":"Error analysis across chemical species"},{"comment":"The M3GNet formation-energy training loss (2.12×10^-2) exceeds its validation loss (7.08×10^-3), which is unusual and may indicate a normalization or logging issue. Please clarify or verify.","section":"Training behavior / Fig. 2"},{"comment":"The caption of Fig. 9 should explicitly state that the '0.0' fraction uses HEPO validation for model selection, while the blind-transfer parity plots do not. This distinction appears only in the main text.","section":"Transferability / Fig. 9 caption"},{"comment":"Several key references (refs 31–34) are 'to be submitted' preprints; if the companion papers become available, the authors should update these references and ideally cite the actual datasets.","section":"Introduction / References"}],"recommendation":"major_revision","confidential_remarks":"The paper's main claims are plausible and the controlled experimental setup is attractive, but the empirical foundation is currently fragile: no multi-seed uncertainty, no trivial baselines, and a central dataset that is unpublished and unreleased. I would advise the editor that major revision is needed; if the authors release the SQS/DFT dataset and add baselines and error bars, the paper could become a solid contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things about arXiv:2607.29510. First, the core result is credible and useful: with a fixed A/B element pool, formation energy transfers from ordered perovskites to SQS-modeled HEPOs about as well as it interpolates within the ordered set, whereas the HOMO-LUMO gap does not; a small HEPO-specific training fraction (20% of the HEPO training split) restores gap accuracy to 0.07 eV. That is a clean, property-dependent transfer benchmark that I don't think is already in the literature. Second, the benchmark stands almost entirely on the authors' own unpublished DFT data for the SQS HEPOs, and no code, data, or checkpoints are released. That is the main thing to fix before trusting the numbers.\n\nWhat the paper does well: the element pools are held fixed between ordered and disordered domains, so the comparison isolates disorder rather than new chemistry. The blind-transfer protocol is honestly described, including the admission that the 0.0-fraction fine-tuning point uses HEPO validation for model selection. The UMAP analysis is suggestive rather than conclusive, but it supports the ALIGNN story without being pushed too hard. The authors also do not oversell; their practical recommendation—use ordered-trained angle-aware models for stability screening, add a small HEPO calibration set for electronic properties—is measured and actionable.\n\nThe soft spots are real but mostly addressable. There are no error bars or multi-seed runs, so claims like 0.04 vs 0.05 eV between models are not statistically grounded. The angle-awareness explanation for ALIGNN is confounded by architecture and training choices; it is plausible, given the known role of octahedral tilting, but not proven by this comparison. The stress-test concern about dataset artifacts is the right worry, even if it is not established: refs 33–34 are to-be-submitted, no external data or independent reproduction is provided, and if the SQS models or the formation-energy reference phases are inconsistent with the ordered set, the property-dependent story could partly evaporate. Also, the Sr and Ba formation-energy labels have tiny variance, so the low MAEs for those families are less impressive than they look; the model still beats a mean predictor (e.g., Sr-family MAE 1.33 meV/atom vs test standard deviation 4.10 meV/atom), so it is not pure mean prediction, but the low variance makes those numbers easier to achieve.\n\nThis deserves a serious referee. I would send it to peer review and ask for data/code release, multi-seed statistics, and a mean-prediction baseline for the low-variance families. The central finding is likely to survive, but the paper should not be accepted in its current form.","headline":"Useful property-dependent transfer benchmark for perovskite GNNs, but the core numbers rest on unpublished SQS/DFT data and there are no error bars; deserves a serious referee with data release and multi-seed runs.","tokens_in":15264,"tokens_out":3116,"would_cite":true,"duration_ms":33783,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that formation-energy prediction learned from ordered perovskites transfers directly to high-entropy perovskite oxides, while HOMO-LUMO gap prediction does not and requires a small HEPO-specific training set.","keywords":["high-entropy perovskite oxides","transfer learning","graph neural networks","formation energy","HOMO-LUMO gap","special quasirandom structures","ALIGNN","chemical disorder"],"falsifier":"Regenerate HEPO structures with an independent disorder sampling method (e.g., a different SQS implementation, random supercells, or experimentally determined cation arrangements), recompute the same DFT targets, and re-run the blind transfer test; if formation-energy MAE rises well above ~2 meV/atom or HOMO-LUMO gap MAE falls below ~0.1 eV, the property-dependent transfer conclusion would not hold.","tokens_in":14368,"feed_emoji":"⚛️","tokens_out":7096,"duration_ms":61397,"temperature":0.7,"pith_summary":"This paper asks whether machine-learning models trained on chemically ordered perovskite oxides can be reused for high-entropy perovskite oxides (HEPOs), where cation disorder makes density-functional-theory calculations expensive. Using a curated dataset of over 12,000 relaxed structures spanning ABO3, A2BB'O6, and SQS-generated HEPO perovskites, the authors show that formation-energy prediction transfers almost unchanged from the ordered to the disordered domain (blind HEPO error 1.98 meV/atom versus 2.47 meV/atom on ordered test structures), while HOMO-LUMO gap prediction degrades roughly tenfold (0.28 eV versus 0.03 eV) and recovers only when a small HEPO-specific training fraction is added. The authors attribute this asymmetry to the gap's sensitivity to local chemical environments and disorder, versus stability's dependence on more transferable bond-length and octahedral-tilting patterns. The practical conclusion is a screening recipe: angle-aware graph neural networks trained on ordered perovskites are ready for HEPO stability screening, but electronic-property screening needs a small HEPO calibration set.","feed_headline":"Formation energy transfers to high-entropy oxides; band gaps won't","feed_subtitle":"A small HEPO-specific dataset fixes band-gap prediction, so costly DFT screening can be targeted.","key_machinery":"The central mechanism is the graph neural network representation of crystal structure, compared across four architectures: CGCNN (pairwise atom-bonds), GATGNN (attention-weighted graphs), ALIGNN (line-graph encoding of bond angles), and M3GNet (three-body geometric features). The load-bearing object is ALIGNN's explicit angular message passing, which allows the model to encode B-O-B angles and octahedral tilting—structural motifs previously shown to control perovskite stability and electronic structure. The transfer protocol itself is the other key piece: training only on chemically ordered ABO3/A2BB'O6 perovskites and evaluating on SQS-generated HEPO structures shares the same elemental poo","core_discovery":"The central discovery is that ordered-to-disordered transfer learning in perovskite oxides is property-dependent: formation energies learned from ordered single and double perovskites transfer directly to disordered high-entropy perovskites (ALIGNN blind HEPO MAE 1.98 meV/atom, similar to its 2.47 meV/atom ordered test error), whereas HOMO-LUMO gap prediction shows a roughly tenfold error increase (0.28 eV versus 0.03 eV) with systematic underestimation. The paper further finds that this transfer deficit is largely correctable: adding only 20% of the HEPO training subset (16% of all HEPO structures) reduces the gap MAE to 0.07 eV, and using the full HEPO training set reaches 0.04 eV. Among f","pith_inferences":["The property-dependent transfer asymmetry likely extends to other local-environment-sensitive electronic properties (band edges, effective masses, optical spectra), so the same ordered-to-disordered protocol could be used to decide where calibration data are needed.","The transfer success for formation energy suggests that bond-length and tilting descriptors are largely disorder-invariant for these chemistries; if true, a descriptor-based model (not necessarily a GNN) might achieve the same transfer with far less training data.","The SQS representation, by construction, captures only a finite set of configurations; the transfer conclusions should be tested against larger supercells or experimentally resolved structures before being used for quantitative HEPO discovery.","The observed non-monotonic formation-energy error with HEPO training fraction (minimum at 0.8) hints that adding disordered data can slightly disturb ordered-domain knowledge, suggesting a possible role for regularization or multi-task learning in future pipelines."],"forward_implications":["Formation-energy models trained on ordered perovskites can be used directly to screen HEPO thermodynamic stability without additional HEPO DFT data.","For HEPO electronic-structure screening, a small HEPO-specific calibration set (20% of a training split) is enough to cut the HOMO-LUMO gap error from 0.28 eV to 0.07 eV.","Angle-aware graph neural networks (ALIGNN) should be preferred over pairwise-only GNNs for perovskite property prediction, because angular encoding improves both accuracy and latent-space organization.","Part of the apparent transfer failure for HOMO-LUMO gaps is a validation-domain mismatch: using HEPO validation data for model selection halves the blind error from 0.28 eV to 0.14 eV even without HEPO training structures."],"fun_headline_variants":["Band gaps fail transfer in high-entropy oxides; small data fixes it","Perovskite transfer learning: formation energy works, band gap doesn't","High-entropy oxides: GNN transfer good for energy, poor for gaps"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The benchmark rests on the authors' own SQS-generated HEPO dataset (1810 structures) and DFT labels from their previous works, several of which are still 'to be submitted'; if those SQS models do not faithfully represent real HEPO disorder, or if the reference phases are inconsistent between the ordered and disordered domains, the transfer comparison collapses.","fun_headline_variants_meta":{"raw":{"variants":["Band gaps fail transfer in high-entropy oxides; small data fixes it","Perovskite transfer learning: formation energy works, band gap doesn't","High-entropy oxides: GNN transfer good for energy, poor for gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000212,"raw_usage":{"total_tokens":1274,"prompt_tokens":786,"completion_tokens":488,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":425}},"tokens_in":530,"tokens_out":488,"duration_ms":5535,"temperature":1.0,"reasoning_tokens":425,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T05:31:35.528943+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate HEPO structures with an independent disorder sampling method (e.g., a different SQS implementation, random supercells, or experimentally determined cation arrangements), recompute the same DFT targets, and re-run the blind transfer test; if formation-energy MAE rises well above ~2 meV/atom or HOMO-LUMO gap MAE falls below ~0.1 eV, the property-dependent transfer conclusion would not hold.","supporting_citations":[],"review_version":1}