{"id":"5fc0480d-4213-4092-b179-cba05dd53c97","arxiv_id":"2607.09910","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"An explainability-plus-DFT audit finds Pt–p-valence entanglement in a composition-only SHC model and a ~30× error in the HfC training label, both confirmed by independent DFT.","lead":"The paper introduces an audit that uses SHAP, counterfactual tests, and cross-model checks, then settles each finding with a few DFT runs, applied to spin Hall conductivity. It shows a composition-only model can match structure-aware networks while exposing a Pt shortcut and a thirtyfold training-label error.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Data-audit pillar treats authors’ PBE+SOC+WannierBerri SHC as ground truth for Zhao labels without a matched recompute of the original HfC structure/settings.","rationale":"The reader correctly isolates the softest load-bearing assumption: treating one DFT pipeline as ground truth for another’s published labels when the paper itself shows large inter-study scatter (W3Ta) and when absolute SHC is method-sensitive. The model-audit pillar (Pt–p_frac entanglement, HgOsPb2 under-prediction 717→2703) is more robust because it is an internal representation diagnosis confirmed by a large directional discrepancy on a Pt-free chemistry the model was expected to under-flag; that does not require absolute label fidelity. The data-audit claim does. A single matched recompute of Zhao’s HfC entry would settle whether the ~30× figure is a true label error or a cross-pipeline effect. Until then the protocol remains valuable as a discrepancy-flagging loop, but the stronger wording (“thirtyfold error… inherited undetectably by every black-box model”) should stay conditional on that control. No stronger internal inconsistency appears; the concern is evidentiary, not conceptual. Verdict therefore stays CONDITIONAL, aligned with the reader.","tokens_in":19128,"tokens_out":673,"duration_ms":5832,"concrete_test":"Take the exact HfC structure (and any documented computational settings) used for the Zhao et al. label; recompute max |σz_xy| with the authors’ QE+PBE+SOC+WannierBerri pipeline on that same structure (and, if possible, with Zhao’s reported k/Wannier settings). If the recomputed value remains ~O(100) rather than ~3.6, the label-error claim stands; if it collapses toward the Zhao number or shows O(10×) pipeline scatter, the data-audit pillar must be restated as “pipeline discrepancy / possible label issue” rather than a confirmed thirtyfold database error.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim’s second pillar (data audit) asserts a ~30× training-label error for HfC (Zhao 3.62 vs authors’ DFT 120). That conclusion requires the discrepancy to be primarily a database/label error rather than a methodological or convention difference. The paper itself documents large literature scatter for the same compound (W3Ta: Zhao 1011 vs dedicated A15 work 2250), and absolute intrinsic SHC is known to be sensitive to structure, k-mesh density, Wannier projection, energy window, and SOC pseudopotentials. The verification panel reports authors’ independent full-window peaks under their own pipeline (QE/PBE/fully-relativistic/WannierBerri 200³) but does not recompute Zhao’s original HfC entry under matched structure and settings, nor does it quantify pipeline-to-pipeline variance on a shared control set. Without that control, the data-audit finding remains under-determined: the model–label disagreement is real and useful as a flag, but the claim that every black-box model “inherits a thirtyfold error” overstates what the present DFT panel can adjudicate.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript proposes a model-agnostic audit protocol for materials ML that combines global SHAP attribution, counterfactual partial-dependence analysis, and Rashomon-style cross-model checks, with each finding adjudicated by targeted DFT. Applied to intrinsic spin Hall conductivity, a composition-only Random Forest (211-D descriptor; no relaxed structure) reaches a test MAE of 114.5 (ℏ/e)(S/cm) on a polymorph-reduced Zhao et al. set, competitive with reported structure-aware CGCNN/Res-CGCNN numbers. The model audit diagnoses statistical entanglement of the average p-valence descriptor with Pt content, consistent across RF and GPR; independent DFT on Pt-free HgOsPb2 yields a peak SHC of 2703 versus an RF prediction of 717. The data audit, triggered by a large model–label disagreement on HfC, reports an independent DFT peak of 120 versus the Zhao label 3.62 (~30×), argued to be inherited silently by black-box models trained on the same labels. Agreement controls (VPt8, BiPt, K2Pb2O3) and under-prediction cases (LiIr, W3Ta) are used to bound reliability.","tokens_in":19481,"tokens_out":1739,"duration_ms":21817,"significance":"If the dual-pillar protocol holds, it addresses a genuine gap: standard held-out metrics do not test whether features track physics versus training-distribution accidents, or whether labels are correct. The composition-only model’s applicability to ~40k Materials Project compositions without structures is practically useful for SHC screening. Strengths include falsifiable, DFT-adjudicated claims (especially HgOsPb2), bootstrap stability of the leading SHAP features, counterfactual PD that quantifies the Pt–p_frac entanglement (gradient drop 3.7→0.53 in Box-Cox units), and Rashomon agreement between RF and GPR. The data-audit idea—treating attribution-explicable model–label outliers as hypotheses about the data rather than noise to regularize away—is a valuable methodological contribution for curated materials datasets where one element dominates the high-property tail.","major_comments":[{"comment":"Abstract, §III.G.c, Table II, and Conclusion: the data-audit claim of a “thirtyfold error” in the HfC training label (Zhao 3.62 vs authors’ DFT 120), “inherited undetectably by every black-box model,” is load-bearing for the second pillar but under-determined by the present evidence. The paper itself documents large literature scatter for the same compound class (W3Ta: Zhao 1011 vs dedicated A15 work 2250; Table II and §III.G.e). Absolute intrinsic SHC is sensitive to structure, k-mesh, Wannier projection, energy window, and SOC pseudopotentials. The verification panel reports independent QE/PBE/fully-relativistic/WannierBerri (200³) peaks under the authors’ pipeline but does not recompute Zhao’s original HfC entry under matched structure and settings, nor quantify pipeline-to-pipeline variance on a shared control set. The model–label disagreement is a useful flag; the stronger claim tha","section":"§III.G.c, Table II, Abstract"},{"comment":"§III.B and comparison to Zhao et al.: the RF test MAE of 114.5 is compared to CGCNN 126.7 and Res-CGCNN 118.7 as if on the same learning problem. The authors apply polymorph collapse (9249→7515/7513), remove zeros, and train under Box-Cox with metrics inverse-transformed; Zhao’s reported numbers are on the original multi-polymorph set without that preprocessing. The manuscript correctly states it does not treat the margin as the contribution, but the abstract and introduction still frame the model as “reaching accuracy competitive with structure-aware graph networks.” Either retrain/report the graph baselines on the same reduced split and target transform, or remove/qualify the numerical competitiveness claim so that the audit—not the MAE race—carries the paper.","section":"§III.B, Abstract"},{"comment":"§II.B and §III.E, Table I: the bias-aware reranking (α0 sweep, additive β) is evaluated on the same held-out set used to choose the correction strength. The text acknowledges this is a diagnostic probe, not a generalizable method, yet Table I and the surrounding narrative still present recall gains as evidence of “screening consequences.” Because the central model-audit claim already rests on SHAP, counterfactual PD, RF–GPR agreement, and the HgOsPb2 DFT result, the probe is not needed for the main argument. Either move it fully to SI as a qualitative illustration, or add a nested hold-out / cross-split protocol so that any recall claim is not circular with the α0 sweep.","section":"§II.B, §III.E, Table I"}],"minor_comments":[{"comment":"Figure 4 caption: Shapley additivity is lost under inverse Box-Cox; panel (a) is shown in original SHC units “for clarity” while (b) retains fλ units. State explicitly how panel (a) was converted (e.g., mean |SHAP| in transformed space mapped approximately) so readers do not treat the two panels as commensurate.","section":"Fig. 4"},{"comment":"§II.A: s_min is introduced after observing that Magpie–AO gap disagreement correlates with strong hybridization in the training set. Briefly address whether this post-hoc engineering was locked before the final train/test split or could have leaked target information into the descriptor design.","section":"§II.A"},{"comment":"HgOsPb2 is reported 0.86 eV/atom above the Materials Project hull (§III.G.b). The under-prediction claim for the relaxed structure is still valid, but the screening narrative should more clearly separate “property of a metastable relaxed structure” from “synthesis-relevant candidate.”","section":"§III.G.b"},{"comment":"Table III mixes RF/GPR predictions with heterogeneous literature conventions (near-EF vs full-window peaks, magnetic vs nonmagnetic). A column noting the energy-window convention for each reference would reduce ambiguity when assessing underestimation († entries).","section":"Table III"},{"comment":"Typographical/spacing issues: “Themodel auditreveals”, “Thedata auditexposes”, “p f rac”, “pf rac” in figure captions and body; standardize p_frac / ⟨p⟩ notation throughout.","section":"Abstract, Fig. 5"},{"comment":"Data availability promises models and ~40k predictions “upon publication.” For a methods/audit paper, depositing the feature matrix, train/val/test indices, and SHAP scripts with the review package would strengthen reproducibility claims.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The model-audit half (Pt–p_frac entanglement + HgOsPb2) is the stronger, cleaner contribution and would stand alone with minor polishing. The data-audit half is conceptually important but currently oversold relative to the DFT controls; if the authors cannot do a matched HfC recompute, insisting on the ~30× “error inherited by every black-box model” language in the abstract risks a correctness challenge in review. Scope fit for a materials/cond-mat journal is good; the work is not pure ML methods theater. I would accept after the data-audit claim is either hardened with matched DFT or clearly demoted to a candidate discrepancy."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this paper is the closed loop, not another SHC predictor. They take SHAP, counterfactual partial dependence, and RF–GPR agreement, treat the findings as hypotheses, and settle them with a few targeted DFT runs. That is the real contribution: a practical QC protocol for property surrogates when one element dominates the high-property tail.\n\nWhat works. The composition-only RF hits test MAE 114.5 on the Zhao set, competitive with the reported CGCNN/Res-CGCNN numbers without needing relaxed structures, so it can screen ~40k Materials Project compositions. The model audit is clean: p_frac is entangled with Pt, the counterfactual PD gradient collapses (3.7 → 0.53), both RF and GPR show the same bias, and DFT on HgOsPb2 (peak 2703 vs RF 717) confirms the under-flagging. Agreement controls (VPt8, BiPt, K2Pb2O3) look sensible. The exploratory reranking is correctly labeled as a probe, not a method. Citations and methods are standard and readable.\n\nSoft spots, in proportion. The data-audit pillar is the weaker one. They flag HfC (Zhao label 3.62 vs their DFT 120) and call it a ~30× training-label error inherited by every black-box model. The model–label disagreement is real and useful as a flag, but the paper itself notes large literature scatter (W3Ta 1011 vs 2250). Absolute SHC is method-sensitive (structure, k-mesh, Wannier window, SOC PP). Without a matched recompute of Zhao’s HfC entry under the same settings, “thirtyfold label error” overclaims what the panel adjudicates. The DFT set is also small and selective. Artifacts are promised “upon publication,” not shipped yet. None of this sinks the protocol; it just means the data-audit language should be tightened.\n\nWho it is for: people building or trusting materials ML screens, especially transport properties. Worth a serious referee. I would engage, cite the protocol and the HgOsPb2 result, and ask for the matched HfC control plus code/data release. Send it to peer review.","headline":"Solid, usable audit loop for materials ML: composition-only SHC model competitive with graph nets, Pt–p_frac shortcut verified by DFT, data-audit claim on HfC slightly overstated without matched recompute.","tokens_in":20142,"tokens_out":558,"would_cite":true,"duration_ms":5096,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A model-agnostic audit with SHAP, counterfactuals, and a few DFT checks can expose both learned shortcuts and bad training labels in materials property models.","keywords":["spin Hall conductivity","machine learning audit","SHAP attribution","training-label error","compositional descriptors","density functional theory","materials informatics","Rashomon sets"],"falsifier":"Recompute the original HfC entry with the same structure, pseudopotentials, and Kubo-Wannier protocol used for the published training set, and check whether the label converges near 3.62 or near the authors’ ~120 (ℏ/e)(S/cm); a match to the low label would collapse the data-audit claim for that compound.","tokens_in":19945,"feed_emoji":"🔍","tokens_out":1054,"duration_ms":13557,"temperature":0.7,"pith_summary":"Machine-learning models used to screen materials rest on two untested assumptions: that their features track the real physics rather than quirks of the training set, and that the training labels themselves are correct. This paper introduces a practical audit that tests both, using feature attribution, counterfactual partial-dependence checks, agreement across very different model types, and a small number of independent first-principles calculations to settle each finding. Applied to spin Hall conductivity with a composition-only model that needs no crystal structure, the audit shows the model has entangled a p-valence descriptor with platinum content, systematically under-flagging high-response platinum-free compounds, and that one published training label is off by a factor of about thirty. A sympathetic reader cares because standard held-out error metrics certify neither assumption, so every black-box model trained on the same data can silently inherit both failures.","feed_headline":"Audit finds Pt shortcut and 30× label error in spin Hall ML","feed_subtitle":"SHAP plus a few DFT checks expose what standard validation never tests","key_machinery":"The audit protocol: global SHAP attribution plus counterfactual partial-dependence analysis plus Rashomon-style agreement between models with different inductive biases (Random Forest and Gaussian Process), with each flagged finding settled by a small number of independent DFT/Kubo-Wannier calculations.","core_discovery":"The authors show that a model-agnostic audit protocol—global SHAP attribution, counterfactual partial-dependence analysis, and Rashomon-style cross-model checks, with every finding adjudicated by targeted DFT—can diagnose both a learned element-as-proxy shortcut and a large training-label error in spin Hall conductivity models. On a composition-only Random Forest competitive with structure-aware graph networks, the audit finds that average p-valence becomes statistically entangled with Pt content; independent DFT confirms a Pt-free compound (HgOsPb2) whose true SHC is nearly four times the prediction. The same protocol flags a roughly thirtyfold error in the HfC training label, an error that","pith_inferences":["Curated materials datasets with sparse high-property tails dominated by one element are the natural regime where this audit pays off; the same pattern should appear for other SOC-driven responses.","A practical next step is to rebuild high-SHC training diversity with deliberately Pt-free chemistries so the p-valence descriptor can track hybridization rather than Pt markers.","Publishing per-entry DFT provenance (structure, k-mesh, Wannier settings) alongside labels would make data audits cheaper and more conclusive than re-running entire pipelines.","Screening shortlists from composition models should treat Pt-free high-tail candidates as higher-priority DFT targets precisely because the model is expected to under-flag them."],"forward_implications":["Composition-only models can match structure-aware graph-network accuracy for SHC while screening the much larger space of compositions that lack relaxed structures.","Wherever one element dominates the high-property tail, attribution-plus-counterfactual checks can reveal element-as-proxy shortcuts that standard MAE never flags.","A conspicuous, attribution-explicable model–label disagreement becomes a hypothesis about the data, not only about the model, and can be settled at the cost of one DFT run.","Black-box models trained on the same corrupted labels will inherit the same errors without any internal mechanism to detect them.","The protocol itself is model-agnostic and is meant to apply to structure-aware networks and other transport properties, not only to this Random Forest or SHC."],"fun_headline_variants":["Audit exposes Pt proxy shortcut and 30× HfC label error in SHC ML","SHAP+DFT audit finds Pt entanglement plus 30-fold training flaw","Composition ML for spin Hall hides Pt shortcut and 30× HfC error","Model audit flags p-valence-Pt link and 30× label error in SHC","DFT-checked protocol uncovers Pt proxy and HfC 30× bug in ML"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That a large gap between the authors’ independent DFT result and a published training label is mainly a label error rather than a difference in methods, conventions, or structure choices between calculation pipelines.","fun_headline_variants_meta":{"raw":{"variants":["Audit exposes Pt proxy shortcut and 30× HfC label error in SHC ML","SHAP+DFT audit finds Pt entanglement plus 30-fold training flaw","Composition ML for spin Hall hides Pt shortcut and 30× HfC error","Model audit flags p-valence-Pt link and 30× label error in SHC","DFT-checked protocol uncovers Pt proxy and HfC 30× bug in ML"]},"model":"grok-4.5","effort":"low","cost_usd":0.004116,"raw_usage":{"total_tokens":1286,"prompt_tokens":852,"num_sources_used":0,"completion_tokens":110,"cost_in_usd_ticks":41160000,"prompt_tokens_details":{"text_tokens":852,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":324,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":852,"tokens_out":110,"duration_ms":3519,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T14:39:37.679608+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Recompute the original HfC entry with the same structure, pseudopotentials, and Kubo-Wannier protocol used for the published training set, and check whether the label converges near 3.62 or near the authors’ ~120 (ℏ/e)(S/cm); a match to the low label would collapse the data-audit claim for that compound.","supporting_citations":[],"review_version":1}