Pith. sign in

REVIEW 3 major objections 6 minor 34 references

A model-agnostic audit with SHAP, counterfactuals, and a few DFT checks can expose both learned shortcuts and bad training labels in materials property models.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:39 UTC pith:J6DKPNEE

load-bearing objection Solid, usable audit loop for materials ML: composition-only SHC model competitive with graph nets, Pt–p_frac shortcut verified by DFT, data-audit claim on HfC slightly overstated without matched recompute. the 3 major comments →

arxiv 2607.09910 v1 pith:J6DKPNEE submitted 2026-07-10 cond-mat.mtrl-sci

Auditing Machine-Learning Models and Their Training Data with Explainability and First-Principles Verification: Application to Spin Hall Conductivity

classification cond-mat.mtrl-sci
keywords spin Hall conductivitymachine learning auditSHAP attributiontraining-label errorcompositional descriptorsdensity functional theorymaterials informaticsRashomon sets
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Machine-learning models used to screen materials rest on two untested assumptions: that their features track the real physics rather than quirks of the training set, and that the training labels themselves are correct. This paper introduces a practical audit that tests both, using feature attribution, counterfactual partial-dependence checks, agreement across very different model types, and a small number of independent first-principles calculations to settle each finding. Applied to spin Hall conductivity with a composition-only model that needs no crystal structure, the audit shows the model has entangled a p-valence descriptor with platinum content, systematically under-flagging high-response platinum-free compounds, and that one published training label is off by a factor of about thirty. A sympathetic reader cares because standard held-out error metrics certify neither assumption, so every black-box model trained on the same data can silently inherit both failures.

Core claim

The authors show that a model-agnostic audit protocol—global SHAP attribution, counterfactual partial-dependence analysis, and Rashomon-style cross-model checks, with every finding adjudicated by targeted DFT—can diagnose both a learned element-as-proxy shortcut and a large training-label error in spin Hall conductivity models. On a composition-only Random Forest competitive with structure-aware graph networks, the audit finds that average p-valence becomes statistically entangled with Pt content; independent DFT confirms a Pt-free compound (HgOsPb2) whose true SHC is nearly four times the prediction. The same protocol flags a roughly thirtyfold error in the HfC training label, an error that

What carries the argument

The audit protocol: global SHAP attribution plus counterfactual partial-dependence analysis plus Rashomon-style agreement between models with different inductive biases (Random Forest and Gaussian Process), with each flagged finding settled by a small number of independent DFT/Kubo-Wannier calculations.

Load-bearing premise

That a large gap between the authors’ independent DFT result and a published training label is mainly a label error rather than a difference in methods, conventions, or structure choices between calculation pipelines.

What would settle it

Recompute the original HfC entry with the same structure, pseudopotentials, and Kubo-Wannier protocol used for the published training set, and check whether the label converges near 3.62 or near the authors’ ~120 (ℏ/e)(S/cm); a match to the low label would collapse the data-audit claim for that compound.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Composition-only models can match structure-aware graph-network accuracy for SHC while screening the much larger space of compositions that lack relaxed structures.
  • Wherever one element dominates the high-property tail, attribution-plus-counterfactual checks can reveal element-as-proxy shortcuts that standard MAE never flags.
  • A conspicuous, attribution-explicable model–label disagreement becomes a hypothesis about the data, not only about the model, and can be settled at the cost of one DFT run.
  • Black-box models trained on the same corrupted labels will inherit the same errors without any internal mechanism to detect them.
  • The protocol itself is model-agnostic and is meant to apply to structure-aware networks and other transport properties, not only to this Random Forest or SHC.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Curated materials datasets with sparse high-property tails dominated by one element are the natural regime where this audit pays off; the same pattern should appear for other SOC-driven responses.
  • A practical next step is to rebuild high-SHC training diversity with deliberately Pt-free chemistries so the p-valence descriptor can track hybridization rather than Pt markers.
  • Publishing per-entry DFT provenance (structure, k-mesh, Wannier settings) alongside labels would make data audits cheaper and more conclusive than re-running entire pipelines.
  • Screening shortlists from composition models should treat Pt-free high-tail candidates as higher-priority DFT targets precisely because the model is expected to under-flag them.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript proposes a model-agnostic audit protocol for materials ML that combines global SHAP attribution, counterfactual partial-dependence analysis, and Rashomon-style cross-model checks, with each finding adjudicated by targeted DFT. Applied to intrinsic spin Hall conductivity, a composition-only Random Forest (211-D descriptor; no relaxed structure) reaches a test MAE of 114.5 (ℏ/e)(S/cm) on a polymorph-reduced Zhao et al. set, competitive with reported structure-aware CGCNN/Res-CGCNN numbers. The model audit diagnoses statistical entanglement of the average p-valence descriptor with Pt content, consistent across RF and GPR; independent DFT on Pt-free HgOsPb2 yields a peak SHC of 2703 versus an RF prediction of 717. The data audit, triggered by a large model–label disagreement on HfC, reports an independent DFT peak of 120 versus the Zhao label 3.62 (~30×), argued to be inherited silently by black-box models trained on the same labels. Agreement controls (VPt8, BiPt, K2Pb2O3) and under-prediction cases (LiIr, W3Ta) are used to bound reliability.

Significance. If the dual-pillar protocol holds, it addresses a genuine gap: standard held-out metrics do not test whether features track physics versus training-distribution accidents, or whether labels are correct. The composition-only model’s applicability to ~40k Materials Project compositions without structures is practically useful for SHC screening. Strengths include falsifiable, DFT-adjudicated claims (especially HgOsPb2), bootstrap stability of the leading SHAP features, counterfactual PD that quantifies the Pt–p_frac entanglement (gradient drop 3.7→0.53 in Box-Cox units), and Rashomon agreement between RF and GPR. The data-audit idea—treating attribution-explicable model–label outliers as hypotheses about the data rather than noise to regularize away—is a valuable methodological contribution for curated materials datasets where one element dominates the high-property tail.

major comments (3)
  1. [§III.G.c, Table II, Abstract] Abstract, §III.G.c, Table II, and Conclusion: the data-audit claim of a “thirtyfold error” in the HfC training label (Zhao 3.62 vs authors’ DFT 120), “inherited undetectably by every black-box model,” is load-bearing for the second pillar but under-determined by the present evidence. The paper itself documents large literature scatter for the same compound class (W3Ta: Zhao 1011 vs dedicated A15 work 2250; Table II and §III.G.e). Absolute intrinsic SHC is sensitive to structure, k-mesh, Wannier projection, energy window, and SOC pseudopotentials. The verification panel reports independent QE/PBE/fully-relativistic/WannierBerri (200³) peaks under the authors’ pipeline but does not recompute Zhao’s original HfC entry under matched structure and settings, nor quantify pipeline-to-pipeline variance on a shared control set. The model–label disagreement is a useful flag; the stronger claim tha
  2. [§III.B, Abstract] §III.B and comparison to Zhao et al.: the RF test MAE of 114.5 is compared to CGCNN 126.7 and Res-CGCNN 118.7 as if on the same learning problem. The authors apply polymorph collapse (9249→7515/7513), remove zeros, and train under Box-Cox with metrics inverse-transformed; Zhao’s reported numbers are on the original multi-polymorph set without that preprocessing. The manuscript correctly states it does not treat the margin as the contribution, but the abstract and introduction still frame the model as “reaching accuracy competitive with structure-aware graph networks.” Either retrain/report the graph baselines on the same reduced split and target transform, or remove/qualify the numerical competitiveness claim so that the audit—not the MAE race—carries the paper.
  3. [§II.B, §III.E, Table I] §II.B and §III.E, Table I: the bias-aware reranking (α0 sweep, additive β) is evaluated on the same held-out set used to choose the correction strength. The text acknowledges this is a diagnostic probe, not a generalizable method, yet Table I and the surrounding narrative still present recall gains as evidence of “screening consequences.” Because the central model-audit claim already rests on SHAP, counterfactual PD, RF–GPR agreement, and the HgOsPb2 DFT result, the probe is not needed for the main argument. Either move it fully to SI as a qualitative illustration, or add a nested hold-out / cross-split protocol so that any recall claim is not circular with the α0 sweep.
minor comments (6)
  1. [Fig. 4] Figure 4 caption: Shapley additivity is lost under inverse Box-Cox; panel (a) is shown in original SHC units “for clarity” while (b) retains fλ units. State explicitly how panel (a) was converted (e.g., mean |SHAP| in transformed space mapped approximately) so readers do not treat the two panels as commensurate.
  2. [§II.A] §II.A: s_min is introduced after observing that Magpie–AO gap disagreement correlates with strong hybridization in the training set. Briefly address whether this post-hoc engineering was locked before the final train/test split or could have leaked target information into the descriptor design.
  3. [§III.G.b] HgOsPb2 is reported 0.86 eV/atom above the Materials Project hull (§III.G.b). The under-prediction claim for the relaxed structure is still valid, but the screening narrative should more clearly separate “property of a metastable relaxed structure” from “synthesis-relevant candidate.”
  4. [Table III] Table III mixes RF/GPR predictions with heterogeneous literature conventions (near-EF vs full-window peaks, magnetic vs nonmagnetic). A column noting the energy-window convention for each reference would reduce ambiguity when assessing underestimation († entries).
  5. [Abstract, Fig. 5] Typographical/spacing issues: “Themodel auditreveals”, “Thedata auditexposes”, “p f rac”, “pf rac” in figure captions and body; standardize p_frac / ⟨p⟩ notation throughout.
  6. [Data Availability] Data availability promises models and ~40k predictions “upon publication.” For a methods/audit paper, depositing the feature matrix, train/val/test indices, and SHAP scripts with the review package would strengthen reproducibility claims.

Circularity Check

1 steps flagged

No load-bearing circularity: model- and data-audit claims are diagnosed from SHAP/Rashomon analysis then independently adjudicated by new DFT, not defined by the training labels or fitted parameters.

specific steps
  1. fitted input called prediction [II.B Exploratory probe of screening consequences; Table I and surrounding text]
    "We emphasize that this is a diagnostic probe, not a proposed screening method: the correction strength is swept over a range and evaluated on the same set, which can reveal whether an exploitable signal exists but cannot yield a generalizable correction (this limitation is discussed with the results). ... bias-aware reranking with a minimal factor α0=1.2 raises recall from 0.680 to 0.760 at k=150 ..."

    α0 (and the additive β) are chosen by sweeping values and measuring recall@k on the identical held-out test set used for evaluation. Any modest lift is therefore partly selected rather than independently predicted; the paper correctly flags the limitation and does not promote the correction as a method, so the circularity is minor and non-load-bearing for the audit claims.

full rationale

The paper’s central chain is: train a composition-only RF (and independent GPR) on the Zhao et al. SHC labels after standard Box-Cox and polymorph reduction; apply global SHAP, counterfactual partial dependence, and RF–GPR agreement to diagnose a Pt–p_frac entanglement that is a property of the learned representation on this distribution; treat large model–label residuals as hypotheses about either the representation or the labels; settle both hypotheses with independent QE/PBE+SOC/WannierBerri calculations on new or re-examined compounds (HgOsPb2 under-prediction, HfC label discrepancy, agreement controls). None of these steps reduces by construction to its inputs. The DFT peaks are first-principles outputs under a stated pipeline, not refits of the Zhao labels or of the RF. The exploratory α0/β re-ranking is explicitly labelled a non-generalizable diagnostic probe evaluated on the same held-out set and is not used as a claimed screening method or as support for the main audit findings. Feature-importance-guided reduction to 211 dimensions is reported only as parity (not gain) relative to the fuller set. There is no self-citation that carries a uniqueness or existence claim, no ansatz smuggled from prior author work, and no renaming of a known empirical pattern as a new derivation. Minor transparency caveats (Box-Cox λ choice, same-set probe, literature scatter on W3Ta) affect correctness risk or generalizability but do not make any reported prediction or first-principles result equivalent to its inputs. Score 1 reflects only the acknowledged same-set diagnostic probe; the load-bearing audit results remain externally falsifiable.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 2 invented entities

The central claims rest on standard DFT/ML tooling plus a few design choices (Box-Cox, polymorph collapse, 211-feature construction, TreeSHAP/Rashomon agreement as diagnostic of entanglement). No new physical entities are postulated. Free parameters are preprocessing/model hyperparameters and the exploratory α0/β corrections, not constants fitted to invent the DFT-verified discrepancies. Domain assumptions include PBE+SOC adequacy for relative SHC comparisons and that Zhao labels share a comparable Kubo-Wannier convention.

free parameters (5)
  • Box-Cox λ (MLE on SHC targets)
    Chosen by maximum-likelihood on the training distribution; all fitting uses the transformed target. Affects loss weighting of the high-SHC tail though metrics are inverse-transformed.
  • RF/XGB/KRR/GPR hyperparameters
    Selected on the validation partition; primary RF performance and SHAP structure depend on these choices (authors report multi-split MAE stability).
  • Exploratory bias correction α0 (and additive β)
    Swept over {1.0…2.0} and β∈{100…700} on the same held-out set; not claimed as a general method but used to quantify screening consequence of the entanglement.
  • s_min engineered descriptor
    Introduced after observing Magpie vs atomic-orbital gap disagreement correlates with hybridization; one hand-designed feature among 211.
  • Feature-importance reduction to 211-D space
    Subset chosen by probing importances from a larger Magpie/combined space; accuracy parity claimed vs 403-D full space across eight splits.
axioms (6)
  • domain assumption Intrinsic SHC is adequately represented by the maximum absolute Kubo-Wannier tensor component as in Zhao et al., enabling direct model–label and model–DFT comparison.
    Training labels and primary DFT comparison use this convention (Methods; Table II).
  • domain assumption PBE GGA with fully relativistic SOC pseudopotentials and WannierBerri integration on dense k-grids yields SHC values reliable enough to adjudicate model and label errors at the reported factors (~4×, ~30×).
    First-principles section; verification pillar of both audits.
  • ad hoc to paper Polymorphs of a composition can be collapsed to one Materials Project lowest-hull entry (or dropped) without destroying the composition–SHC learning problem.
    Dataset construction: 9249→7515 compositions to avoid label noise/leakage in composition space.
  • domain assumption Agreement of qualitatively different models (RF axis-aligned splits vs GPR RBF) on bias direction strengthens that a dependency is representation/data-driven rather than learner-specific (Rashomon-style).
    Methods and local-explanation sections; used to elevate Pt–p_frac entanglement beyond one model class.
  • domain assumption TreeSHAP attributions and counterfactual permutation of p_frac within Pt-containing instances diagnose statistical entanglement usable as a falsifiable screening hypothesis.
    Core of the model-audit protocol; standard XAI tools applied as audit instruments.
  • standard math Standard regression/ML mathematics (Random Forests, GPR with RBF, MAE after inverse Box-Cox) applies without modification.
    Throughout Methods/Results.
invented entities (2)
  • Model-agnostic materials ML audit protocol (SHAP + counterfactual PD + Rashomon cross-model + targeted DFT adjudication) no independent evidence
    purpose: Jointly test whether features track physics vs training accidents and whether individual labels are correct.
    The paper’s primary methodological contribution; not a physical particle/force but a procedural entity. Independent evidence is partial: demonstrated on one property/dataset with a few DFT cases.
  • s_min (Magpie min-gap vs atomic-orbital gap agreement descriptor) no independent evidence
    purpose: Proxy for strong band hybridization in the compositional feature set.
    Engineered after observing correlation in the training set; no external falsifiable prediction beyond model utility.

pith-pipeline@v1.1.0-grok45 · 23067 in / 4312 out tokens · 41301 ms · 2026-07-14T14:39:37.679608+00:00 · methodology

0 comments
read the original abstract

Machine-learning models for materials properties rest on two assumptions that standard validation never tests: that a model's features reflect the physics of the property rather than accidents of the training distribution, and that the training labels are themselves correct. We introduce a model-agnostic audit protocol for both, combining SHAP attribution, counterfactual partial dependence analysis, and Rashomon-style cross-model verification, with every finding adjudicated by targeted density functional theory (DFT). Demonstrated on intrinsic spin Hall conductivity using a composition-only Random Forest, the model needs no relaxed crystal structure, reaching accuracy competitive with structure-aware graph networks while remaining applicable to the far larger space of compositions for which no structure has been computed. The model audit reveals that the average p-valence descriptor becomes statistically entangled with Pt content - a property of the learned representation rather than the physics; DFT confirms the consequence, a Pt-free compound (HgOsPb$_2$) whose true SHC is nearly four times the prediction. The data audit exposes a thirtyfold error in the HfC training label, inherited undetectably by every black-box model trained on the same data. The protocol audits a model and its training data for the cost of a few DFT calculations, wherever one element dominates the high-property regime.

Figures

Figures reproduced from arXiv: 2607.09910 by Mohammed Mahshook, Rudra Banerjee.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4 [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5 [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7 [PITH_FULL_IMAGE:figures/full_fig_p010_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8 [PITH_FULL_IMAGE:figures/full_fig_p012_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 3 linked inside Pith

  1. [1]

    Spin hall effects

    Jairo Sinova, Sergio O Valenzuela, Jörg Wunderlich, Craig H Back, and Tomáš Jungwirth. Spin hall effects. Reviews of modern physics, 87(4):1213–1260, 2015

  2. [2]

    Tuning the spin hall effect of pt from the moderately dirty to the superclean regime

    Edurne Sagasta, Yasutomo Omori, Miren Isasa, Martin Gradhand, Luis E Hueso, Yasuhiro Niimi, YoshiChika Otani, and Fèlix Casanova. Tuning the spin hall effect of pt from the moderately dirty to the superclean regime. Physical Review B, 94(6):060412, 2016

  3. [3]

    Reference

    Edurne Sagasta, Yasutomo Omori, Saül Vélez, Roger Llopis, Christopher Tollan, Andrey Chuvilin, Luis E Hueso, Martin Gradhand, YoshiChika Otani, and Fèlix Casanova. Unveiling the mechanisms of the spin hall ef- 12 TABLE II:First-principles validation of model predictions. SHC values are the maximum absolute component|σz xy|; the full-window peak is compare...

  4. [4]

    Giant intrinsic spin hall effect in w3ta and other a15 superconductors.Science Advances, 5(4):eaav8575, 2019

    E Derunova, Y Sun, C Felser, SSP Parkin, B Yan, and MN Ali. Giant intrinsic spin hall effect in w3ta and other a15 superconductors.Science Advances, 5(4):eaav8575, 2019

  5. [5]

    Dirac surface states, multiorbital dimerization, and superconductivity in nb-and ta-based a15 compounds.Physical Review B, 109(7):075119, 2024

    Raghottam M Sattigeri, Giuseppe Cuono, Ghulam Hus- sain, Xing Ming, Angelo Di Bernardo, Carmine Attana- sio, Mario Cuoco, and Carmine Autieri. Dirac surface states, multiorbital dimerization, and superconductivity in nb-and ta-based a15 compounds.Physical Review B, 109(7):075119, 2024

  6. [6]

    High spin hall conductivity in large-area type-ii dirac semimetal ptte2.Advanced Ma- terials, 32(17):2000513, 2020

    Hongjun Xu, Jinwu Wei, Hengan Zhou, Jiafeng Feng, Teng Xu, Haifeng Du, Congli He, Yuan Huang, Junwei Zhang, Yizhou Liu, et al. High spin hall conductivity in large-area type-ii dirac semimetal ptte2.Advanced Ma- terials, 32(17):2000513, 2020

  7. [7]

    Intrinsic spin hall conductivity of mote2 and wte2 semimetals.arXiv preprint arXiv:1812.06910, 2018

    Jiaqi Zhou, Junfeng Qiao, Arnaud Bournel, and Weisheng Zhao. Intrinsic spin hall conductivity of mote2 and wte2 semimetals.arXiv preprint arXiv:1812.06910, 2018

  8. [8]

    Meng Zhu, Xinlu Li, Fanxing Zheng, Jianting Dong, Ye Zhou, Kun Wu, and Jia Zhang. Crystal facet orientation and temperature dependence of charge and spin hall effects in noncollinear antiferromagnets: A first-principles investigation.Physical Review B, 110(5):054420, 2024

  9. [9]

    Dif- ferent types of spin currents in the comprehensive mate- rials database of nonmagnetic spin hall effect.npj Com- putational Materials, 7(1):167, 2021

    Yang Zhang, Qiunan Xu, Klaus Koepernik, Roman Rezaev, Oleg Janson, Jakub Železn` y, Tomáš Jungwirth, Claudia Felser, Jeroen van den Brink, and Yan Sun. Dif- ferent types of spin currents in the comprehensive mate- rials database of nonmagnetic spin hall effect.npj Com- putational Materials, 7(1):167, 2021

  10. [10]

    Kubo formula for dc conductivity: Generalization to sys- tems with spin-orbit coupling.Physical Review Research, 6(1):L012057, 2024

    IA Ado, M Titov, Rembert A Duine, and Arne Brataas. Kubo formula for dc conductivity: Generalization to sys- tems with spin-orbit coupling.Physical Review Research, 6(1):L012057, 2024

  11. [11]

    Calculation of intrinsic spin hall conductivity by wannier interpolation.Physical Review B, 98(21):214402, 2018

    JunfengQiao, JiaqiZhou, ZheYuan, andWeishengZhao. Calculation of intrinsic spin hall conductivity by wannier interpolation.Physical Review B, 98(21):214402, 2018

  12. [12]

    Accelerating spin hall conductivity predictions via machine learning

    Jinbin Zhao, Junwen Lai, Jiantao Wang, Yi-Chi Zhang, Junlin Li, Xing-Qiu Chen, and Peitao Liu. Accelerating spin hall conductivity predictions via machine learning. Materials Genome Engineering Advances, 2(4):e67, 2024

  13. [13]

    Predicting crystal spin hall conductivity using a multi-modal transformer with real- space and reciprocal-space fingerprints.Physical Review Materials, 9(9):093804, 2025

    Ying Zhang, Yifan Li, Xiwen Zhang, Yi-Ming Zhao, Yi- hong Wu, and Lei Shen. Predicting crystal spin hall conductivity using a multi-modal transformer with real- space and reciprocal-space fingerprints.Physical Review Materials, 9(9):093804, 2025

  14. [14]

    Spin hall conductivity and anomalous hall con- ductivity in full heusler compounds.New Journal of Physics, 24(5):053027, 2022

    Yimin Ji, Wenxu Zhang, Hongbin Zhang, and Wanli Zhang. Spin hall conductivity and anomalous hall con- ductivity in full heusler compounds.New Journal of Physics, 24(5):053027, 2022

  15. [15]

    Lux, Thorben Pürling, Abi- gail Morrison, Stefan Blügel, Daniele Pinna, and Yuriy Mokrousov

    Jonathan Kipp, Fabian R. Lux, Thorben Pürling, Abi- gail Morrison, Stefan Blügel, Daniele Pinna, and Yuriy Mokrousov. Machine learning inspired models for Hall effects in non-collinear magnets.Machine Learning: Sci- ence and Technology, 5(2):025060, 2024. 13 Formula RF (predicted) GP (predicted) GP variance GP confidence Reported SHC Pt[2] 1603.5 834.5 16...

  16. [16]

    Sisso: a compressed-sensing method for identifying the best low- dimensional descriptor in an immensity of offered candi- dates.Physical Review Materials, 2(8):083802, 2018

    Runhai Ouyang, Stefano Curtarolo, Emre Ahmetcik, Matthias Scheffler, and Luca M Ghiringhelli. Sisso: a compressed-sensing method for identifying the best low- dimensional descriptor in an immensity of offered candi- dates.Physical Review Materials, 2(8):083802, 2018

  17. [17]

    Lundberg and Su-In Lee

    Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. InAdvances in Neural Information Processing Systems, volume 30, pages 4765–

  18. [18]

    Curran Associates, Inc., 2017

  19. [19]

    Machine learning in materials sci- ence: From explainable predictions to autonomous de- sign.Computational Materials Science, 193:110360, 2021

    Ghanshyam Pilania. Machine learning in materials sci- ence: From explainable predictions to autonomous de- sign.Computational Materials Science, 193:110360, 2021

  20. [20]

    Statistical modeling: The two cultures (with comments and a rejoinder by the author).Statis- tical Science, 16(3):199–231, 2001

    Leo Breiman. Statistical modeling: The two cultures (with comments and a rejoinder by the author).Statis- tical Science, 16(3):199–231, 2001

  21. [21]

    On the existence of simpler machine learning models

    Lesia Semenova, Cynthia Rudin, and Ronald Parr. On the existence of simpler machine learning models. InPro- ceedings of the 2022 ACM Conference on Fairness, Ac- countability, and Transparency, FAccT ’22, pages 1827– 1858, New York, NY, USA, 2022. Association for Com- puting Machinery

  22. [22]

    A general-purpose machine learning framework for predicting properties of inorganic materials.npj Computational Materials, 2(1):1–7, 2016

    Logan Ward, Ankit Agrawal, Alok Choudhary, and Christopher Wolverton. A general-purpose machine learning framework for predicting properties of inorganic materials.npj Computational Materials, 2(1):1–7, 2016. 14

  23. [23]

    Matminer: An open source toolkit for materials data mining.Computational Materials Science, 152:60– 69, 2018

    Logan Ward, Alexander Dunn, Alireza Faghaninia, Nils ER Zimmermann, Saurabh Bajaj, Qi Wang, Joseph Montoya, Jiming Chen, Kyle Bystrom, Maxwell Dylla, et al. Matminer: An open source toolkit for materials data mining.Computational Materials Science, 152:60– 69, 2018

  24. [24]

    Combina- torial screening for new materials in unconstrained com- position space with machine learning.Physical Review B, 89(9):094104, 2014

    Bryce Meredig, Ankit Agrawal, Scott Kirklin, James E Saal, Jeff W Doak, Alan Thompson, Kunpeng Zhang, Alok Choudhary, and Christopher Wolverton. Combina- torial screening for new materials in unconstrained com- position space with machine learning.Physical Review B, 89(9):094104, 2014

  25. [25]

    From local explanations to global understanding with explainable ai for trees.Nature machine intelligence, 2(1):56–67, 2020

    Scott M Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee. From local explanations to global understanding with explainable ai for trees.Nature machine intelligence, 2(1):56–67, 2020

  26. [26]

    Equality of opportunity in supervised learning

    Moritz Hardt, Eric Price, and Nati Srebro. Equality of opportunity in supervised learning. InAdvances in Neu- ral Information Processing Systems, volume 29, pages 3315–3323. Curran Associates, Inc., 2016

  27. [27]

    Rosenbaum and Donald B

    Paul R. Rosenbaum and Donald B. Rubin. The central role of the propensity score in observational studies for causal effects.Biometrika, 70(1):41–55, 1983

  28. [28]

    P Giannozzi, S Baroni, N Bonini, M Calandra, R Car, and C Cavazzoni. J. phys.: Condens. matter 21 395502. J. Phys.: Condens. Matter, 21:395502, 2009

  29. [29]

    ukbenli, m

    P Giannozzi, O Andreussi, T Brumme, O Bunau, M Buongiorno Nardelli, M Calandra, R Car, C Cavaz- zoni, D Ceresoli, M Cococcioni, et al. ukbenli, m. lazzeri, m. marsili, n. marzari, f. mauri, nl nguyen, h.V. Nguyen, A. Otero-de-la-Roza, L. Paulatto, S. Ponce, D. Rocca, R. Sabatini, B. Santra, M. Schlipf, AP Seitsonen, A. Smo- gunov, I. Timrov, T. Thonhauser...

  30. [30]

    High performance wannier interpola- tion of berry curvature and related quantities with wan- nierberri code.npj Computational Materials, 7(1):33, 2021

    Stepan S Tsirkin. High performance wannier interpola- tion of berry curvature and related quantities with wan- nierberri code.npj Computational Materials, 7(1):33, 2021

  31. [31]

    Giant intrinsic spin hall effect in the taas family of weyl semimetals.arXiv preprint arXiv:1604.07167, 2016

    Yan Sun, Yang Zhang, Claudia Felser, and Binghai Yan. Giant intrinsic spin hall effect in the taas family of weyl semimetals.arXiv preprint arXiv:1604.07167, 2016

  32. [32]

    Intrinsic spin hall effect in topological insulators: A first-principles study

    SM Farzaneh and Shaloo Rakheja. Intrinsic spin hall effect in topological insulators: A first-principles study. Physical Review Materials, 4(11):114202, 2020

  33. [33]

    Tunable giant spin hall conductivities in a strong spin-orbit semimetal: Bi 1-x sb x.Physical review letters, 114(10):107201, 2015

    Cüneyt Şahin and Michael E Flatté. Tunable giant spin hall conductivities in a strong spin-orbit semimetal: Bi 1-x sb x.Physical review letters, 114(10):107201, 2015

  34. [34]

    Current-induced polarization and the spin hall effect at room temperature.Physical review letters, 97(12):126603, 2006

    NP Stern, S Ghosh, G Xiang, M Zhu, N Samarth, and DD Awschalom. Current-induced polarization and the spin hall effect at room temperature.Physical review letters, 97(12):126603, 2006