{"id":"feba987a-f723-44d5-884b-c57c47a48dc1","arxiv_id":"2512.21077","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Active learning guided quantum calculations found Fe2TeSe as a predicted high-spin-Hall-conductivity 2D material, with computed SHC 271.52 ħ/e Ω⁻¹, about 23 times the initial best.","lead":"Researchers used a machine-learning 'active learning' loop paired with expensive quantum calculations to screen about 2,000 two-dimensional materials for a strong spin Hall effect, computing around 41 candidates and finding Fe2TeSe with a predicted spin Hall conductivity roughly 23 times higher than the best of the initial 24. The work shows how AI-guided search could cut the cost of discovering spintronic materials, though the headline number is a theoretical calculation wit","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Computed SHC values for metallic top candidates lack convergence and smearing validation; the 23x discovery claim rests on an unverified numerical protocol.","rationale":"The reader's weakest assumption exactly matches my assessment: the central quantitative claim hinges on the reliability of the computed SHC values. I considered other possible concerns—small training set, lack of a random-selection baseline, feature leakage, or the inconsistency in describing Round 1 as 'random' vs. 'domain-knowledge'—but none are as load-bearing as the numerical robustness of the ground-truth labels. If the SHC values are not converged with respect to k-mesh and smearing, then the 23x improvement, the ranking, and therefore the demonstration of active learning's effectiveness all collapse. The paper's own Figure 5 and text admit extreme sensitivity of Hall conductivity to E_F in metallic candidates, which are exactly the top performers. The lack of any convergence test is a concrete gap. The proposed test directly targets this gap. Since the reader already conditionally accepted with this concern, my verdict remains unchanged (CONDITIONAL). I agree with the reader's emphasis, and no additional major concern outweighs this one.","tokens_in":20231,"tokens_out":3245,"duration_ms":30976,"concrete_test":"Recompute SHC for Fe2TeSe, K2PtTe2, and AuSe using the same workflow while varying: (i) k-grid from 800×800×1 to 1200×1200×1 and 1600×1600×1; (ii) smearing parameter (e.g., 0, 0.01, 0.05 eV or Methfessel-Paxton/Gaussian); (iii) a small Fermi-level shift of ±10 meV (by adjusting electron count). If the SHC values change by more than ~20% or if the rank order changes, the 23x claim and the active-learning success are not robust. A supplementary benchmark against a known SHC material (e.g., bulk Pt or monolayer WTe2) would also calibrate the protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that active learning discovered Fe2TeSe with SHC 271.52 (ħ/e) Ω⁻¹, ~23× higher than the initial best (AuSe, 11.7)—depends entirely on the accuracy of the DFT–Wannier–TB SHC values. Methods report a single 800×800×1 k-grid and no smearing parameter, convergence tests, or error bars. For metallic systems, the intrinsic SHC is notoriously sensitive to the Fermi-level position and k-mesh density (small energy denominators). Figure 5 explicitly shows sharp fluctuations of SHC/OHC near E_F for the top candidates (e.g., K2PtTe2), and the Results acknowledge that 'rapid variation in properties such as Hall conductivity with even minor shifts in E_F' occurs for these metals. Because all top-10 candidates except one are metallic, the computed E_F values are potentially dominated by numerical noise. Additionally, the SOC strength λ is fitted to reproduce the DFT+SOC band structure, but no validation is shown for Fe2TeSe (only AuSe is illustrated in Fig. S6), and no error propagation from the fit is provided. If the true SHC of Fe2TeSe or K2PtTe2 were only modestly different (within a factor of ~2), the claimed monotonic improvement across rounds and the 23x headline could vanish. The paper itself labels trends 'indicative' due to the 41-system dataset, but the discovery claim is stated as quantitative. Thus, the load-bearing assumption is that the computed SHC values are numerically converged and robust; this is currently unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an active-learning workflow that combines DFT band structures, Wannier interpolation, and Kubo-formula SHC calculations with ridge-regression ML models to screen ~2000 2D materials from MC2D. Starting with 24 manually selected systems, three active-learning rounds (sampling 17 additional systems via expected improvement) expand the dataset to 41 systems and identify Fe2TeSe with computed SHC = 271.52 (hbar/e) Ohm^-1, ~23 times the Round-1 best (AuSe, 11.7). The authors also report chemical trends from SHAP and correlation analyses and make data and code publicly available.","tokens_in":20604,"tokens_out":3315,"duration_ms":34638,"significance":"If the numerical results are robust, the paper is a compelling demonstration that active learning can accelerate the discovery of 2D materials with large intrinsic SHC, which is a computationally expensive target. The workflow is reproducible (public GitHub data/code), the ground-truth DFT-Wannier-TB pipeline is state-of-the-art, and the round-over-round improvement is internally consistent. The SHAP analysis connects to physical intuition (d-orbital character, heavy elements, symmetry). However, the central quantitative claim—the 23x enhancement—rests on computed SHC values that are not shown to be converged or validated. For metallic systems, the intrinsic SHC is extremely sensitive to Fermi-level details and k-mesh sampling; the paper itself shows sharp fluctuations in Figure 5. Without convergence tests or comparison to prior SHC data, the ranking of candidates could change materially, undercutting the headline claim. Thus the significance depends on establishing the numerical reliability of the computed SHC values.","major_comments":[{"comment":"The SHC values are computed on a single 800x800x1 k-grid with no smearing parameter, no convergence tests, and no error bars. Figure 5 explicitly shows sharp fluctuations in SHC/OHC near EF for the top candidates (e.g., K2PtTe2), and the Results text acknowledges that these metallic systems are sensitive to minor EF shifts. Since 9 of the top-10 candidates are metallic, the 23x claim and the round-over-round ranking could be an artifact of the unconverged k-integration. Please provide k-grid and smearing convergence tests for at least the top candidates, and report error bars on the SHC values.","section":"Methods, Eq. (2)-(3); Figure 5"},{"comment":"The SOC parameter λ is fitted per material to match the DFT+SOC band structure, but validation of the TB fit is shown only for AuSe (Fig. S6). For Fe2TeSe, K2PtTe2, and the other top candidates, no comparison of TB vs DFT+SOC bands is shown, and there is no error propagation from the λ fit to the SHC. Please display the fit quality for all top candidates and quantify how SHC changes under small variations in λ and Fermi-level position.","section":"Results, Figure 3; Table S1, Figure S6"},{"comment":"The authors state that owing to the limited 41-system dataset, the trends should be treated as 'indicative', yet the Abstract and Results present 'nearly 23 times higher' as a quantitative discovery. The proof-of-concept is weakened by the small number of candidates per round (5-7) and the large ML uncertainties for the top candidates (e.g., Fe2TeSe and BaFe8As2 in Fig. 3 are far from the parity line, and MnC6N4Se2 deviates strongly). Please either temper the claim or add validation, for example by benchmarking the computed SHC values against published high-throughput 2D SHC data or by providing confidence intervals.","section":"Results, Chemical insights"}],"minor_comments":[{"comment":"The label 'Round 4' in Figure S2 is inconsistent with the three active-learning loops described in the main text; the figure should be renamed (e.g., 'Round 3' or 'Final ML model').","section":"Figure S2"},{"comment":"The phase factor in the TB Hamiltonian is written as exp(i k·d_j), which is a sign convention opposite to the usual tight-binding convention; please clarify the convention used and ensure consistency with the Wannier90 output.","section":"Methods, Eq. (1)"},{"comment":"The color/legend for Round 2.1 versus Round 2.2 is difficult to distinguish in grayscale; consider using different symbols or a colorblind-safe palette.","section":"Figure 2(b)"},{"comment":"The initial set is described as 'random but chemically diverse', while the Results section says the 24 systems were 'selected based on domain knowledge and prior literature'. Please reconcile these statements.","section":"Abstract and Results"},{"comment":"The sentence 'The SHC data generated in this work can be accessed at the GitHub Repository or is available as a separate excel sheet' does not provide a specific URL or repository identifier; please include the exact link and a DOI if available.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The work is a useful proof-of-concept, but the lack of convergence validation is a serious gap for a property as sensitive as SHC in metals. If the authors can provide k-grid/smearing convergence tests, error bars, and a comparison to existing 2D SHC data (e.g., Zhou et al., npj 2D Mater. Appl. 2025), the paper would be suitable for publication. There is also a tension between the 'random' initial set and the manual selection described in the text, and the number of materials screened per round is small. The paper would benefit from a more cautious presentation of the discovery claim until numerical robustness is established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core of this paper is a reasonable idea executed without fatal flaws: use expected improvement in an active-learning loop to screen ~2000 2D materials for spin Hall conductivity, with DFT-Wannier-TB calculations as ground truth. The round-over-round improvement in maximum computed SHC (AuSe 11.7 to K2PtTe2 195.56 to Fe2TeSe 271.52) is internally consistent and supported by the reported numbers. The ML workflow—ridge regression with RFE, ensemble uncertainty, and a parity check on screened candidates—is sensible, and the authors get credit for being explicit that the 41-system dataset means the chemical trends are indicative.\n\nThat said, the central quantitative claim is under-supported exactly where the reader's stress test lands. The Kubo calculations are reported with a single 800x800x1 k-grid and no smearing parameter or convergence tests. Figure 5 shows sharp fluctuations in SHC/OHC near E_F for the metallic top candidates, which means the reported values at E_F are sensitive to numerical details. Without a smearing analysis or a k-grid convergence test, I cannot trust the factor-of-23 claim, or even the ranking. This is not a hypothetical worry: 9 of the top 10 candidates are metallic, and small shifts in the Fermi level or SOC parameter can change Berry curvature by large factors.\n\nThe paper would be substantially more convincing with a benchmark against previously computed SHC values for known 2D materials (e.g., PtSe2, WSe2, or the earlier high-throughput dataset they cite). They also claim code/data availability but give no link or DOI—just \"GitHub Repository\"—which is not actionable.\n\nThere are also a few smaller signals that the manuscript needs tightening. The phonon calculation for Fe2TeSe (Figure S5) shows a negative frequency of 0.16 THz, which is usually interpreted as dynamical instability, not stability; calling it \"negligible\" is generous. The initial dataset is described sometimes as \"random\" and sometimes as \"manually curated,\" and the SI Figure S2 refers to a \"Round 4\" while the text says only three active-learning rounds. None of these are disqualifying, but they add up.\n\nIs the work worth a serious referee? Yes. The active-learning approach is new in this application, the internal logic is sound, and the candidate list is a useful starting point for more careful calculations. But the specific discovery claim should not be accepted as stated until the SHC protocol is validated. I would send it out, and I would ask referees to push on convergence, smearing, and external benchmarks.","headline":"The active-learning loop is internally consistent and a legitimate new application, but the headline 23x SHC claim rests on an under-validated numerical protocol; send to review with demands for convergence tests and external validation.","tokens_in":21096,"tokens_out":2158,"would_cite":false,"duration_ms":24930,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports an active-learning material search that, starting from 24 diverse 2D systems, found Fe2TeSe with a computed spin Hall conductivity of 271.52 (ħ/e) Ω⁻¹ — about 23 times the best value in the initial round.","keywords":["spin Hall conductivity","active learning","2D materials","density functional theory","Wannier tight-binding","expected improvement","spintronics","materials discovery"],"falsifier":"Recompute the SHC of Fe2TeSe and K2PtTe2 with a denser k-mesh, a different Wannier projection set, or a small rigid shift of the Fermi level; if the value at EF drops from 271.52 (ħ/e) Ω⁻¹ by a large factor, changes sign, or reorders the top candidates, the central claim fails. A second check is to rerun the active-learning loop from a different random seed or with a different ground-truth method: if it cannot beat the initial round's best candidate, the claimed acceleration is not robust.","tokens_in":20106,"feed_emoji":"🧲","tokens_out":6101,"duration_ms":60413,"temperature":0.7,"pith_summary":"The paper tries to show that active learning — an iterative loop in which a machine-learning model proposes candidates and density functional theory verifies them — can find 2D materials with very large spin Hall conductivity (SHC) without computing all ~2,000 candidates. Starting from 24 chemically diverse systems, the loop selected 41 materials over three rounds using an expected-improvement acquisition function, with SHC ground truth from DFT plus Wannier tight-binding and the Kubo formula. The authors report steady improvement each round, with Fe2TeSe reaching 271.52 (ħ/e) Ω⁻¹, nearly 23 times the best initial-round value (AuSe, 11.7). They also extract chemical trends: the strongest candidates tend to be metallic, dominated by d-orbitals near the Fermi level, and free of rotoinversion symmetry. If the computed values survive experimental scrutiny, this is evidence that active learning can replace broad high-throughput screening in expensive transport-property searches.","feed_headline":"Active-learning search finds 2D spin-Hall material 23x better","feed_subtitle":"An expected-improvement loop screened ~2,000 candidates with just 41 electronic-structure calculations.","key_machinery":"The load-bearing mechanism is the expected-improvement acquisition loop built around a ridge-regression surrogate. Features are 158 symmetry and elemental descriptors; the target is a symmetric-log transform of the computed SHC. Ground truth comes from a DFT-SOC calculation whose bands are fitted by a tight-binding Hamiltonian from maximally localized Wannier functions, with atomic spin-orbit parameter λ tuned to match the ab initio bands; the Kubo formula on an 800×800×1 k-grid then yields the spin (and orbital) Hall conductivity. Expected improvement is what drives the search: it selects candidates with high predicted SHC and high uncertainty, so each round adds the most informative extrem","core_discovery":"The central claim is that an expected-improvement active-learning loop can accelerate the discovery of high-SHC 2D materials by concentrating expensive electronic-structure calculations on the most promising candidates. The loop trains a ridge-regression model on SHC values computed from density-functional-theory bands, Wannier-fitted tight-binding Hamiltonians, and the Kubo formula with explicit spin-orbit coupling; it then scores the remaining ~2,000 candidate materials by expected improvement, which balances predicted SHC against model uncertainty, and sends 5–7 top scorers for full calculation. Across three rounds the computed SHC distribution shifts upward, and the best material, Fe2TeS","pith_inferences":["Editorial inference: The 271.52 value is a computed intrinsic SHC in the clean limit; disorder, temperature, and Fermi-level shifts from doping or gating could substantially alter it, so the 23x claim is a computational hypothesis awaiting transport experiments.","Editorial inference: The paper's Round 2.1 shows that targeting THC did not work as well as targeting SHC; a natural extension is to run the same active loop with the spin Hall angle or OHC-to-SHC conversion ratio as the objective, which might produce different optimal materials.","Editorial inference: Two of the top five candidates contain no heavy elements, which contradicts the usual 'heavy atoms are required for strong spin-orbit effects' heuristic; with only 41 samples this is a weak signal, but it suggests a targeted search over light-element compounds could surprise.","Editorial inference: The sharp Fermi-level sensitivity of the best metallic candidates means the same material could act as a tunable spin Hall switch via electrostatic gating; the authors note this tunability but stop short of proposing a device, and a transport calculation with a gate-induced μ shift would be a direct test."],"forward_implications":["A pool of ~2,000 2D materials can be screened with only 41 full electronic-structure calculations, suggesting this loop is a template for other expensive transport properties such as the anomalous Hall or orbital Hall effect.","The identified candidates, especially Fe2TeSe, are put forward as concrete targets for experimental spin-orbit-torque measurements and device testing.","The design signature 'metallic, d-orbital dominated near EF, no rotoinversion' is offered as a cheap screening filter for future high-SHC searches.","The public dataset and trained model let other groups score any 2D material for SHC before committing to expensive DFT calculations.","Because OHC is computed to dominate SHC in all 41 cases, experiments on these systems should be designed to disentangle spin-current torque from orbital-current conversion."],"fun_headline_variants":["AI-guided active learning finds 2D spin-Hall 23x better","41 calculations, 2000 candidates, 23x spin-Hall conductivity","Smart sampling uncovers 2D material with 23x spin Hall effect","Machine learning loop: 23x boost in spin Hall conductivity","Active learning: 41 calcs to find 2D spin-Hall star"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire ranking rests on the computed DFT–Wannier–Kubo spin Hall conductivity values being accurate and converged, especially for metallic candidates where the SHC fluctuates sharply near the Fermi level; if those numbers change with k-grid, smearing, or the fitted spin-orbit parameter λ, the 23-fold improvement and the candidate ordering collapse.","fun_headline_variants_meta":{"raw":{"variants":["AI-guided active learning finds 2D spin-Hall 23x better","41 calculations, 2000 candidates, 23x spin-Hall conductivity","Smart sampling uncovers 2D material with 23x spin Hall effect","Machine learning loop: 23x boost in spin Hall conductivity","Active learning: 41 calcs to find 2D spin-Hall star"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1624,"prompt_tokens":795,"completion_tokens":829,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":733}},"tokens_in":539,"tokens_out":829,"duration_ms":9045,"temperature":1.0,"reasoning_tokens":733,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T14:12:55.204330+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the SHC of Fe2TeSe and K2PtTe2 with a denser k-mesh, a different Wannier projection set, or a small rigid shift of the Fermi level; if the value at EF drops from 271.52 (ħ/e) Ω⁻¹ by a large factor, changes sign, or reorders the top candidates, the central claim fails. A second check is to rerun the active-learning loop from a different random seed or with a different ground-truth method: if it cannot beat the initial round's best candidate, the claimed acceleration is not robust.","supporting_citations":[],"review_version":1}