Pith. sign in

REVIEW 4 major objections 5 minor

Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides

T0 review · 4 major / 5 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read Coordination-aware tree ensembles predict spinel thermodynamics and magnetism from Materials Project data, but not band gaps.

desk verdict Careful spinel surrogate work with group-aware protocol and an honest negative band-gap result; useful screening numbers, but abstract-only so leakage claims stay unverified. read the letter →

arxiv 2607.11996 v1 pith:S7Z634CN submitted 2026-07-13 cond-mat.mtrl-sci

classification cond-mat.mtrl-sci
keywords spineloxidessurrogatemodelsformationenergyconvexhullmagnetizationbandgapcoordinationchemistryMaterialsProject
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds surrogate models that predict key properties of cubic spinels—oxides, sulfides, selenides, and related compounds—without running new density-functional calculations for every candidate. The authors curated 320 Materials Project entries, assigned cations to tetrahedral-like and octahedral-like sites by local coordination rather than by chemical identity, and trained tree ensembles under strictly group-aware evaluation so that no reduced formula leaks between train and test. On repeated held-out groups the models recover formation energies to roughly 0.12 eV/atom, distances to the convex hull to 0.05 eV/atom, and magnetizations to about 1.3 µB per formula unit, while correctly classifying metals versus non-metals most of the time. Band-gap regression fails to beat a trivial baseline, a negative result the authors attribute to scarce non-metal examples and the known underestimation of gaps by semi-local DFT. The same protocol also shows that a carefully tuned tabular model can outperform an untuned graph network at this data scale, and SHAP analysis plus conformal intervals map which chemistries the surrogates can be trusted for.

What carries the argument

Coordination-resolved tabular descriptors: cations are partitioned into tetrahedral-like and octahedral-like groups by CrystalNN coordination numbers rather than by element identity, then fed to carefully regularized tree ensembles whose entire preprocessing and model-selection pipeline is fit only on training folds of group-aware splits.

What would settle it

A new set of higher-rung (hybrid or GW) calculations, or experimental measurements, on a chemically diverse hold-out of spinels that shows the reported formation-energy or magnetization MAEs rising well above the stated error bars, or that shows the band-gap model suddenly beating a constant predictor once better labels are used.

Watch

Extended reading notes

Core claim

Under repeated grouped holdouts that keep every reduced formula entirely on one side of the split, tree-ensemble surrogates trained on CrystalNN coordination features reach mean absolute errors of 0.121 ± 0.030 eV/atom for formation energy, 0.048 ± 0.013 eV/atom for hull distance, and 1.27 ± 0.19 µB/fu for magnetization, with metallicity accuracy 0.85 ± 0.06; band-gap regression does not beat a trivial baseline on the 19-member non-metal holdout.

Load-bearing premise

The semi-local DFT labels and CrystalNN site assignments from Materials Project are accurate and consistent enough that models trained on them remain useful for real screening of spinels.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript curates 320 cubic (Fd-3m) spinel entries from the Materials Project (nitrides, oxides, sulfides, selenides, including mixed-valence A3X4) and trains tree-ensemble surrogates for formation energy, energy above hull, band gap, total magnetization, and metallicity. Cations are assigned to tetrahedral-like and octahedral-like groups via CrystalNN coordination rather than element identity. Evaluation is group-aware: splits by reduced formula, transforms fit on training folds only, and champions selected on CV before holdout. Over twenty repeated grouped holdouts the reported MAEs are 0.121±0.030 eV/atom (formation energy), 0.048±0.013 eV/atom (hull), 1.27±0.19 μB/fu (magnetization), with metallicity accuracy 0.85±0.06; band-gap regression fails to beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap. On the same split the tabular champion beats an untuned MEGNet trained from scratch on formation energy. SHAP, grouped conformal intervals, permutation nulls, and leave-one-chemistry-out tests are used to interpret models and bound applicability.

Significance. If the leakage-free grouped protocol and reported errors hold under full scrutiny, the work supplies practical, chemistry-aware surrogates for thermodynamic stability and magnetization of cubic spinels at a data scale where graph networks are not yet decisive, and it documents an honest negative result for band-gap regression. Strengths that should be credited include: reduced-formula grouping, train-only transforms, champion selection before holdout, paired bootstrap for the band-gap null, explicit comparison to untuned MEGNet on an identical split, and post-hoc domain bounds (SHAP, conformal, leave-one-chemistry-out). These practices raise the bar for small-N materials ML reporting even if the absolute errors remain modest and label-limited by semi-local DFT.

major comments (4)
  1. The central MAE and accuracy claims (formation energy 0.121±0.030 eV/atom, hull 0.048±0.013 eV/atom, magnetization 1.27±0.19 μB/fu, metallicity 0.85±0.06 over twenty repeated grouped holdouts) are load-bearing and rest on the assertion that reduced-formula grouping, CrystalNN site assignment, feature transforms, hyperparameter tuning, and champion selection all stayed strictly inside training folds. The abstract states this protocol and notes mild optimism from single-seed refits, but does not supply fold-size statistics, per-fold chemistry composition, or leakage diagnostics confirming that coordination numbers and electronegativity descriptors were recomputed solely on train structures. With N=320 and chemistry-grouped splits, even modest global pre-processing would bias the holdout numbers. The full methods must document these steps with enough detail (and preferably code or fold indi
  2. The band-gap negative result is appropriately framed (does not beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap) and attributed to sample scarcity plus semi-local DFT labels. That honesty is a strength, but the same scarcity implies that any screening use case for non-metals is currently unsupported. The manuscript should state explicitly which screening decisions the surrogates are and are not intended to support, and whether metallicity classification (accuracy 0.85±0.06) is the recommended substitute for gap regression.
  3. The MEGNet comparison (tabular champion 0.087 vs untuned MEGNet 0.209 eV/atom formation-energy MAE on the primary holdout) is correctly described as bounding rather than settling the descriptor-versus-graph question. Because the graph model is untuned and trained from scratch on this small set, the comparison cannot be read as evidence that tabular features are intrinsically superior. Either a modest hyperparameter search for MEGNet on the same grouped protocol should be added, or the claim should remain strictly limited to “untuned MEGNet from scratch” throughout the discussion and abstract.
  4. Applicability is described as uneven across anions and cations (grouped conformal intervals, permutation nulls, leave-one-chemistry-out). Those diagnostics are load-bearing for any claim of transfer beyond the curated 320-entry set. The full text must report which anion/cation groups drive the leave-one-chemistry-out failures and whether the conformal coverage is calibrated under the same grouped splits; without that, the domain-of-applicability paragraph remains qualitative.
minor comments (5)
  1. Abstract: clarify whether the twenty repeated holdouts re-draw the grouped splits each time or only re-fit with fixed splits and fixed hyperparameters; the phrase “single-seed refits that reuse the tuned hyperparameters” is slightly ambiguous.
  2. Abstract: “nitrides, oxides, sulfides, and selenides” appears in the opening list while the title emphasizes oxides, sulfides, and selenides; confirm nitride count and whether nitride chemistry is retained in all reported metrics.
  3. Notation: define μB/fu and eV/atom consistently on first use in the main text; the abstract mixes “μBfu” and “eV/atom” styles.
  4. When the full manuscript is prepared, include a small table of dataset composition (counts by anion, by A/B cation family, metal vs non-metal) so readers can judge the 19-member non-metal holdout in context.
  5. Cite the specific Materials Project database version and the CrystalNN settings used for tetrahedral/octahedral assignment so the 320-entry curation is reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: standard supervised ML on external Materials Project labels with group-aware holdouts; abstract-only review finds no self-definitional or fitted-as-prediction reductions.

full rationale

The paper reports tree-ensemble surrogates trained on 320 cubic spinel entries curated from the Materials Project (external DFT labels for formation energy, hull distance, band gap, magnetization, metallicity). Cation site assignments come from CrystalNN coordination numbers; evaluation uses reduced-formula grouped splits, transforms fit only on training folds, and champion selection on CV before holdout examination. Reported MAEs (0.121±0.030 eV/atom formation energy, 0.048±0.013 eV/atom hull, 1.27±0.19 μB/fu magnetization, metallicity accuracy 0.85±0.06) and the negative band-gap result versus a trivial baseline are empirical performance numbers under that protocol, not quantities derived by construction from fitted constants or from self-cited uniqueness theorems. SHAP attributions, conformal intervals, permutation nulls, and leave-one-chemistry-out tests are post-hoc diagnostics. No equations equate a claimed prediction to an input by definition; no ansatz is smuggled via self-citation; no uniqueness theorem is imported from the authors. The only available text is the abstract, which describes a conventional supervised-learning pipeline on external labels. Reader and skeptic concerns about possible leakage or unverifiable protocol details are correctness/reproducibility risks, not circularity. Score 0 is the honest finding for an abstract-only review of this type of work.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

Central claims rest on Materials Project DFT labels as ground truth, CrystalNN coordination assignment as the site descriptor, and standard tree-ensemble inductive bias. No new physical entities. Free parameters are the usual ML hyperparameters and the curated 320-entry set itself; none are numerically disclosed in the abstract.

free parameters (2)
  • tree-ensemble hyperparameters (depth, learning rate, etc.)
    Champion models are selected by cross-validated scores; exact tuned values not given in abstract but are free parameters of the surrogates.
  • curated 320-entry spinel set membership
    Which MP entries are included or excluded (nitrides/oxides/sulfides/selenides, single-cation mixed-valence A3X4) defines the training distribution; selection criteria are summarized but not exhaustively listed in the abstract.
assumptions (3)
  • domain assumption Materials Project semi-local DFT formation energies, hull distances, band gaps, and magnetizations are adequate labels for surrogate training and for the intended screening use.
    All targets are MP-derived; abstract itself attributes band-gap failure partly to these labels.
  • domain assumption CrystalNN coordination numbers correctly assign cations to tetrahedral-like vs octahedral-like groups for Fd-3m spinels.
    Feature construction depends on this assignment rather than element identity.
  • standard math Grouped splits by reduced formula plus train-only transforms yield unbiased estimates of generalization to new compositions.
    Standard ML validation assumption; abstract emphasizes this protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides." pith.science (2026). https://pith.science/paper/S7Z634CN

@misc{pith2026260711996,
  author       = {Pith},
  title        = {Pith review of: Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/S7Z634CN}},
  note         = {Machine review of arXiv:2607.11996}
}
abstract

We curated 320 cubic ($Fd\bar{3}m$) spinel entries from the Materials Project-nitrides, oxides, sulfides, and selenides, including single-cation mixed-valence $A_3X_4$ compounds-and trained tree-ensemble surrogates for the formation energy, energy above the convex hull, band gap, total magnetization, and metallicity. Cations were assigned to tetrahedral-like and octahedral-like groups from CrystalNN coordination numbers rather than from element identity, and the evaluation was group-aware throughout: splits were grouped by reduced formula, every transform was fit on training folds only, and champion models were selected on cross-validated scores before the holdout was examined. Over twenty repeated grouped holdouts (single-seed refits that reuse the tuned hyperparameters, and are therefore mildly optimistic) the champions reach mean absolute errors of $0.121\pm0.030$ eV/atom for the formation energy, $0.048\pm0.013$ eV/atom for the hull distance, and $1.27\pm0.19$~\muBfu{} for the magnetization, with a metallicity accuracy of $0.85\pm0.06$. Band-gap regression does not beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap testing, we report this negative result and trace it to sample scarcity and to the semi-local DFT labels. On the identical grouped split, the tabular champion is more accurate than an untuned MEGNet trained from scratch (0.087 versus 0.209 eV/atom formation-energy MAE on the primary holdout), a comparison that bounds, rather than settles, the descriptor-versus-graph question at this data scale. SHAP attribution ties the magnetization model to octahedral $d$-occupancy and the formation-energy and band-gap models to electronegativity descriptors, and grouped conformal intervals, permutation nulls, and leave-one-chemistry-out tests bound a domain of applicability that is uneven across anions and cations.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.