REVIEW 4 major objections 5 minor
Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides
T0 review · 4 major / 5 minor · reviewed 2026-07-15 · grok-4.5
Pith's one-line read Coordination-aware tree ensembles predict spinel thermodynamics and magnetism from Materials Project data, but not band gaps.
desk verdict Careful spinel surrogate work with group-aware protocol and an honest negative band-gap result; useful screening numbers, but abstract-only so leakage claims stay unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Coordination-resolved tabular descriptors: cations are partitioned into tetrahedral-like and octahedral-like groups by CrystalNN coordination numbers rather than by element identity, then fed to carefully regularized tree ensembles whose entire preprocessing and model-selection pipeline is fit only on training folds of group-aware splits.
What would settle it
A new set of higher-rung (hybrid or GW) calculations, or experimental measurements, on a chemically diverse hold-out of spinels that shows the reported formation-energy or magnetization MAEs rising well above the stated error bars, or that shows the band-gap model suddenly beating a constant predictor once better labels are used.
Extended reading notes
Core claim
Under repeated grouped holdouts that keep every reduced formula entirely on one side of the split, tree-ensemble surrogates trained on CrystalNN coordination features reach mean absolute errors of 0.121 ± 0.030 eV/atom for formation energy, 0.048 ± 0.013 eV/atom for hull distance, and 1.27 ± 0.19 µB/fu for magnetization, with metallicity accuracy 0.85 ± 0.06; band-gap regression does not beat a trivial baseline on the 19-member non-metal holdout.
Load-bearing premise
The semi-local DFT labels and CrystalNN site assignments from Materials Project are accurate and consistent enough that models trained on them remain useful for real screening of spinels.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript curates 320 cubic (Fd-3m) spinel entries from the Materials Project (nitrides, oxides, sulfides, selenides, including mixed-valence A3X4) and trains tree-ensemble surrogates for formation energy, energy above hull, band gap, total magnetization, and metallicity. Cations are assigned to tetrahedral-like and octahedral-like groups via CrystalNN coordination rather than element identity. Evaluation is group-aware: splits by reduced formula, transforms fit on training folds only, and champions selected on CV before holdout. Over twenty repeated grouped holdouts the reported MAEs are 0.121±0.030 eV/atom (formation energy), 0.048±0.013 eV/atom (hull), 1.27±0.19 μB/fu (magnetization), with metallicity accuracy 0.85±0.06; band-gap regression fails to beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap. On the same split the tabular champion beats an untuned MEGNet trained from scratch on formation energy. SHAP, grouped conformal intervals, permutation nulls, and leave-one-chemistry-out tests are used to interpret models and bound applicability.
Significance. If the leakage-free grouped protocol and reported errors hold under full scrutiny, the work supplies practical, chemistry-aware surrogates for thermodynamic stability and magnetization of cubic spinels at a data scale where graph networks are not yet decisive, and it documents an honest negative result for band-gap regression. Strengths that should be credited include: reduced-formula grouping, train-only transforms, champion selection before holdout, paired bootstrap for the band-gap null, explicit comparison to untuned MEGNet on an identical split, and post-hoc domain bounds (SHAP, conformal, leave-one-chemistry-out). These practices raise the bar for small-N materials ML reporting even if the absolute errors remain modest and label-limited by semi-local DFT.
major comments (4)
- The central MAE and accuracy claims (formation energy 0.121±0.030 eV/atom, hull 0.048±0.013 eV/atom, magnetization 1.27±0.19 μB/fu, metallicity 0.85±0.06 over twenty repeated grouped holdouts) are load-bearing and rest on the assertion that reduced-formula grouping, CrystalNN site assignment, feature transforms, hyperparameter tuning, and champion selection all stayed strictly inside training folds. The abstract states this protocol and notes mild optimism from single-seed refits, but does not supply fold-size statistics, per-fold chemistry composition, or leakage diagnostics confirming that coordination numbers and electronegativity descriptors were recomputed solely on train structures. With N=320 and chemistry-grouped splits, even modest global pre-processing would bias the holdout numbers. The full methods must document these steps with enough detail (and preferably code or fold indi
- The band-gap negative result is appropriately framed (does not beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap) and attributed to sample scarcity plus semi-local DFT labels. That honesty is a strength, but the same scarcity implies that any screening use case for non-metals is currently unsupported. The manuscript should state explicitly which screening decisions the surrogates are and are not intended to support, and whether metallicity classification (accuracy 0.85±0.06) is the recommended substitute for gap regression.
- The MEGNet comparison (tabular champion 0.087 vs untuned MEGNet 0.209 eV/atom formation-energy MAE on the primary holdout) is correctly described as bounding rather than settling the descriptor-versus-graph question. Because the graph model is untuned and trained from scratch on this small set, the comparison cannot be read as evidence that tabular features are intrinsically superior. Either a modest hyperparameter search for MEGNet on the same grouped protocol should be added, or the claim should remain strictly limited to “untuned MEGNet from scratch” throughout the discussion and abstract.
- Applicability is described as uneven across anions and cations (grouped conformal intervals, permutation nulls, leave-one-chemistry-out). Those diagnostics are load-bearing for any claim of transfer beyond the curated 320-entry set. The full text must report which anion/cation groups drive the leave-one-chemistry-out failures and whether the conformal coverage is calibrated under the same grouped splits; without that, the domain-of-applicability paragraph remains qualitative.
minor comments (5)
- Abstract: clarify whether the twenty repeated holdouts re-draw the grouped splits each time or only re-fit with fixed splits and fixed hyperparameters; the phrase “single-seed refits that reuse the tuned hyperparameters” is slightly ambiguous.
- Abstract: “nitrides, oxides, sulfides, and selenides” appears in the opening list while the title emphasizes oxides, sulfides, and selenides; confirm nitride count and whether nitride chemistry is retained in all reported metrics.
- Notation: define μB/fu and eV/atom consistently on first use in the main text; the abstract mixes “μBfu” and “eV/atom” styles.
- When the full manuscript is prepared, include a small table of dataset composition (counts by anion, by A/B cation family, metal vs non-metal) so readers can judge the 19-member non-metal holdout in context.
- Cite the specific Materials Project database version and the CrystalNN settings used for tetrahedral/octahedral assignment so the 320-entry curation is reproducible.
Circularity Check
No circularity: standard supervised ML on external Materials Project labels with group-aware holdouts; abstract-only review finds no self-definitional or fitted-as-prediction reductions.
full rationale
The paper reports tree-ensemble surrogates trained on 320 cubic spinel entries curated from the Materials Project (external DFT labels for formation energy, hull distance, band gap, magnetization, metallicity). Cation site assignments come from CrystalNN coordination numbers; evaluation uses reduced-formula grouped splits, transforms fit only on training folds, and champion selection on CV before holdout examination. Reported MAEs (0.121±0.030 eV/atom formation energy, 0.048±0.013 eV/atom hull, 1.27±0.19 μB/fu magnetization, metallicity accuracy 0.85±0.06) and the negative band-gap result versus a trivial baseline are empirical performance numbers under that protocol, not quantities derived by construction from fitted constants or from self-cited uniqueness theorems. SHAP attributions, conformal intervals, permutation nulls, and leave-one-chemistry-out tests are post-hoc diagnostics. No equations equate a claimed prediction to an input by definition; no ansatz is smuggled via self-citation; no uniqueness theorem is imported from the authors. The only available text is the abstract, which describes a conventional supervised-learning pipeline on external labels. Reader and skeptic concerns about possible leakage or unverifiable protocol details are correctness/reproducibility risks, not circularity. Score 0 is the honest finding for an abstract-only review of this type of work.
Assumptions & free parameters
free parameters (2)
- tree-ensemble hyperparameters (depth, learning rate, etc.)
- curated 320-entry spinel set membership
assumptions (3)
- domain assumption Materials Project semi-local DFT formation energies, hull distances, band gaps, and magnetizations are adequate labels for surrogate training and for the intended screening use.
- domain assumption CrystalNN coordination numbers correctly assign cations to tetrahedral-like vs octahedral-like groups for Fd-3m spinels.
- standard math Grouped splits by reduced formula plus train-only transforms yield unbiased estimates of generalization to new compositions.
Cite this review
Pith. "Pith review of Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides." pith.science (2026). https://pith.science/paper/S7Z634CN
@misc{pith2026260711996,
author = {Pith},
title = {Pith review of: Coordination-Resolved Surrogate Models for Thermodynamic Stability, Band Gaps, and Magnetic Moments of Spinel Oxides, Sulfides, and Selenides},
year = {2026},
howpublished = {\url{https://pith.science/paper/S7Z634CN}},
note = {Machine review of arXiv:2607.11996}
}
abstract
We curated 320 cubic ($Fd\bar{3}m$) spinel entries from the Materials Project-nitrides, oxides, sulfides, and selenides, including single-cation mixed-valence $A_3X_4$ compounds-and trained tree-ensemble surrogates for the formation energy, energy above the convex hull, band gap, total magnetization, and metallicity. Cations were assigned to tetrahedral-like and octahedral-like groups from CrystalNN coordination numbers rather than from element identity, and the evaluation was group-aware throughout: splits were grouped by reduced formula, every transform was fit on training folds only, and champion models were selected on cross-validated scores before the holdout was examined. Over twenty repeated grouped holdouts (single-seed refits that reuse the tuned hyperparameters, and are therefore mildly optimistic) the champions reach mean absolute errors of $0.121\pm0.030$ eV/atom for the formation energy, $0.048\pm0.013$ eV/atom for the hull distance, and $1.27\pm0.19$~\muBfu{} for the magnetization, with a metallicity accuracy of $0.85\pm0.06$. Band-gap regression does not beat a trivial baseline on the 19-member non-metal holdout under paired bootstrap testing, we report this negative result and trace it to sample scarcity and to the semi-local DFT labels. On the identical grouped split, the tabular champion is more accurate than an untuned MEGNet trained from scratch (0.087 versus 0.209 eV/atom formation-energy MAE on the primary holdout), a comparison that bounds, rather than settles, the descriptor-versus-graph question at this data scale. SHAP attribution ties the magnetization model to octahedral $d$-occupancy and the formation-energy and band-gap models to electronegativity descriptors, and grouped conformal intervals, permutation nulls, and leave-one-chemistry-out tests bound a domain of applicability that is uneven across anions and cations.
Reviewed July 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.