REVIEW 3 major objections 49 references
Machine learning that seems to read the nuclear equation of state from supernova gravitational waves is only memorising the simulation catalogue, not learning transferable physics.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 22:26 UTC pith:LY75UJX7
load-bearing objection Solid empirical audit: random-CV EoS success on Richers is template leakage; LOEO collapses below mean baselines across models, and the paper is careful not to overclaim impossibility. the 3 major comments →
The Generalization Gap in Machine Learning EoS Inference from Core-Collapse Supernova Gravitational Waves
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under random cross-validation a LightGBM regressor yields R² ≈ 0.70, 0.67, 0.60 for the nuclear incompressibility, symmetry energy and slope parameters (K₀, J, L). Under Leave-One-EoS-Out validation the identical model produces mean absolute errors of 44.57, 3.19 and 30.54 MeV and negative pooled R² scores—worse than simply predicting the catalogue mean. The gap persists across model families and after restriction to hand-crafted physical features, demonstrating that apparent success is catalogue interpolation driven by template leakage rather than a transferable waveform-to-EoS map.
What carries the argument
Leave-One-EoS-Out (LOEO) validation: every waveform generated from one equation-of-state family is withheld from training and used only for testing. This physically grouped split exposes the generalisation gap that random cross-validation conceals.
Load-bearing premise
The public two-dimensional CoCoNuT catalogue—with only eleven unique nuclear-parameter triplets and a short bounce window—is assumed to be a fair test of whether any transferable gravitational-wave map to the equation of state can exist.
What would settle it
Train the same regressors on a denser continuous grid of nuclear parameters (or on full three-dimensional neutrino-radiation-hydrodynamics waveforms) and re-run Leave-One-EoS-Out; positive R² and errors below the mean-predictor baseline on held-out families would falsify the claimed generalisation failure.
If this is right
- Published high accuracies for equation-of-state inference from core-collapse waveforms must be re-evaluated with leave-family-out protocols before they can be trusted for third-generation detectors.
- Future catalogues need wider and denser coverage of nuclear-parameter space; sparse discrete labels are insufficient for continuous regression.
- Physics-aware pipelines that condition on rotation and progenitor structure, or that use multi-messenger priors, become necessary once pure waveform-to-EoS lookup is shown to fail.
- Progenitor-mass classification can generalise across unseen rotation rates, so mass and equation-of-state tasks require different validation designs.
- Detector-noise stress tests that already erase random-CV success reinforce that fragile template leakage will not survive real data.
Where Pith is reading between the lines
- The same leave-family-out audit should be applied to other simulation codes and to long-term three-dimensional catalogues that include neutrino memory and viewing-angle dependence.
- Active-learning loops that preferentially simulate the equation-of-state families where LOEO residuals are largest could close the gap more efficiently than uniform catalogue expansion.
- If bounce and early post-bounce features remain degenerate even on denser grids, external astrophysical priors on rotation or progenitor mass will be indispensable for any GW-only dense-matter constraint.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper audits whether machine-learning models trained on finite CCSN GW catalogues can perform physical EoS inference for families absent from training. Using the Richers CoCoNuT catalogue (1,764 usable bounce-window waveforms; 21 EoS labels mapping to 11 unique (K0,J,L) triplets), the authors show that random 5-fold CV yields high R^{2} for LightGBM on raw time series (≈0.70, 0.67, 0.60 for K0, J, L), while Leave-One-EoS-Out (LOEO) validation produces MAE (44.57, 3.19, 30.54) MeV and negative pooled R^{2}, worse than mean-predictor baselines. The gap persists across Ridge, Random Forest, MLP, and LightGBM, and after restriction to physical features (bounce amplitude, width, peak frequency). Controls include interior vs boundary held-out EoS, Leave-One-Target-Triplet-Out, SHAP frequency importance, a discrete EoS classification leakage check, aLIGO SNR=100 noise stress tests, and a Mitra-catalogue progenitor-mass case study (classification generalises to unseen rotation; continuous LOMO regression compresses toward the catalogue interior). The authors conclude that catalogue interpolation is not robust out-of-catalogue EoS inference and recommend leave-family-out validation and physics-aware pipelines.
Significance. If the reported LOEO failure is representative of current public CCSN GW ML benchmarks, the paper is a timely and load-bearing methodological correction for a growing literature that has largely reported high random-CV accuracies. The multi-architecture comparison, mean baselines, physical-feature ablation, target-triplet superfamily control, and mass contrast study make the central claim falsifiable and useful for 3G detector pipeline design. Strengths include explicit use of public Zenodo catalogues, clear separation of template leakage from physical degeneracy, and appropriately cautious language that does not claim fundamental impossibility of GW-only EoS inference. The result is significant as a validation standard rather than as a new physical measurement.
major comments (3)
- §II and Fig. 1: the regression target space has only 11 unique (K0,J,L) triplets. While LOEO and Leave-One-Target-Triplet-Out both fail, the manuscript should quantify more clearly how much of the negative pooled R^{2} is driven by this discrete grid versus residual waveform-family non-transfer. A short leave-one-triplet residual decomposition (or a synthetic denser-grid thought experiment with the same feature correlations) would strengthen the claim that the failure is not solely sparsity.
- §III and Table I: LOEO MAE for raw LightGBM is worse than the global mean predictor for all three targets, yet several physical-feature models have LOEO MAE comparable to the mean. The paper should state more explicitly, for each target, whether any model+feature combination beats the LOEO training-mean baseline by a statistically meaningful margin, or whether the conclusion is uniformly “no better than mean.”
- §VII.B limitations: the audit is restricted to axisymmetric CFC CoCoNuT bounce windows. The central claim about “current waveform catalogues” is supported, but the discussion should more sharply separate (i) failure of random-CV protocols on existing public benchmarks from (ii) any implication for full 3D neutrino-radiation hydrodynamics catalogues. A one-paragraph statement of what would constitute a decisive positive LOEO test on a denser continuous EoS grid would help readers avoid over-reading the result.
Circularity Check
No significant circularity: empirical LOEO audit with held-out metrics, not a self-defining derivation.
full rationale
The paper is a generalisation audit of ML models on public CCSN GW catalogues. Its central claim is that random CV R^{2} ≈ (0.70, 0.67, 0.60) collapses under Leave-One-EoS-Out to MAE (44.57, 3.19, 30.54) MeV and negative pooled R^{2}, worse than a mean predictor, across model families and after physical-feature restriction. These quantities are computed on waveforms and EoS labels withheld from training; they are not fitted parameters renamed as predictions, nor are they forced by definition of the targets. Self-citations (Mitra et al. 2022/2023 catalogues, Abylkairov et al.) supply public data and prior ML setups; they are not load-bearing uniqueness theorems or ansatzes that force the LOEO failure. Controls (Leave-One-Target-Triplet-Out, interior vs boundary MAE, mean baselines, SHAP, mass classification contrast) are independent empirical checks. Catalogue sparsity and 2D CFC limits are stated as scope, not smuggled premises. No step reduces Eq. X to Eq. Y by construction. Score 0 is the correct honest finding.
Axiom & Free-Parameter Ledger
free parameters (2)
- LightGBM / RF / MLP hyper-parameters
- Bounce-window crop (-2 ms to +6 ms) =
-2 to +6 ms relative to bounce
axioms (4)
- domain assumption Public CoCoNuT axisymmetric CFC waveforms with parameterised deleptonisation constitute a valid benchmark for testing EoS generalisation of early-bounce GW features.
- domain assumption Leave-One-EoS-Out (and Leave-One-Target-Triplet-Out) is the appropriate grouped validation for claiming physical EoS inference rather than catalogue interpolation.
- standard math R^{2} < 0 and MAE worse than a training-mean predictor constitute failure of transferable physical mapping.
- domain assumption The three hand-extracted observables (bounce amplitude, FWHM width, early post-bounce peak frequency) are legitimate low-dimensional physical features for the leakage-control experiment.
read the original abstract
Core-collapse supernova gravitational waves may carry information about the dense matter equation of state (EoS), which describes the relation between pressure, density, temperature, and composition. This work tests a crucial question for physical inference: can a machine learning model trained on a finite simulation catalogue predict EoS parameters for an EoS family that was absent during training? Under standard random cross-validation, a LightGBM regressor appears highly successful, yielding $R^2=(0.70,0.67,0.60)$ for the nuclear incompressibility, symmetry energy, and slope parameter $(K_0,J,L)$. However, under Leave-One-EoS-Out (LOEO) validation, where all waveforms from a single EoS are withheld, the model fails, yielding mean absolute errors of $(44.57,3.19,30.54)$ MeV and negative pooled $R^2$ scores, performing worse than a baseline mean predictor. This generalisation gap persists across linear models, random forests, neural networks, and gradient-boosted trees. Restricting inputs to physical features (bounce amplitude, bounce width, peak frequency) reduces template leakage, the memorisation of related templates shared across random splits, but does not restore reliable EoS extrapolation. In contrast, a progenitor mass case study shows that classification generalises to unseen rotation speeds, while continuous mass regression compresses predictions towards the catalogue interior. These results demonstrate that while machine learning successfully interpolates within current waveform catalogues, this does not imply robust physical inference for unseen EoS models. Future pipelines should adopt leave-family-out validation, wider simulation coverage, and physics-aware inference frameworks.
Figures
Reference graph
Works this paper leans on
-
[1]
Standard random cross validation on raw waveforms yields misleadingly high performance due to template leakage. Future studies should report Leave One EoS Out or leave family 12 out validation, with MAE and mean predictor baselines, to verify physical generalisability
-
[2]
Restricting models to low dimensional physical observables (A bounce,w bounce,f peak) reduces apparent leakage and sometimes narrows the generalisation gap, but does not restore positive out of distribution performance. The feature target correlations remain weak, showing that the relevant observables are not uniquely diagnostic of EoS parameters in this ...
-
[3]
Progenitor mass can be classified robustly across unseen rotation speeds of the same progenitors (up to 76.12% accuracy). However, continuous mass inversion evaluated with Leave One Mass Out validation compresses endpoint masses towards the catalogue interior, confirming that catalogue interpolation is much easier than out of distribution inference. To br...
- [4]
- [5]
-
[6]
Richers et al., Zenodo data release, doi:10.5281/zenodo.201145
S. Richers et al., Zenodo data release, doi:10.5281/zenodo.201145
- [7]
- [8]
-
[9]
D. Vartanyan, A. Burrows, T. Wang, M. S. B. Coleman, and C. J. White, Phys. Rev. D107, 103015 (2023)
work page 2023
-
[10]
L. Choi, A. Burrows, and D. Vartanyan, Astrophys. J. 975, 12 (2024)
work page 2024
- [11]
-
[12]
A. Rusakov, A. S. Burrows, T. Wang, and D. Vartanyan, arXiv:2602.09025 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[13]
A. Burrows and collaborators, CCSN Gravitational Wave Signals data repository,https://www.astro.princeton. edu/~burrows/gw.3d.2024.update/
work page 2024
- [14]
-
[15]
B. P. Abbott et al., Class. Quantum Grav.34, 044001 (2017)
work page 2017
- [16]
-
[17]
V. Srivastava, S. Ballmer, D. A. Brown, C. Afle, A. Burrows, D. Radice, and D. Vartanyan, Phys. Rev. D 100, 043026 (2019)
work page 2019
-
[18]
Inferring physical properties of stellar collapse by third-generation gravitational-wave detectors
C. Afle and D. A. Brown, arXiv:2010.00719 (2020)
work page internal anchor Pith review Pith/arXiv arXiv 2010
- [19]
-
[20]
Mitra et al., Zenodo data release, doi:10.5281/zenodo.7090935
A. Mitra et al., Zenodo data release, doi:10.5281/zenodo.7090935
- [21]
-
[22]
Y. S. Abylkairov, M. C. Edwards, D. Orel, A. Mitra, B. Shukirgaliyev, and E. Abdikamalov, Mach. Learn.: Sci. Technol.5, 045077 (2024)
work page 2024
- [23]
-
[24]
Y. S. Abylkairov, M. C. Edwards, D. Orel, A. Mitra, B. Shukirgaliyev, and E. Abdikamalov, Phys. Rev. D112, 123056 (2025)
work page 2025
- [25]
-
[26]
A. Akhmetali, Y. S. Abylkairov, et al., arXiv:2603.27680 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[27]
A. Akhmetali, Y. S. Abylkairov, et al., arXiv:2605.04896 (2026)
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[28]
Leakage and the Reproducibility Crisis in ML-based Science
S. Kapoor and A. Narayanan, arXiv:2207.07048 (2022)
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[29]
I. E. Tampu, A. Eklund, and N. Haj Hosseini, Sci. Data 9, 580 (2022)
work page 2022
-
[30]
W. J. Engels, R. Frey, and C. D. Ott, arXiv:1406.1164 (2014)
work page internal anchor Pith review Pith/arXiv arXiv 2014
-
[31]
M. C. Edwards, arXiv:2009.07367 (2021)
work page internal anchor Pith review Pith/arXiv arXiv 2009
-
[32]
Y. S. Chao, C. Z. Su, T. Y. Chen, D. W. Wang, and K. C. Pan, arXiv:2209.10089 (2022)
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[33]
H. Dimmelmeier, J. A. Font, and E. M¨ uller, Astron. Astrophys.388, 917 (2002)
work page 2002
-
[34]
H. Dimmelmeier, J. Novak, J. A. Font, J. M. Ib´ a˜ nez, and E. M¨ uller, Phys. Rev. D71, 064023 (2005)
work page 2005
-
[35]
S. E. Woosley, A. Heger, and T. A. Weaver, Rev. Mod. Phys.74, 1015 (2002)
work page 2002
- [36]
-
[37]
S. E. Woosley and A. Heger, Phys. Rep.442, 269 (2007)
work page 2007
-
[38]
A. W. Steiner, M. Hempel, and T. Fischer, Astrophys. J. 774, 17 (2013)
work page 2013
- [39]
-
[40]
A. E. Hoerl and R. W. Kennard, Technometrics12, 55 (1970)
work page 1970
- [41]
-
[42]
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, Nature323, 533 (1986)
work page 1986
-
[43]
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Y. Liu, Adv. Neural Inf. Process. Syst.30 (2017)
work page 2017
-
[44]
S. M. Lundberg and S. I. Lee, Adv. Neural Inf. Process. Syst.30(2017)
work page 2017
-
[45]
C. Pastor Marcos et al., Phys. Rev. D109, 063028 13 (2024)
work page 2024
- [46]
-
[47]
M. J. Szczepa´ nczyk et al., Phys. Rev. D104, 102002 (2021)
work page 2021
-
[48]
S. E. Gossan et al., Phys. Rev. D93, 042002 (2016)
work page 2016
-
[49]
LIGO Scientific Collaboration, Advanced LIGO anticipated sensitivity curves, LIGO Document T0900288 v3 (2010)
work page 2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.