Pith. sign in

REVIEW 3 major objections 49 references

Machine learning that seems to read the nuclear equation of state from supernova gravitational waves is only memorising the simulation catalogue, not learning transferable physics.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 22:26 UTC pith:LY75UJX7

load-bearing objection Solid empirical audit: random-CV EoS success on Richers is template leakage; LOEO collapses below mean baselines across models, and the paper is careful not to overclaim impossibility. the 3 major comments →

arxiv 2607.06736 v2 pith:LY75UJX7 submitted 2026-07-07 astro-ph.HE nucl-th

The Generalization Gap in Machine Learning EoS Inference from Core-Collapse Supernova Gravitational Waves

classification astro-ph.HE nucl-th
keywords core-collapse supernovaegravitational wavesequation of statemachine learninggeneralisation gapleave-one-out validationtemplate leakageproto-neutron star
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Core-collapse supernova gravitational waves are widely hoped to reveal the dense-matter equation of state—the relation between pressure, density and composition inside a newborn neutron star. Recent machine-learning studies report high accuracy when predicting nuclear parameters from simulated waveforms. This paper shows that those high scores appear only under ordinary random train–test splits. When every waveform belonging to one equation-of-state family is held out of training, the same models collapse to worse-than-mean performance. The failure holds for linear models, forests, neural nets and boosted trees, and survives even when the inputs are reduced to three simple physical numbers (bounce amplitude, width and peak frequency). In short, today’s catalogues let models interpolate among known templates; they do not yet support reliable physical inference for an equation of state the network has never seen. The practical consequence is that future pipelines must adopt leave-family-out validation and denser simulation grids before claiming they can measure nuclear physics from gravitational waves.

Core claim

Under random cross-validation a LightGBM regressor yields R² ≈ 0.70, 0.67, 0.60 for the nuclear incompressibility, symmetry energy and slope parameters (K₀, J, L). Under Leave-One-EoS-Out validation the identical model produces mean absolute errors of 44.57, 3.19 and 30.54 MeV and negative pooled R² scores—worse than simply predicting the catalogue mean. The gap persists across model families and after restriction to hand-crafted physical features, demonstrating that apparent success is catalogue interpolation driven by template leakage rather than a transferable waveform-to-EoS map.

What carries the argument

Leave-One-EoS-Out (LOEO) validation: every waveform generated from one equation-of-state family is withheld from training and used only for testing. This physically grouped split exposes the generalisation gap that random cross-validation conceals.

Load-bearing premise

The public two-dimensional CoCoNuT catalogue—with only eleven unique nuclear-parameter triplets and a short bounce window—is assumed to be a fair test of whether any transferable gravitational-wave map to the equation of state can exist.

What would settle it

Train the same regressors on a denser continuous grid of nuclear parameters (or on full three-dimensional neutrino-radiation-hydrodynamics waveforms) and re-run Leave-One-EoS-Out; positive R² and errors below the mean-predictor baseline on held-out families would falsify the claimed generalisation failure.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Published high accuracies for equation-of-state inference from core-collapse waveforms must be re-evaluated with leave-family-out protocols before they can be trusted for third-generation detectors.
  • Future catalogues need wider and denser coverage of nuclear-parameter space; sparse discrete labels are insufficient for continuous regression.
  • Physics-aware pipelines that condition on rotation and progenitor structure, or that use multi-messenger priors, become necessary once pure waveform-to-EoS lookup is shown to fail.
  • Progenitor-mass classification can generalise across unseen rotation rates, so mass and equation-of-state tasks require different validation designs.
  • Detector-noise stress tests that already erase random-CV success reinforce that fragile template leakage will not survive real data.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same leave-family-out audit should be applied to other simulation codes and to long-term three-dimensional catalogues that include neutrino memory and viewing-angle dependence.
  • Active-learning loops that preferentially simulate the equation-of-state families where LOEO residuals are largest could close the gap more efficiently than uniform catalogue expansion.
  • If bounce and early post-bounce features remain degenerate even on denser grids, external astrophysical priors on rotation or progenitor mass will be indispensable for any GW-only dense-matter constraint.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 0 minor

Summary. This paper audits whether machine-learning models trained on finite CCSN GW catalogues can perform physical EoS inference for families absent from training. Using the Richers CoCoNuT catalogue (1,764 usable bounce-window waveforms; 21 EoS labels mapping to 11 unique (K0,J,L) triplets), the authors show that random 5-fold CV yields high R^{2} for LightGBM on raw time series (≈0.70, 0.67, 0.60 for K0, J, L), while Leave-One-EoS-Out (LOEO) validation produces MAE (44.57, 3.19, 30.54) MeV and negative pooled R^{2}, worse than mean-predictor baselines. The gap persists across Ridge, Random Forest, MLP, and LightGBM, and after restriction to physical features (bounce amplitude, width, peak frequency). Controls include interior vs boundary held-out EoS, Leave-One-Target-Triplet-Out, SHAP frequency importance, a discrete EoS classification leakage check, aLIGO SNR=100 noise stress tests, and a Mitra-catalogue progenitor-mass case study (classification generalises to unseen rotation; continuous LOMO regression compresses toward the catalogue interior). The authors conclude that catalogue interpolation is not robust out-of-catalogue EoS inference and recommend leave-family-out validation and physics-aware pipelines.

Significance. If the reported LOEO failure is representative of current public CCSN GW ML benchmarks, the paper is a timely and load-bearing methodological correction for a growing literature that has largely reported high random-CV accuracies. The multi-architecture comparison, mean baselines, physical-feature ablation, target-triplet superfamily control, and mass contrast study make the central claim falsifiable and useful for 3G detector pipeline design. Strengths include explicit use of public Zenodo catalogues, clear separation of template leakage from physical degeneracy, and appropriately cautious language that does not claim fundamental impossibility of GW-only EoS inference. The result is significant as a validation standard rather than as a new physical measurement.

major comments (3)
  1. §II and Fig. 1: the regression target space has only 11 unique (K0,J,L) triplets. While LOEO and Leave-One-Target-Triplet-Out both fail, the manuscript should quantify more clearly how much of the negative pooled R^{2} is driven by this discrete grid versus residual waveform-family non-transfer. A short leave-one-triplet residual decomposition (or a synthetic denser-grid thought experiment with the same feature correlations) would strengthen the claim that the failure is not solely sparsity.
  2. §III and Table I: LOEO MAE for raw LightGBM is worse than the global mean predictor for all three targets, yet several physical-feature models have LOEO MAE comparable to the mean. The paper should state more explicitly, for each target, whether any model+feature combination beats the LOEO training-mean baseline by a statistically meaningful margin, or whether the conclusion is uniformly “no better than mean.”
  3. §VII.B limitations: the audit is restricted to axisymmetric CFC CoCoNuT bounce windows. The central claim about “current waveform catalogues” is supported, but the discussion should more sharply separate (i) failure of random-CV protocols on existing public benchmarks from (ii) any implication for full 3D neutrino-radiation hydrodynamics catalogues. A one-paragraph statement of what would constitute a decisive positive LOEO test on a denser continuous EoS grid would help readers avoid over-reading the result.

Circularity Check

0 steps flagged

No significant circularity: empirical LOEO audit with held-out metrics, not a self-defining derivation.

full rationale

The paper is a generalisation audit of ML models on public CCSN GW catalogues. Its central claim is that random CV R^{2} ≈ (0.70, 0.67, 0.60) collapses under Leave-One-EoS-Out to MAE (44.57, 3.19, 30.54) MeV and negative pooled R^{2}, worse than a mean predictor, across model families and after physical-feature restriction. These quantities are computed on waveforms and EoS labels withheld from training; they are not fitted parameters renamed as predictions, nor are they forced by definition of the targets. Self-citations (Mitra et al. 2022/2023 catalogues, Abylkairov et al.) supply public data and prior ML setups; they are not load-bearing uniqueness theorems or ansatzes that force the LOEO failure. Controls (Leave-One-Target-Triplet-Out, interior vs boundary MAE, mean baselines, SHAP, mass classification contrast) are independent empirical checks. Catalogue sparsity and 2D CFC limits are stated as scope, not smuggled premises. No step reduces Eq. X to Eq. Y by construction. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

The central claim is an empirical performance comparison on fixed public catalogues; it rests on standard ML evaluation axioms and domain assumptions about the simulation lineage rather than on free parameters or invented entities. The sparse target grid and CoCoNuT/CFC approximations are acknowledged limitations, not hidden free knobs that force the result.

free parameters (2)
  • LightGBM / RF / MLP hyper-parameters
    Default or lightly tuned settings are used; the paper does not claim optimality and the LOEO failure is robust across model families, so the claim does not hinge on a particular fit.
  • Bounce-window crop (-2 ms to +6 ms) = -2 to +6 ms relative to bounce
    Hand-chosen analysis window that defines the 524-sample input; different windows could change absolute scores but the random-vs-LOEO gap is the object of study.
axioms (4)
  • domain assumption Public CoCoNuT axisymmetric CFC waveforms with parameterised deleptonisation constitute a valid benchmark for testing EoS generalisation of early-bounce GW features.
    Stated in §II and limitations §VII.B; the audit is performed inside this catalogue family.
  • domain assumption Leave-One-EoS-Out (and Leave-One-Target-Triplet-Out) is the appropriate grouped validation for claiming physical EoS inference rather than catalogue interpolation.
    Core methodological premise of §§III–IV; standard group-wise CV logic applied to EoS labels.
  • standard math R^{2} < 0 and MAE worse than a training-mean predictor constitute failure of transferable physical mapping.
    Standard regression diagnostics used throughout §III and Table I.
  • domain assumption The three hand-extracted observables (bounce amplitude, FWHM width, early post-bounce peak frequency) are legitimate low-dimensional physical features for the leakage-control experiment.
    §IV; chosen for interpretability, not claimed to be a sufficient statistic.

pith-pipeline@v1.1.0-grok45 · 19325 in / 3138 out tokens · 94100 ms · 2026-07-10T22:26:32.211306+00:00 · methodology

0 comments
read the original abstract

Core-collapse supernova gravitational waves may carry information about the dense matter equation of state (EoS), which describes the relation between pressure, density, temperature, and composition. This work tests a crucial question for physical inference: can a machine learning model trained on a finite simulation catalogue predict EoS parameters for an EoS family that was absent during training? Under standard random cross-validation, a LightGBM regressor appears highly successful, yielding $R^2=(0.70,0.67,0.60)$ for the nuclear incompressibility, symmetry energy, and slope parameter $(K_0,J,L)$. However, under Leave-One-EoS-Out (LOEO) validation, where all waveforms from a single EoS are withheld, the model fails, yielding mean absolute errors of $(44.57,3.19,30.54)$ MeV and negative pooled $R^2$ scores, performing worse than a baseline mean predictor. This generalisation gap persists across linear models, random forests, neural networks, and gradient-boosted trees. Restricting inputs to physical features (bounce amplitude, bounce width, peak frequency) reduces template leakage, the memorisation of related templates shared across random splits, but does not restore reliable EoS extrapolation. In contrast, a progenitor mass case study shows that classification generalises to unseen rotation speeds, while continuous mass regression compresses predictions towards the catalogue interior. These results demonstrate that while machine learning successfully interpolates within current waveform catalogues, this does not imply robust physical inference for unseen EoS models. Future pipelines should adopt leave-family-out validation, wider simulation coverage, and physics-aware inference frameworks.

Figures

Figures reproduced from arXiv: 2607.06736 by Ayan Mitra.

Figure 1
Figure 1. Figure 1: FIG. 1. Sparse saturation parameter target space for the EoS [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3. Predictions outside the training sample vs. true values for [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Mean per EoS LOEO MAE for raw time domain [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. MAE comparison for raw time domain LightGBM [PITH_FULL_IMAGE:figures/full_fig_p006_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Spearman correlation matrix for extracted physical [PITH_FULL_IMAGE:figures/full_fig_p006_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. SHAP frequency importance for [PITH_FULL_IMAGE:figures/full_fig_p007_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: FIG. 8. LightGBM accuracy for progenitor mass classification [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: FIG. 9. Confusion matrices for progenitor mass classification under Random 5 Fold CV and GroupKFold by [PITH_FULL_IMAGE:figures/full_fig_p009_9.png] view at source ↗
Figure 11
Figure 11. Figure 11: The striking result is that noise suppresses even [PITH_FULL_IMAGE:figures/full_fig_p009_11.png] view at source ↗
Figure 11
Figure 11. Figure 11: FIG. 11. Time domain LightGBM [PITH_FULL_IMAGE:figures/full_fig_p010_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 49 canonical work pages · 8 internal anchors

  1. [1]

    Future studies should report Leave One EoS Out or leave family 12 out validation, with MAE and mean predictor baselines, to verify physical generalisability

    Standard random cross validation on raw waveforms yields misleadingly high performance due to template leakage. Future studies should report Leave One EoS Out or leave family 12 out validation, with MAE and mean predictor baselines, to verify physical generalisability

  2. [2]

    The feature target correlations remain weak, showing that the relevant observables are not uniquely diagnostic of EoS parameters in this catalogue

    Restricting models to low dimensional physical observables (A bounce,w bounce,f peak) reduces apparent leakage and sometimes narrows the generalisation gap, but does not restore positive out of distribution performance. The feature target correlations remain weak, showing that the relevant observables are not uniquely diagnostic of EoS parameters in this ...

  3. [3]

    Progenitor mass can be classified robustly across unseen rotation speeds of the same progenitors (up to 76.12% accuracy). However, continuous mass inversion evaluated with Leave One Mass Out validation compresses endpoint masses towards the catalogue interior, confirming that catalogue interpolation is much easier than out of distribution inference. To br...

  4. [4]

    Oertel, M

    M. Oertel, M. Hempel, T. Kl¨ ahn, and S. Typel, Rev. Mod. Phys.89, 015007 (2017)

  5. [5]

    Richers, H

    S. Richers, H. P. Pfeiffer, and C. D. Ott, Phys. Rev. D 95, 063019 (2017)

  6. [6]

    Richers et al., Zenodo data release, doi:10.5281/zenodo.201145

    S. Richers et al., Zenodo data release, doi:10.5281/zenodo.201145

  7. [7]

    Abdikamalov et al., Phys

    E. Abdikamalov et al., Phys. Rev. D90, 084007 (2014)

  8. [8]

    O’Connor and S

    E. O’Connor and S. Couch, Astrophys. J.865, 81 (2018)

  9. [9]

    Vartanyan, A

    D. Vartanyan, A. Burrows, T. Wang, M. S. B. Coleman, and C. J. White, Phys. Rev. D107, 103015 (2023)

  10. [10]

    L. Choi, A. Burrows, and D. Vartanyan, Astrophys. J. 975, 12 (2024)

  11. [11]

    Burrows, T

    A. Burrows, T. Wang, and D. Vartanyan, Astrophys. J. Lett.964, L16 (2024)

  12. [12]

    An Exploration of the Equation of State Dependence of Core-Collapse Supernova Explosion Outcomes and Signatures

    A. Rusakov, A. S. Burrows, T. Wang, and D. Vartanyan, arXiv:2602.09025 (2026)

  13. [13]

    Burrows and collaborators, CCSN Gravitational Wave Signals data repository,https://www.astro.princeton

    A. Burrows and collaborators, CCSN Gravitational Wave Signals data repository,https://www.astro.princeton. edu/~burrows/gw.3d.2024.update/

  14. [14]

    Punturo et al., Class

    M. Punturo et al., Class. Quantum Grav.27, 194002 (2010)

  15. [15]

    B. P. Abbott et al., Class. Quantum Grav.34, 044001 (2017)

  16. [16]

    Reitze et al., Bull

    D. Reitze et al., Bull. Am. Astron. Soc.51, 035 (2019)

  17. [17]

    Srivastava, S

    V. Srivastava, S. Ballmer, D. A. Brown, C. Afle, A. Burrows, D. Radice, and D. Vartanyan, Phys. Rev. D 100, 043026 (2019)

  18. [18]
  19. [19]

    Mitra, B

    A. Mitra, B. Shukirgaliyev, Y. S. Abylkairov, and E. Abdikamalov, Mon. Not. R. Astron. Soc.520, 2473 (2023)

  20. [20]

    Mitra et al., Zenodo data release, doi:10.5281/zenodo.7090935

    A. Mitra et al., Zenodo data release, doi:10.5281/zenodo.7090935

  21. [21]

    Mitra, D

    A. Mitra, D. Orel, Y. S. Abylkairov, B. Shukirgaliyev, and E. Abdikamalov, Mon. Not. R. Astron. Soc.529, 3582 (2024)

  22. [22]

    Y. S. Abylkairov, M. C. Edwards, D. Orel, A. Mitra, B. Shukirgaliyev, and E. Abdikamalov, Mach. Learn.: Sci. Technol.5, 045077 (2024)

  23. [23]

    Nunes et al., Phys

    A. Nunes et al., Phys. Rev. D110, 064037 (2024)

  24. [24]

    Y. S. Abylkairov, M. C. Edwards, D. Orel, A. Mitra, B. Shukirgaliyev, and E. Abdikamalov, Phys. Rev. D112, 123056 (2025)

  25. [25]

    A. B. Wang, Y. Yuan, H. Cai, and X. L. Fan, arXiv:2601.01376 (2026)

  26. [26]
  27. [27]
  28. [28]

    Leakage and the Reproducibility Crisis in ML-based Science

    S. Kapoor and A. Narayanan, arXiv:2207.07048 (2022)

  29. [29]

    I. E. Tampu, A. Eklund, and N. Haj Hosseini, Sci. Data 9, 580 (2022)

  30. [30]

    W. J. Engels, R. Frey, and C. D. Ott, arXiv:1406.1164 (2014)

  31. [31]

    M. C. Edwards, arXiv:2009.07367 (2021)

  32. [32]

    Y. S. Chao, C. Z. Su, T. Y. Chen, D. W. Wang, and K. C. Pan, arXiv:2209.10089 (2022)

  33. [33]

    Dimmelmeier, J

    H. Dimmelmeier, J. A. Font, and E. M¨ uller, Astron. Astrophys.388, 917 (2002)

  34. [34]

    Dimmelmeier, J

    H. Dimmelmeier, J. Novak, J. A. Font, J. M. Ib´ a˜ nez, and E. M¨ uller, Phys. Rev. D71, 064023 (2005)

  35. [35]

    S. E. Woosley, A. Heger, and T. A. Weaver, Rev. Mod. Phys.74, 1015 (2002)

  36. [36]

    Heger, S

    A. Heger, S. E. Woosley, and H. C. Spruit, Astrophys. J. 626, 350 (2005)

  37. [37]

    S. E. Woosley and A. Heger, Phys. Rep.442, 269 (2007)

  38. [38]

    A. W. Steiner, M. Hempel, and T. Fischer, Astrophys. J. 774, 17 (2013)

  39. [39]

    Liebend¨ orfer, Astrophys

    M. Liebend¨ orfer, Astrophys. J.633, 1042 (2005)

  40. [40]

    A. E. Hoerl and R. W. Kennard, Technometrics12, 55 (1970)

  41. [41]

    Breiman, Mach

    L. Breiman, Mach. Learn.45, 5 (2001)

  42. [42]

    D. E. Rumelhart, G. E. Hinton, and R. J. Williams, Nature323, 533 (1986)

  43. [43]

    G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Y. Liu, Adv. Neural Inf. Process. Syst.30 (2017)

  44. [44]

    S. M. Lundberg and S. I. Lee, Adv. Neural Inf. Process. Syst.30(2017)

  45. [45]

    Pastor Marcos et al., Phys

    C. Pastor Marcos et al., Phys. Rev. D109, 063028 13 (2024)

  46. [46]

    Sakan et al., arXiv:2511.08010 (2025)

    A. Sakan et al., arXiv:2511.08010 (2025)

  47. [47]

    M. J. Szczepa´ nczyk et al., Phys. Rev. D104, 102002 (2021)

  48. [48]

    S. E. Gossan et al., Phys. Rev. D93, 042002 (2016)

  49. [49]

    LIGO Scientific Collaboration, Advanced LIGO anticipated sensitivity curves, LIGO Document T0900288 v3 (2010)