{"id":"a6ec11b5-da28-4d56-8a07-c6212013cd5b","arxiv_id":"2603.27680","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Using a linear SVM, EOS classification from bounce gravitational waves remains robust to real noise, progenitor diversity, and bounce-time uncertainty in the frequency domain, but collapses in the time domain under timing uncertainty.","lead":"This paper tests whether machine-learning classification of the nuclear equation of state from supernova gravitational waves survives more realistic conditions: real detector noise, multiple progenitor stars, and uncertain bounce timing. It finds the frequency-domain classifier stays accurate while the time-domain classifier fails when bounce time is uncertain.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is that axisymmetric CoCoNuT bounce signals with simplified neutrino treatment are faithful proxies for real CCSN GWs; if not, the robustness results do not transfer.","rationale":"The reader identified the same load-bearing assumption: the simulated waveforms with simplified microphysics and 2D axisymmetry may not represent the real EOS-relevant signal. I agree that this is the most fundamental concern. The paper's internal controls (balanced dataset, frequency-domain invariance, prompt-convection test) address internal robustness but do not establish that the training distribution matches nature. The reader's CONDITIONAL verdict is appropriate; no change is needed. Secondary issues — the swapped real/simulated noise values in Sec. III A versus Figure 3 caption, the abstract's overgeneralized 'none of these effects' given the time-domain failure, and the lack of SNR<50 tests — are real but are either reporting errors or scope limitations; they do not replace the primary concern about simulation fidelity. The concrete cross-code test would directly probe whether the EOS classification transfers across independent simulation frameworks.","tokens_in":18639,"tokens_out":7444,"duration_ms":85084,"concrete_test":"Cross-code transfer test: Train the SVM on the current CoCoNuT waveforms (four EOSs, all rotation configurations) and test on waveforms for the same four EOSs generated by an independent 3D general-relativistic CCSN code with spectral neutrino transport (e.g., CHIMERA or Fosite), using the same 30 ms windows, SNR=200, and Δt_b=20 ms. If accuracy drops from ~85% toward chance (25%), the classifier has learned code-specific features, and the external validity assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — that real noise, multi-progenitor training sets, and bounce-time uncertainty do not significantly degrade EOS classification — is established only within the physics of the CoCoNuT axisymmetric simulations. These use a Y_e(ρ) deleptonization scheme during collapse and a leakage/heating scheme after bounce, and the analysis truncates the signal at 6 ms post-bounce precisely because prompt convection is 'not accurately modeled' (Sec. II.A). The Appendix A test uses another 2D simulation set (Richers et al.), so it does not test whether 3D non-axisymmetric dynamics, turbulent convection, or more sophisticated neutrino transport would alter the EOS-dependent spectral features on which the classifier relies. If the real bounce-ringdown signal has a different spectral shape or different EOS sensitivity, the trained SVM may memorize simulation-specific features rather than physical EOS fingerprints; the reported 85.3% frequency-domain accuracy at SNR=200 with 20 ms timing uncertainty would then be an artifact of the training distribution. The paper's own conclusion acknowledges that the simulations 'do not fully capture physical processes that occur at later stages, such as convection and anisotropic neutrino emission,' which is a direct admission that the external validity condition is not met. This is the single most load-bearing concern because all other issues — the real/simulated noise labeling inconsistency, the time-domain collapse, and the untested SNR 10–30 regime — are internal or reporting caveats, not prerequisites for the claim to hold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends prior machine-learning studies of equation-of-state (EOS) classification from core-collapse supernova gravitational waves by relaxing three simplifying assumptions: simulated Gaussian noise is replaced by real LIGO O4a detector noise; a single progenitor is expanded to four ZAMS-mass models (12–40 M_sun) with multiple rotational configurations; and core-bounce time uncertainty up to 20 ms is introduced. Using axisymmetric CoCoNuT waveforms truncated to an 8 ms window around bounce, the authors inject signals into noise at fixed SNR, train a linear-kernel SVM on time- and frequency-domain representations, and report accuracy averaged over 50 random train/test splits. They conclude that real noise, progenitor diversity, and bounce-time uncertainty do not significantly degrade classification performance, and that the larger multi-progenitor dataset actually improves accuracy. A control with a balanced dataset shows that the improvement is driven by dataset size rather than progenitor diversity. The paper also includes an appendix test with prompt-convection-containing waveforms from Richers et al., finding that the classifier does not rely primarily on the stochastic convection component.","tokens_in":126,"tokens_out":7832,"duration_ms":97875,"significance":"If the results hold, the paper makes a useful incremental contribution to the CCSN GW inference program: it demonstrates within a specific simulation framework that EOS classification remains robust when the training data are expanded to multiple progenitors and real detector noise, and it clearly identifies the frequency-domain representation as necessary when bounce-time uncertainty is present. The strengths are the use of publicly available O4a LIGO noise, the 50-split evaluation, the balanced-dataset control, and the explicit overfitting test regarding prompt convection. These elements make the main empirical claims reproducible and internally well controlled. The main weakness is that the robustness is established only for the axisymmetric, simplified-neutrino CoCoNuT waveform family; the paper's own limitations section acknowledges that later-stage convection and anisotropic neutrino emission are not captured. Therefore the significance is real but should be framed as 'robust within the simulated waveform family,' not as unconditional robustness to real CCSN signals.","major_comments":[{"comment":"The accuracy values are assigned inconsistently between text and figure. The text says at SNR=200, time-domain accuracies are 92.3%±2.9% for real noise and 87.2%±2.9% for simulated noise; the figure caption instead says the time-series classifier achieves 92.3%±2.9% for simulated noise and 91.6%±3.1% for real noise, while the frequency-domain classifier achieves 87.2%±2.9% (simulated) and 85.3%±2.9% (real). Table I and §III.C confirm that with all four progenitors and Δt_b=0 the real-noise results are time=91.6±3.1 and frequency=85.3±2.9. Thus the text labels in §III.A are swapped. Please correct the text and make caption, text, and tables mutually consistent; the conclusion of similar performance may still hold, but as written the reported numbers cannot be trusted.","section":"§III.A and Fig. 3 caption"},{"comment":"The abstract states that 'none of these effects significantly degrades EOS classification performance.' This is contradicted by the paper's own time-domain result: accuracy collapses from 91.6%±3.1% at Δt_b=0 to 39.8% at Δt_b=10 ms and 34.3% at Δt_b=20 ms (Fig. 4). The paper later describes the time-domain classifier as 'struggles.' The claim is only valid for the frequency-domain representation (or for the recommended pipeline using frequency-domain features). Please qualify the abstract and conclusion accordingly; as written, the central claim is overstated.","section":"Abstract and §III.C"},{"comment":"The elevenfold augmentation by applying bounce-time shifts in 2 ms steps must be partitioned in a way that prevents data leakage. If the 80/20 train/test split is performed after augmentation, different shifted copies of the same physical waveform can appear in both training and test sets, inflating the reported 91.6%±2.7% accuracy. Please state explicitly that the split is performed on the original 886 waveforms before augmentation (or use a group-split strategy), and clarify whether the green curve in Fig. 4 is trained only on augmented training data or on a mixture of augmented and unaugmented samples.","section":"§III.C, larger-dataset experiment"},{"comment":"The robustness conclusion is conditional on the axisymmetric CoCoNuT simulation family with a simplified Y_e(ρ) deleptonization scheme, leakage/heating after bounce, and a 6 ms post-bounce cutoff because prompt convection is 'not accurately modeled.' The Appendix A test with Richers et al. uses another 2D dataset and does not test 3D non-axisymmetric dynamics, turbulent convection, or different neutrino transport schemes. Thus the results demonstrate robustness within the simulated waveform family, not transferability to real CCSN signals. Please state this limitation explicitly in the abstract and conclusion, and avoid phrasing that implies the effects are negligible for actual observations.","section":"§II.A, §IV, and Appendix A"}],"minor_comments":[{"comment":"Typos: 'asses' should be 'assess' (§III.A), 'constrainghts' should be 'constraints' (§I), 'intoroduction' in Ref. [100], 'classifies' should be 'classifier' in Appendix A, and 'a injected signal' should be 'an injected signal' in the Fig. 2 caption.","section":"General"},{"comment":"The 'frequency domain counterparts' is not defined. Please specify whether the input is the magnitude spectrum, power spectrum, or complex FFT of the 30 ms window. This affects reproducibility.","section":"§II.C"},{"comment":"Please clarify that the 1024 s O4a segment is a single contiguous noise realization and that the random placement of the signal within 30 ms windows generates different noise samples but from the same segment. Also fix 'three value' to 'three values'.","section":"§II.B"},{"comment":"The description of the balanced dataset is ambiguous: 'subsample signals in each class so that the total number of training examples is equal across all scenarios' could mean per EOS class or per configuration. Specify whether the balancing is applied per EOS class, per progenitor configuration, or to the total training set, and for which SNR the 220-signal count refers.","section":"§III.B"},{"comment":"Add a caption sentence describing the green 'larger dataset' curve; currently it is only explained in the text. Also state the number of training samples after elevenfold augmentation and confirm that the test set remains the same as for the other curves.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a natural extension of the authors' previous work and leans heavily on their own earlier papers (Refs. 75–77, 85) for the simulation pipeline and ML setup. The genuinely new elements are the use of real O4a noise, the multi-progenitor coverage, and the bounce-time uncertainty study. These are sufficient for incremental publication in a specialized journal, but the internal number inconsistencies and the abstract overclaim need to be fixed before acceptance. The paper does not contain machine-checked proofs or novel algorithms; its value is as an empirical robustness study. I would advise the editor to request a revision that addresses the label swap and the leakage ambiguity in the augmented dataset experiment."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid, incremental step, not a breakthrough. It does what it says: relaxes three assumptions from the authors' earlier ML EOS classification work and shows that a frequency-domain SVM stays accurate under real O4a detector noise, four progenitor models, and up to 20 ms bounce-time uncertainty at high SNR. The time-domain classifier collapsing to roughly 34-40% under the same timing uncertainty is a genuinely non-obvious result worth knowing. The methodology is careful. Fixed-SNR injection into real noise, 50 random train/test splits, and a balanced-dataset control that separates dataset-size effects from progenitor-diversity effects are handled properly. The appendix test with a different simulation set (Richers et al.) is a good-faith attempt to address prompt convection, and the overfitting check restricting test bounce times to 18-20 ms is persuasive. The paper also spells out its limitations rather than burying them. The soft spots are real but addressable. First, there is a flat inconsistency between the text in Section III.A and the Figure 3 caption: at SNR=200 the text assigns 92.3% time-domain real noise and 87.2% simulated, while the caption swaps those assignments and gives different numbers for frequency domain. One of these is wrong. Second, the abstract says none of the three effects significantly degrades classification performance, but that is only true for the frequency-domain representation; the time-domain classifier loses more than half its accuracy at 20 ms bounce-time uncertainty. The abstract should say the collapse is limited to time-domain inputs. Third, the SNR range plotted starts at 50, while the real detection regime for current interferometers is SNR 10-30. The paper mentions this but does not test it, a meaningful gap. The deeper concern is external validity: the waveforms come from axisymmetric CoCoNuT simulations with simplified deleptonization and leakage, and the analysis truncates at 6 ms post-bounce because prompt convection is not accurately modeled. Appendix A partially mitigates this, but the test data are also 2D. So robustness claims are established within the simulation family, not guaranteed for real 3D signals. The paper acknowledges this; I do not think it is fatal, but it is another reason to temper the abstract. This paper deserves a serious referee. The methodology is sound, the new tests are useful, and the results are clearly presented apart from the issues above. I would cite it if I worked in this area, and I would bring it to a reading group focused on ML for multimessenger astrophysics. Recommended action: send to peer review, but require fixing the text/caption mismatch and softening the abstract's sweeping claim.","headline":"Useful incremental robustness study; the real-vs-simulated noise comparison has a reporting inconsistency, and the abstract overstates the bounce-time result for time-domain classifiers.","tokens_in":737,"tokens_out":1833,"would_cite":true,"duration_ms":37073,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Supernova equation-of-state fingerprints survive noise and timing jitter","keywords":["supernova gravitational waves","equation of state","machine learning","support vector machine","frequency-domain classification","core bounce","detector noise","progenitor models"],"falsifier":"Train the same frequency-domain classifier on high-resolution, three-dimensional core-collapse simulations that include self-consistent neutrino transport and realistic prompt convection (or, ultimately, on the first coincident neutrino-gravitational-wave detection of a Galactic supernova), and measure whether classification accuracy stays above roughly 80%. If the accuracy drops materially on these more realistic signals, the paper's robustness claim would be falsified for real detections.","tokens_in":18597,"feed_emoji":"🔭","tokens_out":5271,"duration_ms":46950,"temperature":0.7,"pith_summary":"The paper asks whether machine-learning classification of the nuclear equation of state from supernova gravitational waves still works when the idealized assumptions of earlier studies are relaxed. The authors inject simulated signals into real detector noise, expand the training set to four progenitor masses with a range of rotation rates, and allow the core-bounce time to be uncertain by up to 20 milliseconds. They report that none of these complications significantly hurts accuracy: a frequency-domain classifier keeps roughly 85% accuracy at high signal-to-noise ratio, while the time-domain classifier degrades sharply when bounce time is uncertain. The central positive finding is that the equation-of-state fingerprint lives in the spectral content of the bounce and early ring-down signal, and that larger, more diverse training sets help rather than hurt. This matters because a future Galactic supernova will arrive with noise, unknown progenitor structure, and imprecise timing, so a usable classification method must tolerate all three.","feed_headline":"Equation-of-state fingerprint survives supernova noise and timing jitter","feed_subtitle":"Frequency-domain classifier holds ~85% accuracy with 20 ms bounce-time uncertainty; time-domain drops to ~34%.","key_machinery":"The load-bearing mechanism is the frequency-domain (Fourier-amplitude) representation of the short gravitational-wave segment spanning −2 to +6 ms around core bounce. Because this representation discards absolute phase and time alignment, it is nearly invariant to shifts in the bounce time, which is why it remains accurate when the bounce time is unknown to within 20 ms. The classifier itself is a linear-kernel support vector machine (a supervised learning algorithm that finds a separating hyperplane); its simplicity lets the performance differences be attributed to the input representation and the training data. The training set is expanded by injecting each signal into real detector noise","core_discovery":"The paper's central claim is that the EOS of dense nuclear matter can be classified from the gravitational-wave burst of a rapidly rotating core-collapse supernova even under three realistic complications: real detector noise, a diverse set of progenitor models spanning 12 to 40 solar masses, and uncertainty in the core-bounce time of up to 20 ms. The authors find that a linear support-vector-machine classifier acting on the frequency-domain waveform maintains 85.3%±2.9% accuracy at signal-to-noise ratio 200 when bounce time is uncertain to 20 ms, and improves to 91.6%±2.7% when the training set is enlarged elevenfold. The time-domain classifier, by contrast, falls to about 34–40% under the","pith_inferences":["If the frequency-domain fingerprint is truly time-shift invariant, then other shift-invariant representations that preserve more phase information than the magnitude spectrum, such as the bispectrum or wavelet scalogram moduli, might achieve even higher accuracy than the plain Fourier amplitude; this is untested in the paper.","The success of training on multiple progenitors despite their structural differences suggests that the EOS signature is a common shared feature across different masses; this hints that a universal EOS-feature subspace could be learned and transferred to unseen progenitor models, but the paper does not demonstrate generalization to progenitors outside the 12–40 solar-mass range.","The authors' own caveat that their axisymmetric simulations do not model prompt convection faithfully means the extrapolation to real signals rests on the Appendix test using an approximate convection injection; a decisive test would require 3D simulations with self-consistent turbulence, and the paper leaves that to future work."],"forward_implications":["A future Galactic supernova, observed at high SNR by next-generation detectors, could have its nuclear EOS classified using only the bounce and early ring-down signal, without precise knowledge of the bounce time.","Frequency-domain features are the appropriate input for EOS classification; time-domain classifiers should be avoided unless the bounce time is known to well under 10 ms from neutrino timing.","Training-data volume is the main lever: adding progenitors, rotation rates, and time-shift augmentations improves accuracy more than any algorithmic refinement.","The saturation of accuracy above SNR≈100 implies that for third-generation detectors, the limiting factor is the coverage of the waveform model space, not detector sensitivity."],"fun_headline_variants":["Supernova waves still fingerprint dense matter with bounce jitter","EOS readout holds under real noise and bounce-time uncertainty","Dense-matter EOS classification robust to supernova noise and timing","Bounce-time jitter and noise don't break supernova EOS probes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire classification scheme assumes that the gravitational-wave signals produced by axisymmetric core-collapse simulations that use simplified neutrino and convection physics—and that are truncated before prompt convection develops—faithfully represent the observable bounce signal of a real Galactic supernova.","fun_headline_variants_meta":{"raw":{"variants":["Supernova waves still fingerprint dense matter with bounce jitter","EOS readout holds under real noise and bounce-time uncertainty","Dense-matter EOS classification robust to supernova noise and timing","Bounce-time jitter and noise don't break supernova EOS probes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1293,"prompt_tokens":725,"completion_tokens":568,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":469,"tokens_out":568,"duration_ms":6495,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T05:37:46.606067+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same frequency-domain classifier on high-resolution, three-dimensional core-collapse simulations that include self-consistent neutrino transport and realistic prompt convection (or, ultimately, on the first coincident neutrino-gravitational-wave detection of a Galactic supernova), and measure whether classification accuracy stays above roughly 80%. If the accuracy drops materially on these more realistic signals, the paper's robustness claim would be falsified for real detections.","supporting_citations":[],"review_version":2}