{"id":"a214bb13-863a-412a-b54a-019a780a9f71","arxiv_id":"2607.21924","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A CNN can distinguish five supernova equations of state from reconstructed time-frequency images at 1 kpc in LVK O3b noise, but accuracy drops sharply by 5-10 kpc and the reported numbers are internally inconsistent.","lead":"This paper trains a convolutional neural network to tell five nuclear equations of state apart by reading the high-frequency gravitational-wave feature of simulated supernova signals buried in real LIGO-Virgo-KAGRA noise. At 1 kiloparsec the classifier is very accurate, but at 5 and 10 kiloparsec performance collapses, and the paper's own tables disagree about how fast it collapses.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single waveform per EOS means the 1 kpc accuracy may reflect waveform memorization, not EOS identification; the generalization claim is untested.","rationale":"The reader's weakest assumption is exactly the single-simulation-per-EOS design, and the load-bearing concern is that the high 1 kpc accuracy may be a memorization artifact. This is not an ad hominem or a disagreement with consensus; it is an internal validity issue about what the CNN actually learns. The paper provides useful evidence within its design: Monte Carlo cross-validation, cross-window testing, and per-class metrics. However, none of these tests vary the signal waveform, so they cannot rule out memorization. The concern is concrete and testable: train and test on independent simulations of the same EOS. If accuracy collapses, the abstract's 'successful EOS classification' and the scalability claim to 10 kpc would not be supported for real events. If accuracy remains high, the claim is substantially strengthened. The internal accuracy conflict (98.58% vs Tables 7–8) is secondary and does not change the verdict, but it reinforces the need for reproducible code and data. The reader's CONDITIONAL verdict remains appropriate; no change is needed.","tokens_in":26370,"tokens_out":3648,"duration_ms":36511,"concrete_test":"Train the same CNN architecture on cWB likelihood maps built from multiple CCSN simulations per EOS—varying progenitor mass or stochastic seed, e.g., Chimera models with different masses or 3D simulations with the same EOS—and test on a held-out simulation of each EOS that was never used in training, using the same TW1/TW2 noise windows and injection procedure. Report per-class accuracy at 1 kpc. If accuracy on the held-out simulation drops substantially below the 96–99% per-class values in Table 6, the classifier is memorizing the single waveform rather than identifying the EOS.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the CNN classifies EOS in real O3b noise at 1 kpc with 98.58% accuracy and that this could scale to 10 kpc with next-generation detectors—depends on the classifier learning EOS-dependent HFF properties rather than the specific simulated waveform. Section 4.1 injects one 2D Chimera simulation per EOS into real noise; every training and test image for a class is the same deterministic signal superimposed on different noise realizations. The TW1–TW2 cross-window test (Study 3) only changes the noise, not the signal waveform. A real CCSN will have a different stochastic realization, progenitor mass, rotation, magnetic field, 3D structure, and source orientation (the orientation limitation is acknowledged in Section 4.1, and the missing progenitor/rotation/magnetic-field variations are deferred to future work). With ~1000–3800 detections per class all derived from the same waveform, the CNN can memorize the unique time-frequency pattern of that waveform; the high per-class accuracies in Table 6 do not distinguish between 'recognizing this EOS' and 'recognizing this exact simulation.' The abstract's 98.58% accuracy also conflicts with the model accuracies of 0.90–0.95 in Tables 7 and 8, which is a secondary but real inconsistency that further weakens confidence in the headline numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper trains a single-stack convolutional neural network (CNN) to classify five nuclear equations of state (DD2, FSUgold, IUFSU, SFHo, SFHx) using cWB likelihood time-frequency maps of core-collapse supernova gravitational-wave signals injected into real O3b LVK noise. Signals from the Chimera E-series are placed in two one-week data windows (TW1, TW2) at Galactic distances of 1, 5, and 10 kpc, and the CNN is evaluated in three studies: train/test on TW1, train/test on TW2, and train on TW1/test on TW2. The manuscript reports high per-class accuracy at 1 kpc, degraded accuracy at 5 kpc, and near-loss of classification at 10 kpc, and argues that the 1 kpc performance suggests future detectors could classify EOS at about ten times the distance.","tokens_in":26628,"tokens_out":4562,"duration_ms":41073,"significance":"If the central claim holds, the paper would demonstrate that a CNN applied to cWB likelihood maps can separate EOS-dependent high-frequency-feature patterns in realistic interferometric noise, which is a useful proof-of-principle for CCSN parameter estimation. Strengths include the use of real O3b data, an explicit cross-time-window generalization test (Study 3), per-class and OvR metrics, and the SMOTE analysis for class imbalance. However, the significance is conditional on resolving internal inconsistencies in the reported metrics and on demonstrating that the classifier generalizes beyond the specific simulated waveforms used for training; without those, the headline claims about EOS identification in real events are not supported.","major_comments":[{"comment":"The headline accuracy numbers in the abstract do not match the paper's own tables. The abstract reports 98.58% accuracy at 1 kpc and 52.43% at 5 kpc, but Table 7 gives overall accuracy 0.93 at 1 kpc and 0.82 at 5 kpc (TW1, before SMOTE), and Table 8 gives 0.90 and 0.80 (TW2). In addition, Table 6 reports per-class accuracies at 1 kpc that are all above 96%, which is mathematically inconsistent with an overall accuracy of 90–93%; at 5 kpc the per-class values weighted by the class counts in Table 4 give roughly 49–55%, not 82%. The authors must reconcile these numbers and state exactly how the abstract's 98.58% and 52.43% were computed.","section":"Abstract and Sections 5–7 (Tables 6–9)"},{"comment":"The central generalization claim is not supported by the experimental design because each EOS class is represented by exactly one 2D Chimera simulation, injected only at equatorial orientation. Every training and test image for a given class is a noise realization of the same deterministic waveform, so the high 1 kpc accuracy may reflect memorization of that individual simulation rather than identification of EOS-dependent HFF properties. Study 3 changes only the noise window, not the signal waveform. A real CCSN will have a different stochastic realization, progenitor mass, rotation, magnetic field, and orientation, so the abstract's statement that the approach could scale to 10 kpc with next-generation detectors is an extrapolation that the present experiments cannot validate.","section":"Section 4.1 and Table 4"},{"comment":"The macro-averaged OvR AUC values are reported inconsistently. The abstract states macro-averaged OvR AUCs of 0.97 and 0.98 at 1 kpc, but Section 7 reports a macro-average AUC of 0.80 for TW1 and TW2, and Table 9 lists only per-class AUCs (0.96–0.98 at 1 kpc) with no macro-average row. The authors should clarify which number is the macro-average and provide the exact calculation, since the abstract and the text currently contradict each other.","section":"Section 7 and Table 9"},{"comment":"The paper acknowledges that only equatorial source orientation is considered, and it suggests that other orientations can be obtained by modifying the 1/r factor with a cosine of the orientation angle. This is not a substitute for evaluating the classifier at other inclinations, because the detectability of the HFF and the time-frequency morphology of the cWB reconstruction depend on the source orientation in a nontrivial way. A robustness test over inclination angles, or at least an explicit argument for why the equatorial result carries over, is needed before the claims about real Galactic CCSN events can be accepted.","section":"Section 4.1"}],"minor_comments":[{"comment":"The title contains a formatting issue: 'L VK' should be 'LVK'.","section":"Title"},{"comment":"The equation-of-state name is spelled inconsistently as both 'IUFSU' and 'IUSFU' (e.g., Table 2 vs. Section 4.1); the spelling should be unified.","section":"Throughout"},{"comment":"The text says 'we refer to Appendices A and A' but the intended cross-reference is unclear; there is only one Appendix A.","section":"Section 4.2 and Appendix A"},{"comment":"The relationship between the 'globally normalized' confusion matrices in Figure 3 and the 'per-class classification accuracy' in Table 6 should be defined precisely, since the two presentations can lead to different per-class measures.","section":"Figure 3 and Table 6"},{"comment":"The summary refers to 'Table 9' for OvR AUC values, but Table 9 reports per-class ROC AUCs; a sentence clarifying the distinction between per-class and averaged values would improve readability.","section":"Section 8"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reports internally inconsistent headline numbers: the abstract's 98.58% and 52.43% accuracies and the 0.97/0.98 macro AUCs do not match Tables 6–9. If these discrepancies are resolved and the claims are reframed as a proof-of-principle classification benchmark for five specific simulated waveforms, the paper may be suitable; however, the single-simulation-per-EOS design is a substantive limitation that should be stated prominently. I would ask the authors to recompute and report all metrics consistently and to either add generalization tests or soften the claims about applicability to real events."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a proof-of-concept for CNN-based EOS classification from cWB likelihood maps in real O3b noise, and the 1 kpc separation is plausible. But the headline numbers don't line up with the paper's own tables, and because each EOS is represented by a single deterministic waveform, the design does not support the 10 kpc scalability claim.\n\nWhat is actually useful: five Chimera E-series signals are injected into real O3b data, run through cWB, and classified with a CNN on likelihood maps. The cross-window transfer test (train TW1, test TW2) is good practice and shows the classifier is robust to noise non-stationarity. Monte Carlo cross-validation, per-class confusion matrices, OvR AUCs, and a SMOTE check are all there. At 1 kpc per-class accuracies are consistently above 96%, and the degradation at 5 and 10 kpc is monotonic and as expected. As a feasibility demonstration for a very nearby Galactic event, it works.\n\nThe soft spots are real. First, Section 4.1 states one 2D simulation per EOS, all injected at equatorial orientation. Every training and test example for a class is the same signal plus different noise. So the 1 kpc accuracy could reflect memorization of five specific waveforms rather than identification of the EOS. The TW1-to-TW2 test changes noise, not signal morphology. The paper acknowledges orientation and defers mass, rotation, and magnetic-field variation to future work, but the abstract's claim that this may scale to 10 kpc with next-generation detectors goes beyond the evidence. Second, the reported accuracies are internally inconsistent. The abstract gives 98.58% at 1 kpc and 52.43% at 5 kpc; Table 7 gives 0.93 and 0.82. The abstract gives macro-average OvR AUC 0.97/0.98; Section 7 and Table 7 report 0.80. Maybe those are different averaging methods or test splits, but the paper needs to reconcile them. Third, the novelty relative to [69] is not clear: the text says the CNN is 'as in [69]' while also implying previous work used only a DNN. The new accuracy and AUC numbers are new, but the marginal contribution should be stated explicitly.\n\nThis is for readers working on ML classification of CCSN gravitational-wave events. I would not cite the 10 kpc extrapolation, but the 1 kpc feasibility result could be useful once the numbers are reconciled. I would send it to peer review: a serious referee should ask for clearer reporting, code/data release, and tests with varied realizations and orientations before the generalization claim is accepted. Major revision, not desk reject.","headline":"Proof-of-concept CNN EOS classification at 1 kpc is plausible, but the internal accuracy inconsistencies and single-waveform-per-EOS design undercut the generalization claim.","tokens_in":27210,"tokens_out":4824,"would_cite":false,"duration_ms":41269,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CNN can identify the nuclear equation of state from a supernova's gravitational-wave signal at 1 kpc with 98.58% accuracy.","keywords":["core-collapse supernovae","gravitational waves","nuclear equation of state","convolutional neural networks","Coherent WaveBurst","LVK detector noise","high-frequency feature","HFF slope estimation"],"falsifier":"Train the same pipeline on multiple independent simulations per EOS, with different stochastic seeds, 3D structure, progenitor masses, and source orientations, and test on held-out realizations; if 1 kpc accuracy collapses toward the 20% chance level once waveform memorization is excluded, the central claim is refuted. A simpler check is to compare the 10 kpc confusion matrices against chance with a chi-square test, since the paper's own result predicts the diagonal should be statistically indistinguishable from random there.","tokens_in":26178,"feed_emoji":"🔭","tokens_out":9422,"duration_ms":75493,"temperature":0.7,"pith_summary":"This paper sets out to show that a convolutional neural network can read which nuclear equation of state (EOS) governed a core-collapse supernova directly from gravitational-wave data buried in real interferometric noise. The authors take five two-dimensional simulations of the same progenitor that differ only in their EOS (DD2, FSUgold, IUSFU, SFHo, SFHx), inject them into two one-week stretches of O3b detector data at 1, 5, and 10 kpc, and let the Coherent WaveBurst pipeline produce likelihood time-frequency images for a single-stack CNN to classify. At 1 kpc the classifier achieves 98.58% overall accuracy, with per-class scores above 96% and macro-averaged one-vs-rest AUC of 0.97-0.98; at 5 kpc it falls to 52.43% and at 10 kpc it is effectively blind. If the claim holds, it means the initial slope of the high-frequency feature in a supernova's gravitational-wave signal is a usable EOS fingerprint in realistic noise, and that the same approach could extend to about 10 kpc once next-generation detectors deliver their expected order-of-magnitude sensitivity gain.","feed_headline":"CNN names the nuclear equation of state from supernova signals","feed_subtitle":"Trained on real interferometric noise, the classifier names the right equation of state in 98.58% of tests.","key_machinery":"The central object is the cWB likelihood time-frequency map, an image $L_i \\in \\mathbb{R}^{N_{\\mathrm{time}}\\times N_{\\mathrm{freq}}\\times C}$ that grades each pixel by how coherently the two-detector network responds to a transient; once resized to $28\\times 28$ grayscale, the upward-trending high-frequency feature becomes a spatial pattern the CNN can see. The machinery is a single-stack convolutional network: convolution filters, LeakyReLU activations, max-pooling downsampling, a flattening layer, dense layers, and a softmax head that outputs probabilities over the five EOS classes. The physical load-bearing quantity is the HFF initial slope itself, whose noise-free values split the five models into a high-slope group (SFHo, SFHx) and a lower-slope group (DD2, FSUgold, IUSFU); the CNN's job is to recover that split from noisy likelihood maps.","core_discovery":"The central claim is that the initial slope of the high-frequency feature (HFF) in a core-collapse supernova's gravitational-wave signal, reconstructed by Coherent WaveBurst and interpreted by a convolutional neural network, is a practical discriminator among nuclear equations of state in real interferometric noise. Using five Chimera two-dimensional simulations that vary only the EOS, the paper reports per-class accuracy above 96% at 1 kpc in two independent one-week O3b time windows and in a cross-window transfer test, with macro-averaged one-vs-rest AUC of 0.97-0.98. At 5 kpc only the softer EOS models (SFHo and SFHx) remain reasonably identifiable, and at 10 kpc the confusion matrix diagonal falls to the level expected from chance, marking the distance limit of the method. The authors interpret this as evidence that the HFF slope signature survives realistic detector noise and temporal non-stationarity, and they argue that order-of-magnitude sensitivity improvements in next-generation observatories should shift the useful range from roughly 1 kpc to roughly 10 kpc.","pith_inferences":["Because each EOS is represented by a single 2D simulation, the 1 kpc accuracy could reflect waveform memorization rather than a general EOS rule; an obvious stress test is to train on many stochastic noise realizations of several independent simulations per EOS and see whether accuracy survives.","A natural extension the paper does not pursue is to replace the hard five-class label with a continuous regression of the HFF slope, which would turn each detection into a posterior over EOS-relevant physics and could be folded into multimessenger analyses.","The two-cluster structure in the HFF slopes suggests much of the 5 kpc discrimination is effectively soft-versus-stiff EOS classification; recasting the problem as binary or as a continuous compactness estimate might buy extra reach before the 10 kpc floor.","The same image pipeline could be tested on O4-era data or on injections with nonzero rotation and magnetic fields; if the mapping between HFF slope and EOS survives those perturbations, the method becomes a practical early-warning diagnostic for the next Galactic supernova."],"forward_implications":["At 1 kpc, a CNN trained on cWB likelihood maps can separate all five EOS classes with per-class accuracy above 96%, so a single Galactic supernova could plausibly constrain the nuclear EOS from gravitational waves alone.","At 5 kpc overall accuracy drops to 52.43% and at 10 kpc it becomes statistically indistinguishable from chance, setting the current method's reach at roughly the nearest few kiloparsecs.","A model trained on one one-week stretch of O3b data and tested on a later stretch keeps its accuracy, showing the learned EOS features are not tied to a specific noise realization.","With the order-of-magnitude sensitivity gain expected from next-generation detectors, the paper argues the 1 kpc classification performance would extend to roughly 10 kpc, bringing most of the Galaxy into range."],"supporting_citations":[{"why":"Supplies the neural-network method for estimating the HFF initial slope in LVK noise, the physical quantity the CNN is trained to distinguish.","marker":"[68]"},{"why":"Establishes how varying the EOS changes the HFF slope estimates in real interferometric data, providing the class separation this paper exploits.","marker":"[69]"},{"why":"Describes the Chimera code used to produce the five axisymmetric core-collapse supernova simulations that serve as signal templates.","marker":"[85]"},{"why":"Provides the Chimera E-series signals and their EOS-dependent HFF behavior that the paper injects into detector data.","marker":"[87]"},{"why":"Describes the coherent WaveBurst pipeline that generates the likelihood time-frequency maps used as CNN inputs.","marker":"[70]"},{"why":"Releases the O3b open-detector data from which the two one-week time windows (TW1 and TW2) are drawn.","marker":"[75]"}],"fun_headline_variants":["CNN identifies nuclear equation of state from supernova signals","CNN nails supernova's nuclear equation of state","CNN classifies nuclear EOS from supernova gravitational waves","Supernova GW signals reveal nuclear equation of state via CNN","Neural networks read nuclear EOS from supernova gravitational waves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Each EOS class is represented by just one simulated supernova, and every signal is injected at the same orientation relative to the detector, so the network may be matching a memorized waveform rather than learning a general EOS rule; real events with different turbulence, progenitor masses, rotation, magnetic fields, or viewing angles could break that rule.","fun_headline_variants_meta":{"raw":{"variants":["CNN identifies nuclear equation of state from supernova signals","CNN nails supernova's nuclear equation of state","CNN classifies nuclear EOS from supernova gravitational waves","Supernova GW signals reveal nuclear equation of state via CNN","Neural networks read nuclear EOS from supernova gravitational waves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000565,"raw_usage":{"total_tokens":2723,"prompt_tokens":1035,"completion_tokens":1688,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":651,"completion_tokens_details":{"reasoning_tokens":1609}},"tokens_in":651,"tokens_out":1688,"duration_ms":11931,"temperature":1.0,"reasoning_tokens":1609,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:30:20.281134+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same pipeline on multiple independent simulations per EOS, with different stochastic seeds, 3D structure, progenitor masses, and source orientations, and test on held-out realizations; if 1 kpc accuracy collapses toward the 20% chance level once waveform memorization is excluded, the central claim is refuted. A simpler check is to compare the 10 kpc confusion matrices against chance with a chi-square test, since the paper's own result predicts the diagonal should be statistically indistinguishable from random there.","supporting_citations":[{"cited_title":"Antelis, Claudia Moreno, Michele Zanolin, Anthony Mezzacappa, and Marek J","cited_arxiv_id":null,"evidence_quote":"Supplies the neural-network method for estimating the HFF initial slope in LVK noise, the physical quantity the CNN is trained to distinguish."},{"cited_title":"Bruenn, John M","cited_arxiv_id":null,"evidence_quote":"Describes the Chimera code used to produce the five axisymmetric core-collapse supernova simulations that serve as signal templates."},{"cited_title":"Landfield.Sensitivity of Neutrino-Driven Core-Collapse Supernova Models to the Microphysical Equation of State","cited_arxiv_id":null,"evidence_quote":"Provides the Chimera E-series signals and their EOS-dependent HFF behavior that the paper injects into detector data."},{"cited_title":"coherent waveburst, a pipeline for unmodeled gravitational-wave data analysis.SoftwareX, 14:100678, 2021","cited_arxiv_id":null,"evidence_quote":"Describes the coherent WaveBurst pipeline that generates the likelihood time-frequency maps used as CNN inputs."},{"cited_title":"Abbott, H","cited_arxiv_id":null,"evidence_quote":"Releases the O3b open-detector data from which the two one-week time windows (TW1 and TW2) are drawn."}],"review_version":2}