{"id":"7ded6edb-970e-496a-9815-064e9271b07a","arxiv_id":"2411.12524","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A SynthGrav template bank recovers about 88% of simulated core-collapse supernova signals at 1 kpc in real detector noise, but with a high false-alarm rate.","lead":"This paper tests whether a bank of synthetic waveforms can catch gravitational waves from exploding stars in real LIGO-Virgo-KAGRA data, recovering about 88% of injected signals at 1 kiloparsec and about 50% at 2 kiloparsec. It matters because template-based methods could complement existing burst searches and help infer the properties of the newborn neutron star.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported detection efficiencies are quoted at a fixed SNR threshold with FAR ~130/day, not at a fixed false-alarm rate, so the claim of matching state-of-the-art excess-energy searches is unsupported.","rationale":"Good-faith reading: this is a feasibility study, not a deployed search; the authors openly discuss FAR, glitches, envelope dependence, and future improvements. They provide open-source software (SynthGrav), use real LVK O3b data, and clearly describe the pipeline. The strongest claim is the comparison to excess-energy searches, and that comparison requires matching FAR. The reader's CONDITIONAL verdict already captures this concern. I agree with the reader's rationale more than with the stated weakest_assumption: the linear-ramp assumption is a real limitation for generalizing beyond D25, but it is not the main threat to the reported numbers for D25, whose main component is approximately linear and for which the authors explicitly restrict their claims. The uncontrolled FAR is the more direct threat to the headline comparison with excess-energy searches. A single reanalysis with FAR thresholds would settle whether the claim survives. The paper's own caveats are honest, but the central comparison remains unsupported until efficiency is quoted at a fixed false-alarm rate.","tokens_in":16266,"tokens_out":5067,"duration_ms":53368,"concrete_test":"Using the same 3.41 days of O3b noise, measure the distribution of maximum SNR_n from the 150-template bank in the absence of injections, and set thresholds corresponding to FAR = 1/day, 1/month, and 1/100 yr. Recompute detection efficiency at 1, 2, and 5 kpc at each threshold, both with the current signal-derived envelope and with no envelope, and compare efficiency-vs-distance with Fig. 4 of Szczepańczyk et al. (2023) and Fig. 6 of Szczepańczyk et al. (2024) at the same FAR. If the 2-kpc efficiency at FAR = 1/100 yr falls below the excess-energy values (or below roughly 20%), the 'competitive' claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline quantitative claim is that the 150-template bank recovers 88% of D25 injections at 1 kpc and ~50% at 2 kpc (Fig. 6), and that this 'performs relatively well in comparison with state-of-the-art excess energy searches' (Sec. 4.1; Conclusions). These efficiency numbers are computed with a fixed network-SNR threshold of 6 (Sec. 2.2, Eq. 1). The paper itself measures the false-alarm rate at that threshold as ~130/day (Sec. 5, Fig. 8). State-of-the-art excess-energy searches with which the comparison is made (e.g., Szczepańczyk et al. 2023) report efficiencies at controlled FARs, typically 1 per 100 years. Because detection efficiency is a steep function of SNR threshold near the detection boundary, an efficiency quoted at FAR=130/day can be substantially higher than the efficiency at FAR=1/100yr. The paper acknowledges this mismatch but does not provide efficiency at any fixed FAR, so the central claim that the method is competitive with excess-energy searches is not actually established by the reported numbers. This is load-bearing because the abstract and conclusions present the recovery percentages and the comparison as the main result.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether matched filtering with a small theoretical template bank can detect gravitational waves from core-collapse supernovae. Using the 25-solar-mass Chimera model D25, the authors inject its waveform into roughly 3.4 days of O3b LIGO data and search with 150 synthetic templates generated by their new open-source package SynthGrav, whose central frequencies follow a linear time evolution f_c(t)=a t+b. The reported detection efficiencies are about 88% at 1 kpc and 50% at 2 kpc at a fixed network-SNR threshold of 6, with a false-alarm rate of about 130 per day; reconstructed slopes are claimed to be accurate to about 15%. The paper compares these numbers with excess-energy searches and concludes that matched filtering is competitive, while acknowledging limitations and suggesting future improvements.","tokens_in":16481,"tokens_out":7624,"duration_ms":73566,"significance":"The paper is a useful proof of concept: it demonstrates, to my knowledge for the first time, that a small SynthGrav-style template bank can be used with real LIGO data to recover a simulated core-collapse supernova signal, and the open-source release of SynthGrav is a concrete contribution. However, the headline quantitative claims are not yet supported as stated. The template amplitude envelope is taken from the injected signal, detection efficiencies are quoted at a threshold whose false-alarm rate is about 130 per day rather than at a fixed false-alarm rate, only one waveform and one observer orientation are injected, and the reconstruction error is measured against a visually estimated slope. These are load-bearing issues for the claimed comparison with excess-energy searches. The approach is promising and the issues are addressable with additional runs, so the paper merits major revision rather than rejection.","major_comments":[{"comment":"The amplitude envelope of each template is extracted from the injected signal via the Hilbert transform and a Savitzky-Golay filter, so the matched filter is partially constructed from the very signal it is asked to detect. The reported 88% and 50% detection efficiencies (Fig. 6, abstract) therefore are not blind-search efficiencies. The control run without the envelope reports only the mean, maximum, and minimum network SNR over 25 observers (7.0 vs. 6.3, etc.), not the detection efficiency at the SNR>6 threshold or at any fixed false-alarm rate; because detection efficiency is a steep function of SNR near threshold, this control does not establish that the headline efficiencies are unaffected. Please rerun the detection-efficiency calculation without the signal-derived envelope and report both sets of numbers.","section":"Sec. 2.3 / Sec. 4.1"},{"comment":"The detection efficiencies are quoted at a fixed network-SNR threshold of 6, which the paper itself measures to have a false-alarm rate of about 130 per day. The comparisons with Szczepańczyk et al. (2023) and (2024) are made at controlled false-alarm rates (e.g., 1 per 100 years in Szczepańczyk et al. 2023), so the statement that the proposed method outperforms the excess-energy search by almost a factor of 10, and the broader claim of competitive performance, are not supported by the reported numbers. Efficiency should be reported as a function of false-alarm rate, or the comparison should be restricted to the same false-alarm rate.","section":"Sec. 4.1 / Sec. 5 / Fig. 6 and Fig. 8"},{"comment":"Only a single simulated waveform (D25) and a single observer direction (phi, theta) = (35 deg, 0 deg) are used for the efficiency and reconstruction claims; no other model or code is injected, despite the abstract mentioning three models simulated with three different codes. The orientation sensitivity is large (network SNR varies from 2.7 to 12.5 over 25 observers at 1 kpc, Sec. 4.1), so the 88% and 50% numbers are not representative of an orientation-averaged search. Please inject a broader set of waveforms and average over observer orientations, or explicitly qualify all claims as applying to this one waveform and orientation.","section":"Sec. 2.1 / Sec. 4.1 / Fig. 6"},{"comment":"The claimed reconstruction accuracy of about 15% is measured relative to a true slope of approximately 2700 Hz/s that is estimated visually from the spectrogram in Fig. 2 (the dashed white line), rather than from a quantitative definition of the injected signal's instantaneous frequency. This informal reference makes the stated reconstruction error difficult to interpret. The bias should be quantified against a well-defined time-frequency measure of the injected waveform, such as a ridge estimate from the spectrogram or the known mode frequency evolution from the simulation.","section":"Sec. 4.2"}],"minor_comments":[{"comment":"The arXiv abstract differs from the full text: it says signals from three models simulated with three codes are considered, but only D25 is injected; please align the abstract with the analysis actually performed.","section":"Abstract / Sec. 2.1"},{"comment":"The network SNR is defined as the square root of the product of the single-detector SNRs, which is not the usual quadrature-sum network SNR; please justify this choice and state how the threshold of 6 maps to single-detector sensitivities.","section":"Eq. (1)"},{"comment":"The procedure used to count false alarms should be described explicitly, including how injection times are excluded and how the false-alarm rate is estimated from the analyzed stretch of data.","section":"Sec. 5"},{"comment":"Minor typographical errors appear, including 'a matched-filtering methods' in the abstract and 'complexcomplex conjugate' in Section 3.","section":"Abstract / Sec. 3"}],"recommendation":"major_revision","confidential_remarks":"The main risk is that the headline comparison to excess-energy searches is not yet supported, because the efficiency numbers use a signal-derived envelope and a threshold with a much higher false-alarm rate than the comparison papers. I would advise the editor that the revision must provide no-envelope efficiencies and false-alarm-rate-controlled efficiencies; if those are not provided, the comparison claims should be removed. The open-source code and the use of real LIGO data are strengths that make the paper worth a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2411.12524. It is the first generic template-bank matched-filtering study for the main stochastic CCSN GW component, and it comes with an open-source template generator, SynthGrav, that will be useful to the field. Second, the headline efficiencies—88% at 1 kpc and ~50% at 2 kpc—are computed at a fixed network-SNR threshold of 6, which the paper itself measures as a false-alarm rate of ~130/day. That is not the FAR (1 per 100 years) used by the excess-energy searches they compare against.\n\nWhat is genuinely good: the analysis uses real O3b LVK data, injects a realistic Chimera D25 signal, builds a 150-template bank of linear frequency ramps, and measures detection efficiency, reconstruction accuracy, and FAR. The reconstruction of the slope parameter to ~15% is a real result. The authors are transparent about the main cheat—the template amplitude envelope is extracted from the injected signal via Hilbert transform—and they show that dropping it reduces network SNR by roughly 10%. They do not, however, recompute detection efficiencies without the envelope, so the reported percentages are still partly optimistic. They also openly discuss glitches and the need for vetoing.\n\nThe softest spot is the comparison to excess-energy searches. They compare to Szczepańczyk et al. (2023), which reports efficiencies at FAR=1/100yr, while their own efficiencies sit at FAR ~130/day. Efficiency is steep near threshold, so the 'factor of 10' outperformance is not supported. The paper acknowledges this mismatch, and the multi-messenger argument—neutrino timing tolerates a high FAR—is legitimate. But that makes the claim a triggered-search claim, not a blind-search comparison. That distinction should be in the abstract. The envelope leakage is the second issue: the matched filter is partially built from the target, so the efficiencies are upper bounds. The control is partial, measuring SNR not efficiency. Third, only one injected model (D25) is used; the abstract in the arXiv listing mentions three models and three codes, which does not match the full text.\n\nThis paper deserves serious peer review. The feasibility case is real and the code and data are public. The revisions needed are major but tractable: quote efficiencies at a fixed FAR, measure efficiency without the envelope, and inject more than one signal. I would send it out, and I would expect the authors to be able to address these points.","headline":"A genuinely useful feasibility study with a public template generator, but the headline efficiencies are quoted at a false-alarm rate of ~130/day, so the 'competitive with excess-energy searches' claim is not yet established.","tokens_in":17025,"tokens_out":5313,"would_cite":true,"duration_ms":49725,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A template bank of 150 synthetic signals, generated by the open-source package SynthGrav, recovers 88% of injected core-collapse supernova gravitational-wave signals at 1 kpc and about 50% at 2 kpc in real LIGO-Virgo-KAGRA noise, with the…","keywords":["gravitational waves","supernovae: general","methods: data analysis","matched filtering","template bank","SynthGrav","core-collapse supernovae","LIGO-Virgo-KAGRA"],"falsifier":"Inject a numerical supernova waveform with a strongly curved or multi-branch frequency evolution (for example, a model with vigorous SASI activity, which produces narrowband emission near 100 Hz alongside the main component) into the same O3b LIGO noise at 1 kpc, and run the same 150-template linear-ramp bank with the same network-SNR threshold of 6; if the detection efficiency falls substantially below the 88% reported for D25, the single-linear-ramp template assumption is falsified.","tokens_in":16013,"feed_emoji":"💥","tokens_out":22074,"duration_ms":173727,"temperature":0.7,"pith_summary":"Matched filtering has been the standard way to detect gravitational waves from binary mergers, but core-collapse supernova signals are so stochastic that the field has relied on excess-energy searches that do not use waveform templates. This paper claims that template-based matched filtering is now feasible: a bank of 150 synthetic signals built by the new open-source package SynthGrav recovers 88% of simulated 25-solar-mass supernova signals injected at 1 kpc into real LIGO-Virgo-KAGRA noise, and about 50% at 2 kpc. For more than half of the recovered events, the signal's dominant frequency slope—which traces the proto-neutron star's compactness—is reconstructed to within about 15%. The detection efficiency is similar to that of current excess-energy searches, with the added benefit that the method returns physical parameter estimates. The main limitation is a false-alarm rate of about 130 per day, driven mostly by noise glitches, which the authors argue is acceptable when neutrino detectors provide timing information for a Galactic supernova.","feed_headline":"Matched filtering recovers 88% of supernova signals at 1 kpc","feed_subtitle":"A 150-template search matches excess-energy methods and recovers the supernova's frequency slope to ~15%.","key_machinery":"The pivotal object is SynthGrav, an open-source Python package that synthesizes supernova gravitational-wave signals as collections of modes, each built from overlapping pulses of colored random noise whose power spectral density is a Gaussian centred on a time-dependent frequency $f_c(t)$. For this analysis, SynthGrav is used to create 150 templates with $f_c(t) = a t + b$, sampling the slope $a$ in 50 steps from 250 to 3000 Hz/s and the intercept $b$ at 100, 200, and 300 Hz. The template amplitudes are shaped by an envelope extracted from the injected signal via the Hilbert transform and smoothed with a Savitzky–Golay filter, which improves the network signal-to-noise ratio by about 10% compared with using unshaped templates. The matched-filtering search is carried out with standard software: the data are whitened with the power spectral density of each 4096 s frame, the output signal-to-noise time series are clustered into 1 s bins, and the network SNR is defined as the geometric mean of the Livingston and Hanford SNRs, with a detection registered when this quantity exceeds 6. This machinery carries the argument because the template bank encodes the physical assumption that the dominant emission (the proto-neutron-star g-mode) sweeps its frequency linearly over the roughly 0.4 s signal, and the best-matching template's $(a,b)$ parameters are the reconstructed signal characteristics that the paper compares with the injection.","core_discovery":"The paper's central claim is that a small template bank of physically motivated synthetic signals can bring matched filtering to core-collapse supernova gravitational-wave detection, which was previously thought impractical because the signals are noisy and irregular. Using the D25 waveform—a 25-solar-mass, non-rotating solar-metallicity progenitor—injected into O3b data from the two LIGO detectors, the authors find that a 150-template bank with a network SNR threshold of 6 recovers 88% of injections at 1 kpc and about 50% at 2 kpc, with no detections at 5 kpc. The reconstructed templates cluster at a frequency slope of about 2250 Hz/s against a true value of roughly 2700 Hz/s, an underestimate of about 15%, and at a frequency offset of 100 Hz in most cases. The authors further show that if the signal itself were used as the template, it could be detected at 10 kpc under favorable orientations, which they interpret as the performance ceiling of the method. Their conclusion is that matched filtering with this template family performs comparably to excess-energy searches for Galactic distances and, unlike those searches, offers a route to measuring proto-neutron-star properties through the relation between the emitted frequency and the PNS mass and radius.","pith_inferences":["An implication not spelled out by the authors is that the reported detection efficiencies are probably optimistic for a real search, because the template amplitudes use an envelope extracted from the injected signal itself; a blind search would need to treat the amplitude evolution as unknown, which would enlarge the bank and likely lower the recovery rates.","The comparison with excess-energy searches is not apples-to-apples: the paper quotes a false-alarm rate of about 130 per day, while the cited excess-energy studies enforce false-alarm rates of roughly one per hundred years; if the matched filter were run at a comparable threshold, its detection efficiency would be lower than the raw numbers reported.","A natural next experiment, left for future work in the paper, is to inject a waveform with a curved (polynomial) frequency evolution—using the fitting formulas already implemented in SynthGrav—and rerun the linear-ramp bank; if the recovery efficiency at 1 kpc drops sharply, the linear-ramp family is the binding constraint.","Because most false triggers are associated with ~1 s noise glitches, a relatively simple time-frequency veto that checks whether the trigger follows the expected linear frequency ramp over the signal duration could reduce the false-alarm rate substantially without sacrificing sensitivity to genuine supernova signals; the authors mention glitch rejection but do not implement such a veto."],"forward_implications":["A Galactic supernova at 1 kpc would be detectable with ~88% efficiency by a 150-template matched-filter search in current LIGO detectors, and the recovered frequency slope would provide an estimate of the proto-neutron star's compactness.","Because the method returns parameter estimates rather than just a detection flag, a single nearby supernova could yield simultaneous information on the proto-neutron star's mass and radius when combined with neutrino measurements of the anti-electron neutrino energy.","The steep distance dependence (88% at 1 kpc, ~50% at 2 kpc, and none at 5 kpc) means that template-based matched filtering is a near-field technique limited to Galactic and very nearby extragalactic events, not an all-sky survey.","The high false-alarm rate (~130 per day) rules out standalone blind searches, but the probability of a false trigger coinciding with a neutrino signal is about $10^{-3}$, so the method is already viable as a confirmatory and parameter-estimation tool for neutrino-triggered Galactic supernova searches.","Because the current bank contains only linear frequency ramps, supernova signals with additional components (e.g., SASI or other PNS oscillation modes) will be recovered less efficiently; expanding the bank with physically motivated frequency evolutions, as the authors propose, should directly improve both detection and reconstruction."],"supporting_citations":[{"why":"Provides the D25 core-collapse supernova waveform that is injected into the detector noise.","marker":"Mezzacappa et al. 2023"},{"why":"Describes the three-dimensional simulation code used to produce the D25 model.","marker":"Bruenn et al. 2020"},{"why":"Identifies the proto-neutron-star g-mode as the main emission component whose linear frequency evolution the templates mimic.","marker":"Andresen et al. 2017"},{"why":"Supplies the fitting formulas for mode central frequencies that are implemented in SynthGrav.","marker":"Torres-Forné et al. 2019a"},{"why":"Derives the relation between gravitational-wave frequency and proto-neutron-star mass and radius that motivates the reconstruction accuracy.","marker":"Müller et al. 2013"},{"why":"Establishes the matched-filter signal-to-noise formalism used to rank templates.","marker":"Allen et al. 2012"},{"why":"Provides the O3b LIGO-Virgo-KAGRA open data into which the signals are injected.","marker":"Abbott et al. 2023b"},{"why":"Serves as the excess-energy search baseline whose detection efficiencies are compared with the matched-filter results.","marker":"Szczepańczyk et al. 2021"},{"why":"Reports detection efficiencies from an optically-triggered excess-energy search that the authors compare with their matched-filter results.","marker":"Szczepańczyk et al. 2024"}],"fun_headline_variants":["Matched filtering now practical for core-collapse supernova GWs","150 synthetic templates catch 90% of supernova GWs at 1 kpc","Matched filter recovers supernova GW frequency slope to 15%","Synthetic supernova templates enable matched filter at 1 kpc","Template bank turns supernova GW detection into a matched filter"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the main gravitational-wave emission of a core-collapse supernova has a single dominant frequency that rises linearly with time over the detector-relevant duration; if the real signal has a curved frequency evolution, multiple simultaneous emission components, or a very different amplitude envelope, the template bank will not match and the reported detection and reconstruction rates will not hold.","fun_headline_variants_meta":{"raw":{"variants":["Matched filtering now practical for core-collapse supernova GWs","150 synthetic templates catch 90% of supernova GWs at 1 kpc","Matched filter recovers supernova GW frequency slope to 15%","Synthetic supernova templates enable matched filter at 1 kpc","Template bank turns supernova GW detection into a matched filter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001038,"raw_usage":{"total_tokens":4453,"prompt_tokens":1116,"completion_tokens":3337,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":732,"completion_tokens_details":{"reasoning_tokens":3243}},"tokens_in":732,"tokens_out":3337,"duration_ms":22711,"temperature":1.0,"reasoning_tokens":3243,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:25:45.715068+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject a numerical supernova waveform with a strongly curved or multi-branch frequency evolution (for example, a model with vigorous SASI activity, which produces narrowband emission near 100 Hz alongside the main component) into the same O3b LIGO noise at 1 kpc, and run the same 150-template linear-ramp bank with the same network-SNR threshold of 6; if the detection efficiency falls substantially below the 88% reported for D25, the single-linear-ramp template assumption is falsified.","supporting_citations":[],"review_version":1}