{"id":"0483daa3-8a53-415d-acc8-fbcfc46e269f","arxiv_id":"2412.19883","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Applying the GWAK autoencoder search to LIGO-Virgo O3 data recovers known compact binary mergers and glitches but finds no statistically significant unmodeled burst events.","lead":"This paper applies the GWAK neural network, trained on real LIGO and Virgo noise, to search for short gravitational-wave transients in the third observing run. It recovers known black-hole mergers and glitches but finds no statistically significant new burst events, and it reports sensitivity for generic waveforms.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported FARs and Table II sensitivities may be in-sample: the autoencoders are trained on O3 background data and then evaluated on the same O3 data (and timeslides built from it), with no stated train/search split.","rationale":"The reader's weakest assumption concerns the heuristic reweighting model being trained on a subset of the same timeslide pool used for FAR evaluation. That is a real and concrete leakage risk, and I agree it must be clarified. However, the more load-bearing and more general version of the same problem is that the autoencoders themselves are trained on O3 background data and then evaluated on the same O3 data and timeslides derived from it. Section III.B describes the training data as the O3 background after glitch excision and does not mention any train/test split; Section IV and Figure 2 analyze 'all O3 data.' Since the timeslides are built from exactly that data, every per-detector segment in the FAR pool is potentially in-sample for the autoencoders. This affects not only the heuristic reweighting but the core GWAK scores and the entire background model. The heuristic subset issue is thus a special case of a broader train/evaluation overlap. I keep the verdict CONDITIONAL (hence UNCHANGED relative to the reader) because the concern is addressable: the authors could demonstrate that training windows were excluded from the search, or they could rerun the analysis on a held-out portion. The null result itself may survive such a rerun, but the sensitivity numbers and the claim of enhanced search capability would not be trustworthy without this check. I do not see grounds for rejection: the detection of known CBCs provides independent evidence that the pipeline can find real signals, and the potential bias direction is not one that would manufacture the null result.","tokens_in":11208,"tokens_out":7778,"duration_ms":89752,"concrete_test":"Reproduce the analysis with strict data separation: train all autoencoders and the heuristic model only on a disjoint subset of O3 (e.g., the first half of the run), then compute FARs and Table II injection efficiencies on the remaining held-out half (or on timeslides built only from held-out data). Quantify the change in the 1/100-yr hrss thresholds and in the loudest candidate's FAR. If the thresholds shift by more than the quoted statistical uncertainties, in-sample training is confirmed and the reported sensitivity and FARs must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central null result and the sensitivity claims in Table II rest entirely on the false-alarm-rate estimates computed from ~10,000 years of timeslide data (Section III.C). Those timeslides are constructed by time-shifting the same H1/L1 O3 strain data that, according to Section III.B, was used to train the background autoencoder, the glitch autoencoder, and the three signal autoencoders. The paper never states that the O3 segments used for training were excluded from the segments scanned in the search or from the timeslide pool. If the autoencoders have seen the exact noise realizations they are later scored on, their reconstruction errors on those segments are artificially favorable, biasing the background score distribution and therefore the mapping from score to FAR. In addition, Section III.D says the heuristic reweighting model was trained on ~1000 years of timeslides 'randomly sampled' from the timeslides analyzed for the FAR calculation, with no statement that the training subset was held out of the 10,000-year evaluation pool. Even if the subset was held out, the heuristic's engineered features were designed after inspecting the 10,000-year background, so the evaluation pool informed the model selection. The likely effect is understated FARs: reported 1/100-yr thresholds are too lenient, making the Table II hrss sensitivities look better than they would be on truly unseen data. For the null conclusion, this bias is conservative in the sense that no candidate crossed even a potentially too-low FAR threshold, but the paper's broader claim of enhancing searches beyond existing pipelines is not supported if the background model is in-sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the GWAK semi-supervised neural-network framework to a search for unmodeled short-duration gravitational-wave transients in the H1 and L1 data from the third LIGO-Virgo-KAGRA observing run (O3). Five recurrent autoencoders are trained on background noise, glitches, and three injected signal classes (BBH, low-frequency sine-Gaussians, high-frequency sine-Gaussians); a linear classifier combines reconstruction losses and a frequency-domain correlation into a detection statistic, and a heuristic model reweights that statistic to suppress coincident and single-detector glitches. The search claims to recover three known CBCs and numerous glitches, to find no statistically significant unmodeled burst candidates after excluding known CBCs, and to achieve the sensitivity levels listed in Table II for several ad hoc burst morphologies.","tokens_in":11502,"tokens_out":3185,"duration_ms":37411,"significance":"If the methodology is sound, this is a useful demonstration that a semi-supervised machine-learning search can run on real GW detector data, recover known CBCs, and set burst-sensitivity benchmarks that are directly comparable with standard pipelines. The external verifiability of the recovered CBC events and the consistency of the cumulative FAR distribution with Poisson background are genuine strengths. The sensitivity numbers in Table II and the null result, however, inherit their validity from the false-alarm-rate estimation, and that estimation is exactly where the manuscript leaves a critical ambiguity. The paper therefore has the potential to be a valuable reference for ML-based unmodeled searches, but the central claims require an explicit and verifiable out-of-sample treatment of the training data.","major_comments":[{"comment":"The autoencoders are trained on O3 background data, and the search is then run on the same O3 data and on timeslides built from that same data, but the manuscript never states that the segments used for training were excluded from the analyzed 203.3 days or from the timeslide pool used for the FAR calculation. If the networks have seen the exact noise realizations they are later scored on, their reconstruction errors are in-sample and the background score distribution is artificially favorable; this biases the mapping from detection statistic to FAR, which is load-bearing for the null result in Fig. 2 and for the sensitivities in Table II. The authors should either state explicitly that disjoint train and search segments were used, or rerun the evaluation on segments and timeslides that were never used to train the autoencoders.","section":"Section III.B and III.C"},{"comment":"The heuristic reweighting model is trained on a randomly sampled subset of the timeslides analyzed for the FAR calculation, but the text does not say that this subset was held out from the 10,000-year evaluation pool. Even if the subset was held out, the engineered features and the model architecture were chosen after inspecting the full 10,000-year background, so the evaluation pool informed model selection. This is not merely a cosmetic issue: the final detection statistic is the product of the GWAK score and the heuristic score, so any in-sample tuning of the heuristic directly changes the reported FARs and the Table II hrss thresholds. The authors need to clarify the holdout status and, if the subset was not held out, repeat the FAR estimation with a properly separated training and evaluation set.","section":"Section III.D"},{"comment":"The cumulative FAR plot in Fig. 2 compares the observed events with the expected mean background and Poisson bands derived from the same timeslide pipeline. If the training/evaluation overlap described above exists, the background model itself is biased and the agreement shown in Fig. 2 is not an independent validation of the null result. A clean test would be to recompute the FARs using timeslides built only from data segments never used in training, or to show that the background distribution is stable when the training set is resampled. Without such a check, the quoted 1-in-100-year threshold and the conclusion that no candidate passes it are not yet established.","section":"Section IV and Fig. 2"}],"minor_comments":[{"comment":"The reader is told that the method improves on [43] by using 'real background data instead of simulated noise,' but the paper does not specify how much real background data was used for each autoencoder or how the training set was split by time; adding those numbers would make the in-sample risk much easier to assess.","section":"Section III.A / III.B"},{"comment":"The GPS range for O3b is printed as '125665561–1269363618'; the first number appears to be missing a digit (presumably 1256655618).","section":"Section III.C"},{"comment":"The text uses 'iFAR' without defining it, and in one place says 'iFAR≥ 100 years' where it presumably means an inverse false-alarm rate of 1 per 100 years; please define the quantity and correct the phrasing.","section":"Section IV"},{"comment":"The sensitivity entries are quoted without statistical uncertainties or the number of injections per morphology; adding these would make the comparison with other pipelines more meaningful.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"The central concern raised by the stress-test is real and is directly visible in the manuscript text: no train/search split is stated for the autoencoders, and the heuristic model is trained on a random subset of the very timeslides used for the FAR calculation without an explicit holdout. This is the kind of issue that can be fixed within the scope of the paper by clarifying the split and, if necessary, re-running the FAR estimation on held-out data. I do not see a fundamental flaw in the method itself, so I am recommending major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this paper takes the GWAK autoencoder search, previously validated on simulated noise, and runs it on real O3 data. It recovers three known CBCs, sees no other significant unmodeled bursts, and gives sensitivity estimates for several burst morphologies. That is honest progress for the ML-meets-LIGO crowd, and the null result is plausible.\n\nWhat is actually new: using real background data to train the autoencoders, a frequency-domain correlation statistic that replaces the expensive time-shifted Pearson term, and a heuristic reweighting model to suppress coincident glitches. The paper also includes CAT2 data that most pipelines veto, and the loudest non-CBC candidate is shown to be a glitch. Those are legitimate extensions of the prior GWAK paper, and the recovery of known events is a good sanity check.\n\nThe soft spot, and it is the one that matters, is the train/search overlap. The autoencoders are trained on O3 background and then evaluated on timeslides built from the same O3 data, with no stated split. The heuristic model is trained on a randomly sampled subset of the very timeslides used for the FAR calculation, and the text never says that subset was held out. That is an in-sample evaluation risk. If the models have seen the exact noise realizations, the background score distribution is artificially compact, the FARs are understated, and the sensitivities in Table II look better than they would on truly unseen data. The stress-test note is right about that. I would soften one part: the null conclusion is probably still safe, because the loudest non-CBC candidate has an inverse FAR of about a month, far from the 1-per-100-years threshold, so even a factor of a few bias would not flip it. But Table II is exactly where the bias bites.\n\nMinor issues: no code or model weights are released, and there is no direct comparison to cWB or BayesWave, so the abstract's claim of \"enhancing searches beyond the limits of existing pipelines\" is not actually demonstrated. That claim should be tempered unless they put the numbers side by side.\n\nThis paper deserves a serious referee. The central issue is fixable: a clean train/test split, or an explicit statement that the heuristic training timeslides were excluded, plus a sensitivity check on a held-out chunk of O3. If the numbers survive that, it's a solid foundation paper for the next observing run.\n\nSend it to review, with a request that the referee push on the split.","headline":"A useful first application of an autoencoder-based unmodeled search to O3 data, but the missing train/search split could bias the FAR and sensitivity numbers.","tokens_in":12138,"tokens_out":2995,"would_cite":false,"duration_ms":29930,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["04.80.Nn","07.05.Mh"],"model":"deepseek-v4-flash","headline":"A neural-network search of LIGO-Virgo-KAGRA O3 data finds no statistically significant unmodeled gravitational-wave bursts after excluding known compact binary coalescences.","keywords":["gravitational waves","unmodeled transients","neural networks","autoencoders","anomaly detection","LIGO-Virgo-KAGRA","O3 observing run","false alarm rates"],"falsifier":"Re-run the analysis with the heuristic model's training background strictly removed from the background pool used for the final false-alarm rates; if any non-CBC candidate then reaches an inverse false alarm rate of 100 years or more, the null conclusion would be overturned.","tokens_in":11008,"feed_emoji":"🔭","tokens_out":9851,"duration_ms":89238,"temperature":0.7,"pith_summary":"This paper asks whether a neural-network search that does not assume a specific signal template can find short gravitational-wave transients in the third LIGO-Virgo-KAGRA observing run that standard pipelines missed. Using the GWAK method, five recurrent autoencoders compress 50 ms detector segments into a low-dimensional space, and a classifier turns reconstruction loss and inter-detector coherence into a detection score. Run on 203.3 days of coincident Hanford-Livingston data, the search recovers three known compact binary coalescences and many detector glitches. After removing those known CBCs, it finds no statistically significant unmodeled burst, and it gives the strain sensitivity at which 50% of injected burst morphologies would be found at a false alarm rate of 1 per 100 years. The result matters because it shows a semi-supervised learning pipeline can reach a standard burst-search null conclusion while staying sensitive to generic waveforms.","feed_headline":"Neural-network search finds no new bursts in LIGO-Virgo O3 data","feed_subtitle":"Autoencoder GWAK recovers known black-hole mergers, maps burst sensitivity, and sees nothing significant beyond them.","key_machinery":"The carrier of the argument is the GWAK embedded space: five recurrent autoencoders are trained separately on background noise, glitches, binary-black-hole waveforms, and low- and high-frequency sine-Gaussian injections, and their reconstruction losses plus a frequency-domain correlation between the two detectors form the detection features. A linear classifier combines these features into the GWAK score. A small heuristic model, trained on about 1,000 years of time-shifted background events and injected signals, reweights that score using one-second context features and per-detector score asymmetries, suppressing coincident glitches. False alarm rates are then computed by running the reweighted statistic over about 10,000 years of time-shifted background data.","core_discovery":"The central claim is that a semi-supervised autoencoder search, GWAK, can be applied directly to real O3 data and both reproduce established detections and bound what remains. Five recurrent autoencoders are trained separately on background noise, glitches, binary-black-hole mergers, and low- and high-frequency sine-Gaussian injections; their reconstruction losses, combined with a frequency-domain correlation between the two detectors, define a detection metric. A small heuristic model reweights events using one-second context and per-detector score asymmetries to suppress coincident glitches. Analyzing 203.3 days of Hanford-Livingston data, GWAK detects the known CBCs, most prominently GW190828_063405, and the loudest non-CBC candidate has an inverse false alarm rate of about one month and is identified as a glitch. The paper's conclusion is that, once known CBCs are excluded, the O3 data contain no statistically significant unmodeled gravitational-wave bursts, with 50% detection efficiency at a 1-per-100-year false alarm rate reaching $h_{\\rm rss} \\approx 0.55 \\times 10^{-22}\\,\\mathrm{Hz}^{-1/2}$ for high-frequency sine-Gaussians.","pith_inferences":["The sensitivity values in Table II could be converted into population upper limits for burst sources such as cosmic-string cusps or supernovae during O3, but the paper stops at reporting efficiency.","Because the frequency-domain correlation replaces the costly time-sliding correlation, a low-latency version of the same search could plausibly run during O4, although real-time operation is not demonstrated here.","Retraining the five autoencoders on each future run's own background would let the method track non-stationary detector noise without changing the architecture, a step the paper leaves for later work."],"forward_implications":["GWAK recovers the known compact binary coalescences in O3, including GW190828_063405, so the method can serve as an independent check on template-based detections.","The search remains background-consistent even when vetoed CAT2 periods are included, meaning future analyses need not discard those data.","At a false alarm rate of one per hundred years, 50% detection efficiency is reached at $h_{\\rm rss} \\approx 0.55 \\times 10^{-22}\\,\\mathrm{Hz}^{-1/2}$ for high-frequency sine-Gaussians, giving a concrete benchmark for other burst pipelines.","After excluding known CBCs, no non-CBC event reaches the 1-in-100-year significance threshold, so the O3 data show no statistically significant unmodeled transient within GWAK's sensitivity."],"supporting_citations":[{"why":"introduces the GWAK semi-supervised autoencoder method that this analysis applies to real O3 data.","marker":"[43]"},{"why":"supplies the O3 all-sky short-burst search result and the 1-in-100-year inverse false alarm rate threshold the conclusion is evaluated against.","marker":"[14]"},{"why":"provides the catalog of known compact binary coalescences used as the check set and then excluded from the null result.","marker":"[5]"},{"why":"identifies the excess-power glitches that are split off into a dedicated training class.","marker":"[58]"},{"why":"is the public O3 strain data release from which the analyzed Hanford and Livingston segments are taken.","marker":"[51]"},{"why":"generates the binary black hole and sine-Gaussian injections with the stated priors used to train the signal autoencoders and to measure sensitivity.","marker":"[59]"}],"fun_headline_variants":["GWAK neural net finds no new GW bursts in O3","Autoencoder search: no new bursts beyond known CBCs","O3 unmodeled bursts: none found by neural network","Neural network sets limits on unmodeled O3 transients","GWAK recovers known mergers, finds no new bursts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported false-alarm rates assume the roughly 1,000 years of time-shifted background events used to train the glitch-reweighting model were not also counted in the roughly 10,000 years of time-shifted background used to compute the final rates, and the paper does not explicitly say they were excluded.","fun_headline_variants_meta":{"raw":{"variants":["GWAK neural net finds no new GW bursts in O3","Autoencoder search: no new bursts beyond known CBCs","O3 unmodeled bursts: none found by neural network","Neural network sets limits on unmodeled O3 transients","GWAK recovers known mergers, finds no new bursts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00024,"raw_usage":{"total_tokens":1529,"prompt_tokens":968,"completion_tokens":561,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":476}},"tokens_in":584,"tokens_out":561,"duration_ms":6053,"temperature":1.0,"reasoning_tokens":476,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:49:43.551116+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis with the heuristic model's training background strictly removed from the background pool used for the final false-alarm rates; if any non-CBC candidate then reaches an inverse false alarm rate of 100 years or more, the null conclusion would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"is the public O3 strain data release from which the analyzed Hanford and Livingston segments are taken."},{"cited_title":"Nitz et al., gwastro/pycbc: Pycbc release v1.16.9 (2020)","cited_arxiv_id":null,"evidence_quote":"generates the binary black hole and sine-Gaussian injections with the stated priors used to train the signal autoencoders and to measure sensitivity."}],"review_version":1}