{"id":"7b800cbd-ba4a-4234-8c0b-cefcfc4c2ce2","arxiv_id":"2509.10505","paper_version":7,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An unsupervised CWT-LSTM autoencoder trained only on LIGO O4 noise reaches 97.0% precision and 96.1% recall on 102 confirmed gravitational wave events, and the study shows cross-run calibration differences distort multi-run training.","lead":"A noise-only neural network trained on LIGO O4 data flags 98 of 102 confirmed gravitational wave events with 97% precision and 96% recall. The paper also documents a calibration-related batch effect between observing runs and shows that single-run training removes it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported 97.0%/96.1% operating point is selected on the test set; headline metrics may be optimism-biased and need a held-out threshold.","rationale":"The reader's weakest assumption focused on O4 internal homogeneity, but their rationale explicitly identified the more direct problem: the detection threshold appears to be fitted on the test set. My read agrees with that concern. The paper contains no validation split for threshold selection: Section 4.2 describes only a train/test split for noise and reserves all signals for test, while Section 3 promises validation-based threshold optimization. Section 4.4 then gives a specific τ selected to maximize F1, with no indication of a separate validation set. This is load-bearing because the strongest claim is the precise 97.0%/96.1%/F1=96.6% operating point on O4 data. If τ is chosen on the test set, these numbers are not honest out-of-sample estimates. The ROC-AUC of 0.994 is threshold-independent and would survive a stricter protocol, which is why I do not recommend moving to REJECT; the method may still genuinely separate signals from noise. But the headline operating-point claim needs correction or re-evaluation. The reader's CONDITIONAL verdict already captures this uncertainty, so I leave the verdict unchanged. The concrete test I propose—using a proper validation split for τ and reporting held-out metrics—directly settles whether the reported operating point is overfit.","tokens_in":11314,"tokens_out":3489,"duration_ms":39815,"concrete_test":"Re-run the evaluation with a strictly held-out threshold-selection split: randomly partition the 399 noise segments into validation/test (e.g., 200/199) and the 102 signals into validation/test (e.g., 51/51). Select τ on the validation split only (maximizing F1 or maintaining precision ≥90%) and then report precision/recall/F1 on the held-out test split. If the held-out metrics differ from 97.0%/96.1% by more than ~5 points, the reported operating point is overfit; report the held-out ROC-AUC as well. This single change would settle whether the headline numbers are reproducible under an honest evaluation protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 states that the detection threshold is 'optimized using precision-recall analysis on validation data,' but no validation split is described in Section 4.2: noise is divided into train (1592) and test (399), and all 102 signals are reserved for test. Section 4.4 then reports τ=0.667 'selected to maximize F1-score' and uses it to obtain 98 TP, 4 FN, 396 TN, 3 FP. With no independent validation set, τ must have been chosen on the same test set used to compute the headline precision/recall/F1. This makes 97.0% precision and 96.1% recall post hoc selections rather than unbiased estimates, so the central claim of 'exceptional performance' at a specific operating point is not yet supported. The ROC-AUC of 0.994 is threshold-independent and remains encouraging, but the exact F1-based comparison to supervised methods is inflated by test-set threshold fitting. This is a methodological leak, not an astrophysical one, and it directly affects the strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an unsupervised, template-free gravitational-wave detection method that combines continuous wavelet transform (CWT) preprocessing with an LSTM autoencoder. The autoencoder is trained on LIGO Hanford O4 noise segments only, and anomaly scores are reconstruction errors. The authors report that on a test set of 102 GWTC-4.0 O4 signals and 399 noise segments, the model achieves 97.0% precision, 96.1% recall, F1 96.6%, and ROC-AUC 0.994. They also report a discovery of cross-run calibration batch effects in multi-run training, motivating a single-run O4 training strategy that improved recall from 52% to 96%. The paper includes synthetic-data validation, a comparison with matched filtering and supervised/unsupervised baselines, and public code/data availability.","tokens_in":11541,"tokens_out":4322,"duration_ms":51252,"significance":"If the central performance claim were properly supported, this would be a meaningful result: a noise-only-trained, template-free anomaly detector performing close to supervised methods on real LIGO data would support discovery-oriented searches. The batch-effect analysis is also a useful cautionary result for machine-learning analyses of multi-epoch gravitational-wave datasets. Strengths of the paper include the explicit public release of code, data-processing scripts, and trained checkpoints, and the threshold-independent ROC-AUC/AP metrics, which are less affected by the main methodological issue. However, the headline precision/recall/F1 numbers are currently fitted to the test set, so the 'exceptional performance' claim is not yet established by the evaluation protocol.","major_comments":[{"comment":"The detection threshold τ is selected on the test set. Section 3 says τ is 'optimized using precision-recall analysis on validation data,' but Section 4.2 describes only a train/test split of noise (1592/399) and reserves all 102 signals for test; no validation split is described. Section 4.4 then reports τ=0.667 'selected to maximize F1-score' and uses that same test set to compute 97.0% precision and 96.1% recall. The headline operating point is therefore a post hoc fitted statistic, not an unbiased performance estimate. This directly undermines the central claim. The authors should hold out a validation set (or use nested/cross-validated threshold selection) and report metrics at that threshold, or restrict headline claims to threshold-independent metrics such as ROC-AUC and average precision.","section":"Sections 3, 4.2, 4.4"},{"comment":"The claimed improvement from multi-run training to single-run training, 'recall from 52% to 96%,' is not an apples-to-apples comparison. Section 6.1 describes multi-run training on O1–O4 data but does not specify the exact test set used for the 52% recall figure, its size, or whether it is the same O4 test set used in Section 4.4. If the multi-run evaluation included O1–O3 events with different calibration properties, the comparison conflates domain shift with detection performance. The authors should report both models on the same held-out O4 test set, and, if possible, per-run recall for the multi-run model.","section":"Abstract and Section 6.1"},{"comment":"The central assumption that single-run O4 H1 data are internally homogeneous is not directly validated. The autoencoder is trained on 1592 O4 noise segments, and the 102 GWTC-4 events are all O4. But O4 spans multiple calendar periods and possible calibration epochs, and the paper does not report any diagnostic for residual within-run drift, glitches, or non-stationarity. If the 399 test noise segments and the 102 signal windows differ in such artifacts, the reconstruction-error gap could be inflated beyond astrophysical signal content. A simple check would be to stratify noise and signal reconstruction errors by sub-period within O4, or to compare a noise-only holdout from a different O4 epoch.","section":"Sections 4.2 and 6.1"},{"comment":"The comparison in Table 2 is not methodologically transparent. Matched filtering is assigned 99.8% precision and >99% recall with no citation or definition of the operating point; the CNN and 'Unsupervised AE (Raw)' rows are described only as 'representative' of literature results with no quantitative sources. Because the CWT-LSTM row uses the test-set-fitted threshold from Section 4.4, the table does not support the claim of competitiveness with supervised methods. The authors should compare on the same O4 dataset and evaluation protocol, or clearly state that the baselines are illustrative only.","section":"Table 2 and Section 5"}],"minor_comments":[{"comment":"Typo: 'Rather that searching for known signal templates' should be 'Rather than searching.'","section":"Section 1"},{"comment":"The caption says 'The bright vertical band at ∼15.2 seconds marks the merger event,' but the top panel is described as a full 4-second window. This time value is inconsistent with a 4-second window; please clarify the time axis or correct the caption.","section":"Figure 3 caption"},{"comment":"Reference [16] is cited as 'GWTC-4... in preparation' with a placeholder arXiv number '2407.xxxxx'. This is not a complete citable reference; please update to the published or arXiv version.","section":"Reference [16]"},{"comment":"The preprocessing description says the data are whitened to 'zero mean and unit variance' and later that per-scale z-score normalization is computed on training noise. The relationship between these two normalization steps should be stated more precisely, since it is important for reproducibility.","section":"Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The central methodological flaw is fixable within the manuscript's scope: with public code and data, the authors can create a proper validation split for threshold selection or report threshold-independent metrics as primary. If that is done, the paper could become a useful contribution. I would also encourage the editor to verify that the GitHub repository is indeed accessible and contains the claimed scripts and checkpoints, since the reproducibility claim is a notable strength but should be checked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, the batch-effect finding is real and worth taking seriously: reconstruction errors cluster by observing run, not by astrophysical parameters, and single-run training recovers most of the missed signals. That is a legitimate, actionable warning for anyone doing ML on multi-epoch GW data. Second, ignore the headline 97% precision / 96% recall numbers. They are selected on the test set. The threshold is chosen to maximize F1 on the same 102 signals and 399 noise segments used to report the confusion matrix, so your optimistic bias is baked in. The ROC-AUC of 0.994 is threshold-independent and encouraging, but the exact operating point is not evidence of \"exceptional performance.\"\n\nWhat the paper does well: it is honest about the dead ends. The re-whitening attempt that killed the model and the negative results on global normalization are reported in detail. The architecture is a standard CWT + LSTM autoencoder, nothing exotic, but the combination is sensible. The code repository is promised with checkpoints and batch-effect scripts, which is more than many papers do.\n\nNow the soft spots, in proportion. The multi-run vs single-run comparison (52% to 96% recall) is not apples-to-apples because the multi-run test set draws from O1\\u2013O4 signals while the single-run evaluation only uses O4. The improvement is confounded by the change in test distribution. The \"template-free discovery\" claim is also overstated: every positive test example is a GWTC-4 catalog event, so the model is doing template-free recognition, not discovery of new signals. The synthetic validation AUC of 0.806 is modest, which makes the jump to 0.994 on real data suspicious and worth probing. Minor issues: the GWTC-4 reference has a placeholder arXiv ID, and the repository was not independently run.\n\nWho gains from this: anyone building ML pipelines for GW data, especially those doing multi-run analyses. The batch-effect warning is the citable core. The detection-performance claim needs rework.\n\nRecommendation: send to peer review, but only after the authors fix the evaluation protocol. Require a held-out threshold selection (e.g., a validation split or the synthetic data to set tau), a matched test set for the single-run/multi-run comparison, and a genuine search over O4 noise for new candidates to back the discovery claim. If the batch-effect section were extracted and published on its own, I would accept it as is.","headline":"A useful case study of cross-run batch effects in GWOSC data, but the headline detection metrics are inflated by test-set threshold fitting.","tokens_in":12021,"tokens_out":1748,"would_cite":true,"duration_ms":21516,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A CWT-LSTM autoencoder trained only on O4 noise detects confirmed gravitational waves at 97.0% precision and 96.1% recall, once run-dependent calibration differences are removed.","keywords":["gravitational waves","anomaly detection","LSTM autoencoder","continuous wavelet transform","template-free detection","LIGO O4","batch effects","GWOSC calibration"],"falsifier":"Re-run the exact O4 experiment after applying a single common re-whitening PSD to all segments (the paper's own multi-run attempt of this caused ROC-AUC to fall to 0.44); if single-run O4 performance survives unified re-whitening, the gap is astrophysical, and if it collapses, the reported separation depends on run-specific whitening artifacts. Alternatively, test the trained model on O4 events with network SNR near the catalog threshold: if recall falls sharply there, the model may only be recognizing loud, already-obvious events.","tokens_in":11184,"feed_emoji":"📡","tokens_out":6898,"duration_ms":68271,"temperature":0.7,"pith_summary":"The paper tries to show that a gravitational-wave detector can be built without waveform templates: an unsupervised autoencoder learns what LIGO O4 noise looks like from a continuous-wavelet-transform scalogram, then flags anything that does not reconstruct well. On 102 confirmed O4 events and 399 noise segments, the model reports 97.0% precision and 96.1% recall, with 4 missed signals and 3 false alarms. The paper's second claim is about data hygiene: training on combined O1–O4 data makes reconstruction errors cluster by observing run rather than by astrophysical source, a batch effect traced to GWOSC's evolving calibration and whitening. Restricting training to O4 data removes that effect and raises recall from 52% to 96% at the same precision. If this holds, template-free anomaly detection can be competitive with supervised matched filtering and can search for signal morphologies no one has modeled.","feed_headline":"Noise-trained model catches 96% of O4 gravitational waves","feed_subtitle":"Trained only on detector noise, the autoencoder matches supervised detection once calibration drift across runs is removed.","key_machinery":"The key mechanism is the reconstruction-error anomaly score. Each 32-second strain segment is high-passed at 15 Hz, low-passed at 1024 Hz, whitened, downsampled to 1024 Hz, and transformed with a Morlet continuous wavelet transform over 8 scales (20–512 Hz) into an 8×4096 scalogram; log compression and z-score normalization use statistics computed only on training noise. A bidirectional LSTM autoencoder with 64 hidden units and a 32-dimensional bottleneck is trained to minimize mean squared reconstruction error on noise segments. At test time, mean squared error between input and reconstruction is computed, and segments above threshold 0.667 are flagged as gravitational-wave candidates. The","core_discovery":"Central claim: a noise-trained LSTM autoencoder on CWT scalograms can separate confirmed gravitational-wave events from detector noise without templates. Trained on 1,592 O4 H1 noise segments, it gets 97.0% precision, 96.1% recall, 99.2% specificity, and ROC-AUC 0.994 (threshold 0.667) on 102 events plus 399 noise segments. Error distributions are unimodal per class: noise mean 0.48, signal mean 0.77. Trained on O1–O4 instead, errors cluster by observing run (Spearman ρ=0.68) not astrophysical parameters (|r|<0.15), and re-whitening collapses discrimination (AUC 0.44). The paper concludes per-run training is the correct practice for multi-epoch gravitational-wave machine learning.","pith_inferences":["The paper's own re-whitening failure hints that the O4 model may be using run-specific spectral fingerprints as part of its signal discriminator; a calibration-invariant assessment, e.g., training on one run's noise and testing on another run's confirmed events, would show how much of the 0.29 error gap is truly astrophysical.","Because all 102 test events are already confirmed detections, the reported recall measures recognition of known signals, not discovery of new ones; the template-free discovery claim remains untested until the same model is run over unlabeled O4 data and its outliers are followed up.","The method still encodes a frequency-band assumption through the chosen CWT scales (20–512 Hz); signals outside that band, such as long-duration or very-high-frequency transients, may be invisible to it despite the absence of templates.","A straightforward extension would be a two-detector coincidence rule: flag only events with high reconstruction error in both LIGO Hanford and Livingston O4 data, which should cut the 3 false alarms and sharpen precision."],"forward_implications":["Template-free detection becomes a practical complement to matched filtering, able in principle to flag signals whose waveforms are not in any template bank.","Per-run training emerges as a reusable rule for machine learning on multi-epoch gravitational-wave data; combined-run models should either be domain-adapted or treated with caution.","The O4-trained model can be retrained on earlier runs to run high-recall archival searches of O1–O3 data without inheriting cross-run calibration artifacts.","Tuning the reconstruction-error threshold trades false alarms against sensitivity, so the method could be used as a low-latency candidate generator for follow-up.","Because training needs only noise, the method can be deployed before a new signal catalog exists for a new observing run."],"supporting_citations":[{"why":"Supplies all LIGO H1 strain data used for training noise segments and test events.","marker":"[13]"},{"why":"Defines the 102 confirmed O4 events from GWTC-4 used as the labeled test set.","marker":"[16]"},{"why":"Provides the GWpy data-access method used to fetch and segment strain time series.","marker":"[14]"},{"why":"Establishes matched filtering as the standard detection method whose template requirement the paper claims to remove.","marker":"[1]"},{"why":"Supplies the time-frequency decomposition rationale that motivates CWT preprocessing.","marker":"[7]"},{"why":"Supplies the LSTM architecture used to model temporal dependencies in the scalograms.","marker":"[12]"},{"why":"Represents prior unsupervised autoencoder gravitational-wave detection that this work extends from raw time series to CWT representations.","marker":"[10]"}],"fun_headline_variants":["Noise-trained autoencoder catches 96% of O4 waves","Unsupervised model hits 96% recall after fixing run drift","Template-free LSTM autoencoder detects 96% of O4 waves","Autoencoder trained on noise alone matches supervised detection"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that O4 H1 GWOSC strain, after the paper's filtering and whitening, is internally uniform enough that an autoencoder trained on 1,592 O4 noise segments captures all normal noise; if calibration drift, glitches, or unresolved weak signals within O4 remain, the error gap between the 102 known signals and noise could be inflated by artifacts rather than by astrophysical signal content.","fun_headline_variants_meta":{"raw":{"variants":["Noise-trained autoencoder catches 96% of O4 waves","Unsupervised model hits 96% recall after fixing run drift","Template-free LSTM autoencoder detects 96% of O4 waves","Autoencoder trained on noise alone matches supervised detection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1441,"prompt_tokens":897,"completion_tokens":544,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":472}},"tokens_in":641,"tokens_out":544,"duration_ms":6224,"temperature":1.0,"reasoning_tokens":472,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:10:49.943453+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the exact O4 experiment after applying a single common re-whitening PSD to all segments (the paper's own multi-run attempt of this caused ROC-AUC to fall to 0.44); if single-run O4 performance survives unified re-whitening, the gap is astrophysical, and if it collapses, the reported separation depends on run-specific whitening artifacts. Alternatively, test the trained model on O4 events with network SNR near the catalog threshold: if recall falls sharply there, the model may only be recognizing loud, already-obvious events.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies all LIGO H1 strain data used for training noise segments and test events."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 102 confirmed O4 events from GWTC-4 used as the labeled test set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the GWpy data-access method used to fetch and segment strain time series."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes matched filtering as the standard detection method whose template requirement the paper claims to remove."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the time-frequency decomposition rationale that motivates CWT preprocessing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the LSTM architecture used to model temporal dependencies in the scalograms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents prior unsupervised autoencoder gravitational-wave detection that this work extends from raw time series to CWT representations."}],"review_version":1}