{"id":"497f4feb-6cde-4037-816b-6459311dd304","arxiv_id":"2509.05283","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AresGW model 1's injection detection count at a false-alarm rate of 1/month varies with noise dataset by up to 39% coefficient of variation, while sensitive distance varies by only a few percent.","lead":"This paper measures how stable a machine-learning gravitational wave search's sensitivity is when test datasets with real detector noise are swapped. It finds that noise choice causes large swings in detection counts at low false-alarm rates, while sensitive distance is steadier, and leftover real signals in noise can distort scores.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CV comparison rests on 10 non-independent, overlapping one-month windows; the 'up to 39%' CI is an artifact of treating a deterministic grid as a random sample.","rationale":"The reader's weakest_assumption concerns the AresGW catalog used for contamination cleaning. That is a legitimate concern for the paper's secondary claim about contamination, but it does not directly threaten the central empirical finding about metric variability: the CV comparison in Tables III–VI is computed on contaminated datasets and remains similar after cleaning (Tables VIII–IX). The more load-bearing issue for the central claim is the statistical foundation of the CV estimates. With only 10 datasets per category, selected as a deterministic grid from a single 81-day noise file, and with several datasets overlapping in time, the assumption of independent random sampling in Appendix C is not met. The confidence intervals, and especially the 'up to 39%' upper bound, are therefore not reliable. The qualitative direction—counts more variable than sensitive distance—is plausible and likely robust, but the quantitative magnitude and the claim that sensitive distance is preferable when few datasets are available need to be validated on a properly independent sample. This does not change the overall conditional verdict, but it identifies a different condition than the reader did: the empirical sample must be representative and large enough for stable variance estimation.","tokens_in":26206,"tokens_out":13932,"duration_ms":146207,"concrete_test":"Construct a much larger set of independent one-month datasets from the full O3a (or O3a+O3b) observing run, using non-overlapping 30-day windows with identical fixed injections (same source parameters, appropriate GPS times). For each dataset, evaluate AresGW model 1 and compute N_F and S at FAR=1/month. Recompute sample CVs and bootstrap CIs on this larger sample (e.g., n=50). If the CV gap narrows substantially (e.g., count CV <10% or sensitive-distance CV >10%), the paper's central recommendation is not robust. Alternatively, apply a block-bootstrap to the existing 10 datasets that accounts for overlapping windows; if the upper bound of the count CV drops below ~15%, the 'up to 39%' headline is an artifact of the small non-independent sample.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim—that the number of detected injections at FAR=1/month has a coefficient of variation up to ~39% while the sensitive distance has CV up to ~4%—is derived from 10 datasets per category. These datasets are not independent random draws from a common distribution. They are a deterministic grid of offsets and seeds on a single 81-day O3a noise file, with several windows overlapping substantially (offsets 0, 10, 20, 30, 40, 46.3 days for 30-day datasets). Appendix C computes confidence intervals using t, chi-square, and BCa bootstrap methods that assume independent sampling; with n=10 and overlapping windows, the nominal coverage is not valid. The observed point CV for detection counts is about 21%, and the 39.1% upper bound is a small-sample chi-square artifact. The spread is also driven largely by two quiet months at offsets 40 and 46.3. The sensitive distance CV, by contrast, is a weighted integral over thousands of injections, so it is mechanically smoother. If the datasets are not representative independent one-month realizations, the quantitative superiority of sensitive distance is not established at the claimed precision.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates the robustness of AresGW model 1, an ML-based gravitational-wave search pipeline, across 28 one-month datasets constructed from O3a LIGO noise. Three dataset categories are considered: identical noise with varying injections, varying noise with identical injections, and both varying. For each category the authors compute the number of detected injections and the sensitive distance at FAR = 1, 10, and 100 per month, with confidence intervals for the mean, standard deviation, and coefficient of variation using t, chi-square, and BCa bootstrap methods after Shapiro-Wilk normality checks. The central claims are that detection counts are much more variable across noise realizations than the sensitive distance (CV upper bound ~39% versus ~4% at FAR = 1/month), that noise variability dominates over injection variability, and that contamination by real GW events—especially GW190511_125545—biases detection-count metrics, so clean datasets and multiple metrics are recommended.","tokens_in":26518,"tokens_out":4228,"duration_ms":44957,"significance":"If the conclusions hold, the paper provides practically useful benchmarking guidance for ML-based GW search pipelines and makes concrete reproducibility artifacts public (seeds, offsets, injection files). The use of real O3a noise and the comparison of two standard sensitivity metrics address a real methodological gap. The paper is also commendable for reporting explicit statistical procedures and for identifying a specific contamination effect that could affect earlier evaluations. However, the central quantitative claim relies on statistical inference from ten small, overlapping, non-randomly sampled datasets, and the contamination analysis uses the authors' own catalog as ground truth; both issues need to be addressed before the conclusions can be considered robust.","major_comments":[{"comment":"The ten datasets in the varying-noise categories are not independent random draws from a common distribution. They are a deterministic grid of offsets and seeds on a single 81-day noise file, with several 30-day windows overlapping substantially (e.g., offsets 40 and 46.3 days overlap by ~23.7 days, and multiple offset-0 rows use the same time span). Appendix C explicitly assumes independent observations for the t, chi-square, and bootstrap intervals. With n=10 and overlapping windows, the nominal 95% coverage is not valid. This directly affects the headline quantitative contrast: the upper CV bound of 39.1% for detection counts at FAR=1/month is a small-sample chi-square artifact, and the point CV itself is driven by two low-count windows. The authors should either construct non-overlapping independent month-long datasets, use a method that accounts for the dependence (e.g., block boots","section":"Sec. III; Tables III, V; Appendix C"},{"comment":"The contamination-cleaning step uses events from the AresGW catalog—produced by the authors' own model 2—as ground truth for which signals are real and should be removed. If some or all of those events are false positives of model 2, then removing them removes the model's own noise artifacts, and the measured increase in detection counts after cleaning is at least partly an artifact. This is load-bearing for the claim that dataset contamination by real GW events substantially biases detection-count evaluations and for the recommendation in Sec. VI that all known events be removed. The authors should redo the cleaning using only independently confirmed events (e.g., GWTC-2.1, OGC, IAS) or otherwise provide evidence that the AresGW candidates are genuine signals.","section":"Sec. V; Table X; Sec. VI"},{"comment":"The comparison of the coefficient of variation of detection counts with that of sensitive distance is not like-for-like. The sensitive distance is a weighted integral over thousands of injections (Appendix D), so it is expected to be smoother than a raw count even under identical underlying detection performance. The claim that 'sensitive distance is more robust' should be framed as a property of the estimator, not necessarily as evidence about the stability of the model's detection capability. A more convincing analysis would compare, on the same non-overlapping datasets, the bootstrap distributions of both metrics, or explicitly discuss the bias-variance trade-off introduced by the chirp-mass weighting in the sensitive distance.","section":"Sec. IV B; Sec. IV C; Appendix D"}],"minor_comments":[{"comment":"The text after Table I states that the CIs for sigma_N/mu_N are [0.4%,1.0%], [0.4%,1.0%], and [0.5%,1.2%], but Table I reports [0.8%,2.0%], [0.7%,1.9%], and [0.7%,2.0%]. The text values match Table II instead. This inconsistency should be corrected.","section":"Sec. IV A; Table I"},{"comment":"The presentation of the mean CI in Table IV is inconsistent with other tables: it is given as an interval rather than mean +/- margin. Also, 'sigma_S/mu_s' uses a lowercase 's' in the header, and dimensionless CV intervals are labeled with 'Mpc'.","section":"Table IV and Table VI"},{"comment":"The text says the datasets contain 'approximately 40 real signals,' while Table X lists 37 events. The count should be made consistent and the source of the discrepancy clarified.","section":"Sec. IV B; Table X"},{"comment":"The datasets are described as 'independent one-month datasets,' but they are overlapping windows drawn from one 81-day file. The wording should be adjusted to avoid overstating independence, especially in the abstract and introduction.","section":"Sec. I and Sec. III"},{"comment":"The category 'varying noise and identical injections' is not strictly identical in received signal: because the injections are placed at different GPS times, the antenna response changes. The text acknowledges this but later attributes all differences to noise; the caveat should be carried through the interpretation.","section":"Sec. IV B"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a data-analysis/methods journal but is more of a benchmarking study than a new astrophysical result. The most serious concern is statistical: the central quantitative claim is built on CIs computed from ten overlapping, deterministically chosen windows, which violates the stated independence assumption. The contamination analysis also relies on the authors' own catalog as ground truth, which is a circularity risk even if unintentional. These issues are fixable with additional analysis or a more cautious presentation, so I do not recommend rejection, but the current version should not be accepted without substantial revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thanks for the report. I read the paper with your notes and the stress-test in hand. My take: this is a worthwhile paper, but it needs a serious statistical rework before its numbers can be believed as stated.\n\nWhat's genuinely new: the paper actually runs the AresGW model over 28 one-month windows of O3a noise and shows that the detection count at a fixed FAR fluctuates much more across noise realizations than the sensitive distance does. That's a real, quantified observation, and it directly extends Nagarajan and Messenger's single-month analysis. The contamination result—that a single high-ranking event (GW190511_125545) can push the FAR threshold and suppress injection recovery by ~560 detections—is striking and, if the event is real, an important caution for anyone benchmarking on MLGWSC-1 data. The authors also ship seeds, offsets, and code, which is good practice.\n\nThe soft spots are real, and the stress-test correctly identifies the biggest one. The 'varying noise' categories are not 10 independent draws from a common distribution; they're a deterministic grid of offsets and seeds on a single 81-day noise file, with several 30-day windows overlapping substantially. The chi-square and BCa intervals in Table III and Appendix C assume independent random sampling. With n=10 and overlapping windows, those intervals have no valid coverage. The 'CV up to 39%' is a small-sample chi-square upper bound; the point CV is about 21%, and the spread is driven almost entirely by the two quiet months at offsets 40 and 46.3. That doesn't mean the qualitative conclusion is wrong—detection counts are clearly more noise-sensitive than sensitive distance—but the precision claimed is not supported.\n\nThe second issue is the contamination-cleaning. The key event is from the authors' own AresGW catalog (model 2), not confirmed independently. If that event is a false positive, the 'contamination bias' is partly an artifact of removing your own model's noise artifact. The authors should at least acknowledge this or seek independent confirmation.\n\nThe sensitive distance being smoother is also partly mechanical, since it averages over thousands of injections with a chirp-mass weighting; the paper's claim that it's 'preferable' is reasonable, but the comparison is less clean than the presentation implies.\n\nWho's this for? Anyone working on ML-GW search benchmarking. The paper deserves a serious referee, but I would not accept it as-is. The statistical framework needs to be rebranded as a sensitivity analysis over a grid—not inferential statistics—or the authors need to generate genuinely independent, non-overlapping datasets (e.g., from multiple noise files) and redo the CIs. The contamination discussion is valuable but needs a stronger ground-truth defense.","headline":"Worthwhile empirical study of metric variance for ML-GW search, but the headline CV numbers rest on non-independent datasets and the contamination claim depends on the authors' own unverified catalog.","tokens_in":26922,"tokens_out":4411,"would_cite":true,"duration_ms":42659,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Detection counts for a machine-learning gravitational-wave search swing up to ~39% between one-month noise files, while the sensitive distance changes only ~4%; a single leftover real event can bias counts by hundreds, but barely moves the","keywords":["gravitational-wave detection","machine-learning search","sensitive distance","false alarm rate","benchmark variability","dataset contamination","AresGW model","LIGO O3a noise"],"falsifier":"Run the paper's 28-dataset protocol (same O3a noise, injections, FARs) with a standard matched-filtering search such as PyCBC: if its detection-count coefficient of variation at FAR=1/month is near its sensitive-distance value (~4%) rather than near 39%, the metric-stability ranking is specific to AresGW model 1, not general. Separately, rebuild the clean datasets keeping only AresGW-catalog events that a matched-filter search also confirms; if the ~560-injection gap at FAR=1/month vanishes, the contamination claim depends on those events being genuine.","tokens_in":26153,"feed_emoji":"🔭","tokens_out":11570,"duration_ms":108653,"temperature":0.7,"pith_summary":"This paper asks how much a machine-learning gravitational-wave search's measured sensitivity depends on which month of real detector noise you test it on. Running AresGW model 1 on 28 one-month LIGO O3a datasets, the authors find that the number of detected injections at a fixed false-alarm rate is highly noise-dependent (coefficient of variation up to about 39% at FAR=1/month), while the sensitive distance is much more stable (coefficient of variation up to about 4%). They also show that real gravitational-wave signals left in the evaluation noise can seriously bias count-based results: a single high-ranking event, GW190511 125545, suppresses detection counts by roughly 560 injections at FAR=1/month while barely moving the sensitive distance. The paper concludes that sensitive distance is the more reliable benchmark when only a few datasets are available, that known signals must be removed from evaluation data, and that single-month evaluations at low false-alarm rates should be treated with caution.","feed_headline":"Gravitational-wave search counts swing 39% between noise months","feed_subtitle":"The sensitive distance metric stays within ~4%, and one leftover signal can skew benchmark results.","key_machinery":"Two metrics are compared at fixed false-alarm rates (1, 10, 100/month): the number of detected injections, and the sensitive distance — the radius of a sphere equal to the sensitive volume, with each found injection reweighted by (Mc/Mmax)^(5/2). Variability is quantified by the coefficient of variation with 95% confidence intervals across three dataset categories — identical noise/varying injections, varying noise/identical injections, both varying — that isolate the sources of scatter. The contamination result hinges on 37 real events (Table X) left in the MLGWSC-1 reference noise, above all GW190511 125545, discovered by the authors' AresGW model 2 and highly ranked by model 1; its presen","core_discovery":"With noise fixed and only injections changed, AresGW model 1's detection counts and sensitive distance are both stable (coefficient of variation under ~2%). With noise varied and injections fixed — the condition that matters in practice — detection counts at FAR=1/month vary with coefficient of variation up to 39%, while sensitive distance stays within ~3%; the pattern holds when both vary. The reference dataset used in earlier evaluations still contained ~40 real signals; removing them raises its detection count at FAR=1/month from 2892 to 3475, driven mainly by GW190511 125545, while sensitive distance moves only from 1574.6 to 1589.9 Mpc. Conclusion: detection counts are fragile benchmark","pith_inferences":["The contamination bias is strongest for events discovered by the same pipeline family being tested: if future benchmarking datasets are curated using only independently confirmed catalogs, the measured ~560-injection effect at FAR=1/month may shrink or disappear — a testable prediction of the paper's own logic.","Because the sensitive-distance reweighting scales as chirp mass to the 5/2 power, its stability advantage may not extend to injections outside AresGW model 1's effective training range (chirp mass below 10 or above 40 solar masses); binning the same 28 datasets by chirp mass would settle this.","If the noise-dominance finding generalizes, next-generation detectors with stronger non-stationarity will make single-month benchmarking even less representative, pushing evaluation protocols toward longer effective durations or explicitly noise-marginalized metrics.","Applying the same protocol to a matched-filtering pipeline would reveal whether the 39%-versus-4% gap between metrics is specific to the AresGW architecture or a general property of count-based metrics; the paper does not include that control."],"forward_implications":["At FAR=1/month, a single one-month evaluation of detection counts can misrepresent AresGW model 1's sensitivity by tens of percent; reporting the sensitive distance instead cuts the variability to a few percent.","Prior comparisons based on the MLGWSC-1 real-noise month without removing non-GWTC-2 signals (such as the detection-count numbers in [78]) are biased: cleaning the data raises the count at FAR=1/month from 2892 to 3475 for the same model and dataset.","Benchmarking reports should state mean, standard deviation, coefficient of variation, and confidence intervals computed over multiple independent datasets rather than relying on one dataset.","Because sensitive distance has its own chirp-mass dependence, the paper recommends reporting both metrics together instead of either alone.","Variability at low FAR is dominated by noise realization rather than injection realization, so drawing more injection sets cannot substitute for testing across different noise months."],"supporting_citations":[{"why":"Supplies the reference one-month O3a real-noise dataset, the injection-generation script used to build all 28 test datasets, and the sensitive-distance definition the paper evaluates.","marker":"[69]"},{"why":"The prior reanalysis whose single-month, count-based comparison at FAR=1/month the paper re-evaluates and shows to be contaminated and noise-sensitive.","marker":"[78]"},{"why":"The AresGW model 2 catalog is the source of GW190511 125545 and other unremoved events; this event drives the measured contamination bias.","marker":"[63]"},{"why":"Defines which events were already removed from the original MLGWSC-1 reference noise, establishing the paper's contaminated baseline.","marker":"[70]"},{"why":"Provides GWTC-2.1 events that the paper additionally removes when constructing the clean evaluation datasets.","marker":"[71]"},{"why":"Describes the Virgo-AUTH predecessor of AresGW model 1 and its earlier variance estimate; the model under test belongs to this lineage.","marker":"[76]"},{"why":"Supplies the bias-corrected-and-accelerated bootstrap method used for confidence intervals when normality is rejected.","marker":"[82]"}],"fun_headline_variants":["GW search counts vary 39% month to month","Detection counts fragile, sensitive distance stable","Leftover real signal skews GW detection benchmark","Counts swing 39%, distance robust in GW search","GW benchmark: counts vary, distance holds"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The 'clean' evaluation datasets are made by deleting 37 events that include candidates reported only by the authors' own machine-learning pipeline (AresGW model 2, reference [63]); if any of those candidates are not real gravitational waves, the measured jump in detection counts after cleaning would partly reflect deletion of the pipeline's own noise artifacts rather than removal of genuine signals.","fun_headline_variants_meta":{"raw":{"variants":["GW search counts vary 39% month to month","Detection counts fragile, sensitive distance stable","Leftover real signal skews GW detection benchmark","Counts swing 39%, distance robust in GW search","GW benchmark: counts vary, distance holds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000103,"raw_usage":{"total_tokens":867,"prompt_tokens":747,"completion_tokens":120,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":63}},"tokens_in":491,"tokens_out":120,"duration_ms":2259,"temperature":1.0,"reasoning_tokens":63,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T05:23:50.528025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's 28-dataset protocol (same O3a noise, injections, FARs) with a standard matched-filtering search such as PyCBC: if its detection-count coefficient of variation at FAR=1/month is near its sensitive-distance value (~4%) rather than near 39%, the metric-stability ranking is specific to AresGW model 1, not general. Separately, rebuild the clean datasets keeping only AresGW-catalog events that a matched-filter search also confirms; if the ~560-injection gap at FAR=1/month vanishes, the contamination claim depends on those events being genuine.","supporting_citations":[{"cited_title":"Abbott et al","cited_arxiv_id":null,"evidence_quote":"The prior reanalysis whose single-month, count-based comparison at FAR=1/month the paper re-evaluates and shows to be contaminated and noise-sensitive."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The AresGW model 2 catalog is the source of GW190511 125545 and other unremoved events; this event drives the measured contamination bias."},{"cited_title":"Convolutional Neural Networks for signal detection in real LIGO data","cited_arxiv_id":"2402.07492","evidence_quote":"Defines which events were already removed from the original MLGWSC-1 reference noise, establishing the paper's contaminated baseline."},{"cited_title":"Neural network time-series classifiers for gravitational-wave searches in single-detector periods","cited_arxiv_id":"2307.09268","evidence_quote":"Provides GWTC-2.1 events that the paper additionally removes when constructing the clean evaluation datasets."}],"review_version":1}