{"id":"3aec7911-570f-4c6c-a03b-19202c935b77","arxiv_id":"2512.04516","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hybrid matched-filter/deep-learning pipeline recovers 31 known O3 events and reports a new tentative high-mass candidate, with sensitivity comparable to existing searches only for chirp masses above 25 solar masses.","lead":"This paper applies a hybrid matched-filter/deep-learning pipeline to LIGO's third observing run, recovering 31 previously known events and reporting a new tentative candidate (GW190929 091722). The work shows comparable sensitivity to existing searches only for chirp masses above 25 solar masses, with a notable drop at lower masses.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Candidate GW190929 091722's p_astro=0.63 is inconsistent with the paper's own foreground model: with network SNR 6.7 < O3a threshold ρ_th=8.5, Eq. (3) gives p1=0, so p_astro should be 0; the reported value implies an off-support evaluation.","rationale":"The reader's verdict (CONDITIONAL) identifies the p_astro calibration and the candidate's subthreshold SNR as the main weakness. My concern is more specific and more severe: the reported p_astro=0.63 is not merely uncertain, it is inconsistent with the paper's own Eq. (3) when applied to a trigger with SNR below the tuned threshold. This is a concrete, checkable internal inconsistency rather than a generic 'near edge of training distribution' worry. If the inconsistency is confirmed, the central claim of a previously unreported promising candidate collapses, though the methodological contribution (hybrid pipeline sensitivity, unique injection recovery, missed-candidate analysis) can stand on its own. The paper already provides a GitHub data release, so the proposed test is feasible. I do not recommend outright rejection because the paper's main scientific value may not depend solely on this one candidate, and the authors could correct the p_astro implementation or soften the claim. I recommend CONDITIONAL with the specific condition that the p_astro for GW190929 091722 be recomputed with the proper support cutoff and either confirmed or the candidate re-characterized as a sub-threshold trigger. This partially agrees with the reader's weakest_assumption, but the reader's framing as 'calibration uncertainty' understates the issue: the model as written assigns zero probability to this trigger, so the onus is on the authors to demonstrate how p_astro=0.63 was obtained.","tokens_in":33324,"tokens_out":7343,"duration_ms":67603,"concrete_test":"Recompute p_astro for GW190929 091722 using the code and inputs in the GitHub release (github.com/damonbeveridge/new_BBH_O3_candidate-UWA), explicitly enforcing the support condition p1(ρ)=0 for ρ<ρ_th (ρ_th=8.5 for O3a). If the recomputed p_astro drops below 0.5 (or to 0), the candidate claim is unsupported. As a cross-check, run the O3a injection set and count how many injections with network SNR < 8.5 receive p_astro > 0.5; a substantial fraction would confirm the foreground model is being evaluated off-support. Also inspect the p_astro code to verify whether the SNR threshold is applied as a hard cutoff.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central novel claim is the new candidate GW190929 091722 with p_astro=0.63 (Table IV). The p_astro model in Sec. II E 2 uses a foreground SNR distribution p1(ρ) = 3ρ_th^3/ρ^4 (Eq. 3), which is the density of detected signals above a threshold ρ_th. For any trigger with ρ < ρ_th, this density is zero: the trigger lies outside the support of the foreground population. For O3a the tuned threshold is ρ_th=8.5 (Sec. II E 2). The candidate occurred on 2019-09-29, which is in O3a, and its network SNR is reported as 6.7 (Table IV). Therefore Eq. (3) gives p1(6.7)=0, so the Bayes factor K=0 (Eq. 4) and p_astro=0. The reported p_astro=0.63 can only arise if the implementation evaluates the power-law formula below its support (i.e., extrapolates the ρ^{-4} tail into a region where no foreground trigger should exist) or if the variable ρ in Eq. (3) is not the network SNR but some other ranking statistic (e.g., the unbounded deep-learning output) while the text and tuning description call it SNR. Either way, the stated significance of the candidate is not justified by the model as written. The paper acknowledges uncertainty in Sec. IV B ('reasonable amount of uncertainty') but does not address this hard cutoff. This is not a matter of calibration near the edge of the training distribution; it is an internal inconsistency between the reported p_astro and the stated statistical model. If p_astro is actually ~0, the claim of a 'promising new candidate' loses its evidential basis, although the sensitivity and missed-candidate analyses may remain valid.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a hybrid gravitational-wave search pipeline for binary black hole (BBH) mergers, combining matched filtering with a deep-learning classifier that acts on the signal-to-noise ratio (SNR) time series. The pipeline is applied to Advanced LIGO O3 data from Hanford and Livingston. The authors first benchmark sensitivity using the official GWTC-3 injection set and report sensitivity comparable to existing pipelines for source-frame chirp masses above ~25 M_sun, with decreased sensitivity at lower masses. They then conduct an offline O3 search, recovering 31 GWTC-3 candidates, one IAS higher-order-mode candidate, and one previously unreported candidate (GW190929 091722) with a reported p_astro of 0.63. Parameter estimation for this new candidate yields a high total mass (median 171 M_sun) and a substantial probability for an intermediate-mass primary black hole. The paper also emphasizes the large number of uniquely detected injections produced by the deep-learning ranking statistic compared to template-bank-only searches.","tokens_in":33766,"tokens_out":8517,"duration_ms":77159,"significance":"The work is potentially significant as an independent, open-data search that demonstrates a deep-learning ranking statistic can recover known events and identify new candidates. The injection study uses the official GWTC-3 injection set and includes comparisons with multiple established pipelines, which is a strength. The reported new candidate, if real, would be a high-mass BBH possibly containing an intermediate-mass black hole, making it astrophysically interesting. However, the central novelty—the p_astro=0.63 for GW190929 091722—relies on a p_astro model that appears internally inconsistent with the stated foreground distribution, as detailed in the major comments. The sensitivity and unique-detection claims may also be affected by this issue for low-SNR injections. These concerns must be resolved before the results can be accepted.","major_comments":[{"comment":"The reported p_astro=0.63 for GW190929 091722 is inconsistent with the foreground model as written. Eq. (3) defines p1(ρ) = 3ρ_th^3/ρ^4 as the SNR distribution of detected signals above an SNR threshold ρ_th. For O3a the tuned threshold is ρ_th=8.5. The candidate occurred in O3a and has network SNR 6.7 (Table IV). Since 6.7 < 8.5, the support of Eq. (3) excludes this trigger, so p1(6.7)=0, K=0 in Eq. (4), and p_astro=0 by Eq. (1). The reported value 0.63 implies either that the implementation evaluates the power-law formula below its support or that the quantity used in Eq. (3) is not the network SNR reported in Table IV. Either way, the stated statistical model does not justify the candidate's significance. This is a load-bearing issue because the new candidate is a central claim of the paper.","section":"Section II E 2, Eqs. (3)-(4); Table IV"},{"comment":"The same support issue likely affects the injection study. The paper states that the p_astro model is tuned on injections with a minimum network SNR of 6 (Section IV B), yet the O3a ρ_th is 8.5. If Eq. (3) is applied to any injection or candidate trigger with SNR between 6 and 8.5, the foreground density is zero on its support, meaning p_astro should be zero; if instead the implementation extrapolates the ρ^{-4} tail, the resulting p_astro values are not normalized and may be inflated. Since the detection threshold used throughout the sensitivity study is p_astro≥0.5, the reported total and unique detection counts (e.g., Table II) and ⟨VT⟩ values could be biased for low-SNR injections. The authors should quantify how many of the 23,238 detected injections have network SNR < ρ_th and demonstrate that their p_astro calculation is valid for those triggers, or recompute the sensitivity resul","section":"Section II E 2; Section III, Figs. 3-4; Table II"},{"comment":"The new candidate's estimated parameters place it near or beyond the edge of the deep-learning model's training distribution. The training set uses component masses 2–100 M_sun, while the posterior for GW190929 091722 gives 44.5% probability for m1 > 120 M_sun and a median total mass 171 M_sun. The paper notes this issue but does not provide a quantitative test of the model's ranking statistic for such out-of-distribution signals. Given that the Hanford peak is below the usual single-detector SNR threshold of 4 and the model was not trained on such samples, the significance measure (p_astro) is not reliable without additional validation—e.g., applying the pipeline to the LVK IMBH injection set or a dedicated high-mass, low-SNR injection campaign. The current qualitative caveat is insufficient to support the 'promising new candidate' claim.","section":"Section IV B; Section IV B 1; Table V"}],"minor_comments":[{"comment":"The header uses the same symbol 'M' for both total mass and chirp mass; the second column should be labeled with a distinct symbol, e.g., \\mathcal{M}, to avoid ambiguity.","section":"Table V"},{"comment":"The vector x is used generically in Eq. (1), but it is not explicitly defined as the pair (ρ, R) until later. Please define the ranking-statistic vector clearly before Eq. (1).","section":"Section II E 2, Eq. (1)"},{"comment":"The 'p-astro4 package' is referenced by URL only; please provide a formal citation or a version number to support reproducibility.","section":"Section II E 2"},{"comment":"The star symbols marking the trigger peaks are hard to discern, especially for the Hanford subthreshold peak. Consider adding an inset or a zoom of the SNR time series around the trigger.","section":"Figure 9"},{"comment":"In the sentence 'the candidate has a low SNR, resulting in unconstrained priors,' specify whether this refers to the network SNR, the original search SNR, or the reweighted SNR.","section":"Section IV A 2"},{"comment":"The ellipsis notation for absent triggers is explained only in the caption; repeating this explanation in the text just before the tables would improve readability.","section":"Tables VI and VII"},{"comment":"The sentence 'A common concern in developing deep learning models is explainability...' is somewhat disconnected from the preceding discussion; consider integrating it more smoothly or moving it to the model description section.","section":"Section IV B"}],"recommendation":"major_revision","confidential_remarks":"The injection study and the recovery of known GWTC-3 events provide a useful benchmark, and the authors have made their parameter-estimation products publicly available. However, the new candidate's p_astro is calculated from a model that, as written, assigns zero foreground probability to its SNR. This is an internal inconsistency, not just a calibration concern, and it directly undermines the paper's central claim. The authors need to either correct the p_astro implementation, justify the extrapolation below ρ_th with a concrete statistical argument, or remove the candidate claim. The same issue may affect the injection-based sensitivity estimates, since many detected injections likely have SNR below the tuned ρ_th. I would not reject the paper outright because the method and most of the sensitivity analysis are sound and potentially valuable, but the required fixes touch the core results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my take on arXiv:2512.04516.\n\nThe paper's real contribution is the sensitivity benchmarking. They run their matched-filter/deep-learning hybrid over the official GWTC-3 BBH injection set and show that at source-frame chirp masses above roughly 25 solar masses they are comparable to the LVK pipelines, and that they add a meaningful number of unique detections to the combined search. The missed-candidate analysis is honest and useful, and they release posterior samples and Asimov configs for their one new event. That part is solid and deserves a careful read by anyone building ML search pipelines.\n\nThe problem is the new candidate, GW190929 091722, and its p_astro. As written, Eq. (3) defines the foreground density as a truncated power law with support only above rho_th = 8.5 for O3a. The candidate is in O3a and has network SNR 6.7. Plug that into the model and you get p1 = 0, hence p_astro = 0, not 0.63. The paper reports 0.63 and only says there is \"a reasonable amount of uncertainty.\" Unless the implementation actually uses the ranking statistic (the deep-learning output) in place of the SNR in Eq. (3)—which would contradict the text—the reported p_astro is an off-support extrapolation. The authors need to clarify what variable enters Eq. (3) and why the threshold does not apply to the candidate. This is not a minor calibration quibble; it is the evidential basis for the \"previously unreported promising new candidate\" claim.\n\nThe rest of the paper holds up reasonably well. The p_astro tuning has free parameters, but that's standard practice and they use an external injection set for the sensitivity comparison. The high-mass sensitivity comparison and the unique-detection analysis are the strongest parts. The lack of a full pipeline release with code and commit hash is a minor weakness.\n\nBottom line: I'd send it to peer review, not desk-reject it, but I'd tell the referee to make the p_astro issue the central question. If the authors clarify or correct it, the injection study alone is worth publishing. If they can't, the new-candidate claim should be withdrawn or heavily downgraded. I'd bring it to our reading group when discussing ML search methods, but not because the new candidate is likely real.","headline":"Solid sensitivity benchmarking for a hybrid ML search pipeline; the one new candidate's p_astro is internally inconsistent with the paper's own foreground model and needs clarification before the claim can be taken seriously.","tokens_in":34299,"tokens_out":5068,"would_cite":true,"duration_ms":48590,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["83C35","68T07"],"pacs":["04.30.-w","04.80.Nn"],"model":"deepseek-v4-flash","headline":"A hybrid matched-filter/deep-learning search finds a previously unreported high-mass binary black hole merger candidate with 63% probability of astrophysical origin.","keywords":["gravitational waves","binary black hole mergers","deep learning","matched filtering","signal-to-noise ratio","intermediate-mass black hole","candidate event","O3 observing run"],"falsifier":"A dedicated intermediate-mass black hole search of the same data that requires both detectors to see a coincident peak above an SNR of 4 and finds nothing at GPS time 1253783860.8 would settle that the candidate is not a real merger; the paper notes that an existing dedicated IMBH search found no coincident candidate.","tokens_in":33198,"feed_emoji":"🕳️","tokens_out":5582,"duration_ms":51955,"temperature":0.7,"pith_summary":"The paper argues that a hybrid search—matched filtering to generate signal-to-noise ratio time series, then a deep-learning model to score them—can find binary black hole mergers that standard analytic searches miss, while matching their sensitivity for high-mass systems. Applied to the third observing run of the Hanford and Livingston detectors, the pipeline recovers 31 previously reported candidates and also flags one event seen only by a search including higher-order harmonics and one completely new candidate. The new candidate, GW190929 091722, is estimated to have a total mass of about 171 solar masses with a 44.5% chance that the primary black hole exceeds 120 solar masses, placing it in the intermediate-mass black hole range. The authors present parameter estimation and data-quality checks that find no instrumental or glitch cause, so if the candidate is real it would be a previously missed massive merger.","feed_headline":"New deep-learning pipeline flags an unreported 170-solar-mass merger","feed_subtitle":"The hybrid pipeline also recovers 31 known gravitational-wave events and matches standard searches above 25 solar masses.","key_machinery":"The pipeline's ranking statistic is the maximum, over the ten highest-SNR templates in a given second, of the average output of a convolutional residual network that sees one-second SNR time series from both detectors; sixteen staggered 1/16-second views are averaged to make the model insensitive to trigger timing. Significance is assigned through a false-alarm rate estimated from time-shifted backgrounds and a probability of astrophysical origin (p_astro), computed by comparing foreground and background trigger densities with a foreground model assuming a uniform-in-volume source distribution. The deep-learning score replaces the analytical SNR ranking, and the paper shows this score—not th","core_discovery":"The central claim is that a ranking statistic built from a deep-learning model's average prediction over the signal-to-noise ratio time series—rather than from analytical SNR thresholds—detects binary black hole mergers with sensitivity comparable to existing pipelines for source-frame chirp masses above about 25 solar masses, and identifies a distinct population of events those pipelines miss. In the offline search, 31 of 33 candidates above p_astro ≥ 0.5 and false-alarm rate < 2/day match previously reported events; one (GW190605 025957) was previously reported only by a search including higher-order harmonics; and one (GW190929 091722) is new. Parameter estimation for the new candidate gi","pith_inferences":["If confirmed by a targeted intermediate-mass black hole search, the new candidate would suggest that current catalogs are missing a population of very massive mergers, with implications for the black hole mass gap and formation channels.","The fact that this search's unique detections do not overlap with other pipelines' suggests that ensemble searches with diverse ranking statistics may be more sensitive than any single pipeline—an idea the paper gestures at but does not develop.","Because the new candidate has a Hanford peak below the usual single-detector threshold, the model may be exploiting features other than a coincident SNR peak; testing the model on single-detector masks would clarify which features drive the detection.","A targeted re-search of O3 data with an IMBH template bank and a p_astro model tuned at higher masses would provide a sharper test of whether p_astro = 0.63 is reliable."],"forward_implications":["Including this search in a combined detection scenario increases the sensitive volume, notably for systems with component masses at or above 20 solar masses.","The new candidate, if astrophysical, would be among the most massive binary black hole systems known and would support an intermediate-mass black hole population.","The method's sensitivity is stable over months, so retraining between observing runs rather than weekly is sufficient.","Independent deep-learning searches can recover events missed by matched-filter-only pipelines, strengthening multi-pipeline detection strategies.","Expanding training to lower masses and other source types is expected to close the sensitivity gap below 25 solar masses, given earlier work on neutron-star binaries."],"fun_headline_variants":["AI pipeline discovers a new 170-solar-mass merger candidate in LIGO data","Deep learning finds a new black hole merger in LIGO data","Hybrid AI pipeline uncovers a new gravitational wave candidate","AI-assisted search finds a possible 170-solar-mass black hole merger","AI search uncovers an unreported black hole merger in LIGO data"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The significance estimate for the new candidate assumes that a model tuned on simulated signals with stronger, two-detector peaks still works correctly for a weaker signal with only one clear detector peak and a mass near the edge of its training data.","fun_headline_variants_meta":{"raw":{"variants":["AI pipeline discovers a new 170-solar-mass merger candidate in LIGO data","Deep learning finds a new black hole merger in LIGO data","Hybrid AI pipeline uncovers a new gravitational wave candidate","AI-assisted search finds a possible 170-solar-mass black hole merger","AI search uncovers an unreported black hole merger in LIGO data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001159,"raw_usage":{"total_tokens":4665,"prompt_tokens":798,"completion_tokens":3867,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":3771}},"tokens_in":542,"tokens_out":3867,"duration_ms":24482,"temperature":1.0,"reasoning_tokens":3771,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:33:25.326367+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A dedicated intermediate-mass black hole search of the same data that requires both detectors to see a coincident peak above an SNR of 4 and finds nothing at GPS time 1253783860.8 would settle that the candidate is not a real merger; the paper notes that an existing dedicated IMBH search found no coincident candidate.","supporting_citations":[],"review_version":1}