{"id":"7b92b21c-19f5-4c0b-ac59-08502bdca992","arxiv_id":"2411.13058","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Matched-filtering searches recover strongly lensed gravitational waves at rates far below optimal-SNR predictions, with detection efficiency dropping to under one percent.","lead":"The paper simulates lensed gravitational-wave signals in real detector noise and shows that matched-filtering search pipelines recover them far less often than simpler SNR-based estimates suggest. It matters because current constraints on exotic signals, including compact dark matter, may be overly optimistic if based on those simpler estimates.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 90%-to-<1% detection drop likely depends on the reduced 14,606-template bank; a denser bank cannot worsen the best-match template, so the paper's defense of the small bank is unsupported.","rationale":"The paper's central quantitative claim is the drop in detection efficiency from ~90% (unlensed) to <1% (lensed) for the strongest wave-optics distortion. That number comes from a GstLAL run with a deliberately small template bank (14,606 templates, Sec. V.A). The paper's only justification for the small bank is the assertion in Sec. VI.A that a larger bank would only broaden the recovered parameter space and worsen the xi^2 time series. This is not a trivial point: in matched filtering, the maximum SNR over a denser bank is at least as large as over a sparser bank that it contains, and the template that maximizes the ranking statistic for a given injection should, if anything, match the lensed signal better. A better match reduces the autocorrelation mismatch in Eq. (7), lowering xi^2 and improving the FAR. The small bank also changes the noise-trigger distribution used to calibrate the false-alarm rate in Eq. (13); the paper does not quantify this effect. Thus the quantitative efficiency could be a bank-size artifact in either direction, and the paper's direction claim (worse with larger bank) is unsupported. A concrete rerun with a denser bank, at least for the headline set (set 4), would settle this. If the result survives, the qualitative conclusion is robust; if not, the headline numbers are not. The reader's conditional verdict already flags the bank-size issue; our concern is a specific, testable version of that caveat, so the verdict should remain conditional/unchanged.","tokens_in":21943,"tokens_out":7194,"duration_ms":68394,"concrete_test":"Re-run injection campaign 1 set 4 (y = 0.01, MLz = 1e5 Msun, dL = 125 Mpc, 19432 injections) with a substantially denser template bank, ideally the full O3 LVK bank or a bank that contains the injected 30-30 Msun template plus many nearby templates (e.g., 10x density in m1-m2-chi_eff). Compare the number of found injections (FAR < 3.85e-7 Hz), the recovered-parameter distributions (Fig. 3), and the xi^2 distribution (Fig. 4). If the found count stays near 1, the reduced-bank concern is settled; if it rises by more than a factor of a few (e.g., >10 found), the headline 90%-to-<1% drop and the paper's bank-size justification are not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. V.A uses a reduced template bank of 14,606 templates (m1 = 10-90 Msun, m2 = 10-40 Msun, aligned spins -0.99 to 0.99) to search for 30-30 Msun spinless injections. The headline result (Sec. VII, Table I) is that for y = 0.01, MLz = 1e5 Msun at 125 Mpc, only 1/19432 injections are found (FAR < 3.85e-7 Hz), versus 17020/19432 unlensed. The paper justifies the small bank in Sec. VI.A: a larger bank 'permits greater scope for inaccurate recovery of source parameters' and 'leads to a more inaccurate SNR time series for xi^2.' This is not demonstrated and is likely backwards for signal injections: adding templates cannot decrease the maximum match (Eq. 5) or the recovered SNR, and a better-matched template should reduce the autocorrelation mismatch entering xi^2 (Eq. 7). The small bank also changes the noise-trigger distribution that calibrates C(ln L|noise) in Eqs. (10)-(13), so the FAR threshold is not the same quantity as in the full LVK search. Consequently, the quantitative efficiency drop (90% to <1%) is not robust to a basic pipeline configuration choice. The qualitative conclusion that optimal-SNR proxies overestimate detectability in the wave-optics regime may survive, but the specific numbers and the statement that current dark-matter constraints are overoptimistic rest on an untested and questionable bank-size assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the first injection campaign using the GstLAL matched-filtering pipeline to study the detectability of gravitational waves lensed by point masses in the wave-optics regime. It injects 30-30 solar-mass spinless BBH signals, both lensed and unlensed, into one week of O3 Hanford-Livingston data, searches with a reduced 14,606-template bank, and counts injections with FAR below 1/30 days. The central result is that strongly lensed signals (y=0.01, M_Lz=10^5 solar masses) at 125 Mpc are found in only 1/19432 injections versus 17020/19432 for the unlensed case, despite higher optimal SNR, because of SNR loss and elevated signal-consistency test values. The paper argues that optimal-SNR proxies overestimate detectability and that current compact dark matter constraints based on such proxies are overoptimistic.","tokens_in":22203,"tokens_out":6562,"duration_ms":69209,"significance":"If the quantitative result is robust, this is an important correction to detectability estimates used in lensing-rate calculations and dark matter constraints, and it provides concrete motivation for lensed template banks and alternative search pipelines. The paper is careful in separating the roles of matched-filter SNR and the signal-consistency statistic, and it demonstrates that higher optimal SNR does not automatically imply higher detectability. The main limitation is that the headline numbers are obtained for one pipeline configuration, one reduced template bank, one week of data, and one source/lens model; the generalization to full LVK searches is asserted rather than demonstrated. The qualitative conclusion is likely to survive, but the specific efficiency numbers and the claim about overoptimistic constraints need further support.","major_comments":[{"comment":"The reduced 14,606-template bank is load-bearing for the headline result, but the paper's defense of it is not convincing. Adding templates to a bank cannot reduce the maximum match defined in Eq. (5) or the quality of the best available template, and a better-matched template should reduce the autocorrelation mismatch entering Eq. (7); the argument that a larger bank 'permits greater scope for inaccurate recovery' addresses noise-trigger misrecovery, not the recovery of injected signals. In addition, the FAR in Eqs. (10)-(13) is calibrated on the noise-trigger distribution of this specific bank, so the same trigger would receive a different FAR with a different bank; the found counts in Table I are therefore not robust until this is tested. Please provide a control with a denser bank around the injection parameters and quantify the change in C(ln L|noise).","section":"Sec. V.A and Sec. VI.A, Eqs. (5), (7), (10)-(13), Table I"},{"comment":"The analysis uses one week of O3 Hanford-Livingston data and a single noise realization. The FAR threshold of 3.85e-7 Hz corresponds to one event per 30 days; with one week of data, the empirical C(ln L|noise) distribution has limited statistics near the threshold, and the found counts (17020, 12878, 6217, and the lensed counts 1, 8, 13) carry unquantified uncertainty. The central trend is plausible, but the specific detection-efficiency numbers and the Sec. VII conclusion that current constraints are overoptimistic extrapolate beyond the tested configuration. Please provide error estimates on the found counts or analyze additional data stretches, and qualify the constraint statement accordingly.","section":"Sec. V.A and Table I"},{"comment":"The statement that current compact dark matter constraints are 'overoptimistic' is a quantitative population-level claim that does not follow directly from the single-source, single-lens-model injection study. Computing the impact on constraints requires a selection function over the source and lens populations, the full production template bank, and the pipeline's actual FAR calibration. The paper itself notes that a full injection campaign is required; as written, the conclusion outruns the evidence. Please either reframe this as a direction for future work or provide the necessary population-level calculation.","section":"Sec. VII and Conclusion"}],"minor_comments":[{"comment":"The text states that 'only 1 out of 19421 injections is successfully detected,' but Table I lists 19432 injections; please correct the number for consistency.","section":"Sec. VI.A"},{"comment":"There are typos 'inacccurate modelling' and 'mathced filtering' that should be corrected.","section":"Sec. II.B and Sec. IV"},{"comment":"The sentence preceding Eq. (13) mentions N as the total number of triggers, but N does not appear in the equation; please remove or clarify the definition.","section":"Sec. III.B, Eq. (13)"},{"comment":"The right-panel caption has an unmatched parenthesis in 'ρL,opt) and unlensed (ρUL,opt'; please fix the formatting.","section":"Fig. 1 caption"},{"comment":"The terms 'matched-filtering' and 'match-filtered' are used inconsistently; please standardize the terminology.","section":"Abstract and throughout"},{"comment":"The FAR color scale and threshold tick are difficult to read in printed form; consider adding a clear threshold line or annotation in each panel.","section":"Figs. 4 and 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a gravitational-wave journal and reports a useful first step. My main concern is that the headline efficiency drop is presented as a property of matched-filtering searches generally, while the evidence is for one reduced bank and one week of data; the authors should either add robustness checks or soften the generalization. The paper's own acknowledgement that a full injection campaign is required should be reflected in the conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The first thing to know: this is the first injection-campaign study of lensed CBC detectability inside a real matched-filtering pipeline. The main qualitative claim—that optimal-SNR proxies badly overestimate how often strongly lensed signals are found, because they ignore the signal-consistency test—is credible and likely right. The paper does real work: standard point-mass diffraction kernels, the actual GstLAL ranking statistic, real O3 noise, no fitted parameters making the result. The 90%-to-<1% drop in detection efficiency is not built in by construction; it falls out of the pipeline behavior. That is a genuine contribution.\n\nThe soft spots are about the quantitative reach. The headline number comes from a reduced 14,606-template bank, one week of Hanford-Livingston data, and a single 30-30 M_sun spinless source. The paper's defense of the small bank is weak: it claims a larger bank 'permits greater scope for inaccurate recovery of source parameters,' but adding templates cannot decrease the best-match SNR, and a better-matched template should reduce the xisquared penalty, not increase it. The FAR threshold is calibrated on the noise-trigger distribution of this small bank, so the absolute efficiencies are not directly the LVK search's numbers. The paper also leaves the xi-squared integration window delta-t unreported, which is a minor but real omission.\n\nI would not, however, accept the stress-test's stronger version that the bank choice could erase the effect. The xi-squared penalty comes from the mismatch between the lensed signal's SNR time series and the unlensed template's autocorrelation; a denser bank changes the boundary of the best-matching template but does not remove that mismatch for heavily distorted signals. The qualitative conclusion that optimal SNR proxies overestimate detectability in the wave-optics regime should survive a bank-size change. The exact 90%-to-<1% numbers, and the implied statement that current dark-matter constraints are overoptimistic, are extrapolations the paper does not test.\n\nWho this is for: people computing lensing detection rates or setting compact dark matter constraints, and pipeline developers who care about template completeness. It deserves a serious referee. The referee should ask for a denser-bank robustness check or an explicit demonstration that the FAR calibration is insensitive to bank size; a comparable run with a few times more templates would settle it. I would cite this paper for the qualitative point, flagging the numbers as provisional.","headline":"First real pipeline-level demonstration that lensed GWs can be badly missed by matched-filtering searches; the qualitative result is credible, but the headline numbers are tied to a reduced template bank and should not be taken as generic LVK efficiencies.","tokens_in":22780,"tokens_out":2582,"would_cite":true,"duration_ms":25577,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Matched-filtering searches catch almost none of the most strongly lensed gravitational-wave signals, despite their high optimal signal-to-noise ratio.","keywords":["gravitational waves","gravitational lensing","wave optics","matched-filtering search","detectability","injection campaign","signal-consistency test","compact dark matter"],"falsifier":"Run the same $y=0.01$, $M_{Lz}=10^5\\,M_\\odot$ injection set at 125 Mpc through the full production template bank and through an independent matched-filtering pipeline; if the detection fraction stays near the unlensed level rather than falling below 1%, the claimed pipeline rejection does not generalize. A complementary check is to search the same week of data with an unmodeled burst algorithm: a strongly magnified lensed event that the template search misses but the burst search finds would directly confirm that the loss is caused by template mismatch rather than by the signal being absent.","tokens_in":21701,"feed_emoji":"🔭","tokens_out":10848,"duration_ms":93825,"temperature":0.7,"pith_summary":"Gravitational waves magnified and distorted by compact intervening masses (wave-optics lensing) are usually assumed to be easy to detect because lensing raises their optimal signal-to-noise ratio. This paper tests that assumption by injecting simulated lensed binary-black-hole signals into real detector noise and running a standard matched-filtering search pipeline. It finds that the strongest lensing distortions (lens mass $10^5$ solar masses, impact parameter 0.01) drop the detection efficiency from about 90% for unlensed signals to under 1%, because the distorted waveforms no longer match the unlensed templates: they lose recovered SNR and score badly on the signal-consistency test, pushing their false-alarm rates up by many orders of magnitude. The result implies that current searches are systematically missing the most highly lensed events, and that constraints on compact dark matter derived from optimal-SNR estimates are overoptimistic.","feed_headline":"Strongly lensed wave signals: detection drops from 90% to under 1%","feed_subtitle":"Lensing distorts waveforms so much that standard template searches reject the loudest magnified events.","key_machinery":"The machinery is the matched-filtering detection statistic itself: the recovered signal-to-noise ratio $\\rho$, the autocorrelation-based signal-consistency value $\\xi^2$ (a measure of how well the SNR time series of the data matches the one expected from the best template), and the likelihood-ratio ranking that turns $\\rho$, $\\xi^2$, and detector information into a false-alarm rate. Against this sits the point-mass lens amplification factor $F(w,y)$, which imprints frequency-dependent beating on the waveform when the dimensionless frequency $w \\sim 1$; the paper shows that the lensed-to-unlensed optimal-SNR ratio is always $>1$ for this model while the match $M$ between lensed and unlensed waveforms can fall to $0.7$. The competition between those two quantities — amplification raising SNR, mismatch lowering recovered SNR and inflating $\\xi^2$ — is what determines whether an event passes the false-alarm threshold, and the injection campaigns map that competition across lens mass and source position.","core_discovery":"The paper's central claim is that optimal signal-to-noise ratio fails as a proxy for the detectability of gravitational waves lensed by compact masses, and that the failure is large and systematic. Simulated binary-black-hole signals at 125 Mpc with a point-mass lens of redshifted mass $10^5$ solar masses and impact parameter $y=0.01$ are magnified to higher optimal SNR than their unlensed counterparts, yet the search pipeline recovers only $1$ of $19{,}432$ of them at the standard false-alarm threshold, versus $17{,}020$ of $19{,}432$ unlensed injections. The loss is driven by two pipeline effects: the distorted waveform makes the pipeline recover wrong source parameters, cutting the matched-filter SNR, and the residual mismatch inflates the signal-consistency statistic, which by itself pushes most lensed events past the false-alarm threshold. The paper also finds that lowering the signal strength by moving the source further away partly masks the distortion, so $8$ and $13$ lensed injections are found at $625$ and $1250$ Mpc, respectively, the opposite of what a louder-is-easier optimal-SNR treatment predicts.","pith_inferences":["A direct test of the mechanism across other missing-physics scenarios (eccentric orbits, higher-order modes, modified gravity) would likely show the same qualitative pattern: waveform distortion lowers recovered SNR and worsens signal-consistency scores, so optimal-SNR detectability estimates are probably biased for any physics absent from the template bank.","Because the $\\xi^2$ test cannot tell waveform mismatch from noise, a ranking statistic that explicitly models lensing distortion (rather than only adding lensed templates) might recover part of the lost efficiency, an option the paper does not explore.","If the SNR-dependence holds, next-generation more-sensitive detectors could make strongly lensed events harder to find even as the absolute number of loud sources grows; rate estimates for future runs should fold in the mismatch penalty as a function of SNR."],"forward_implications":["Lensing-rate forecasts and detection-rate estimates built on optimal SNR thresholds will overpredict how many lensed events current searches can actually claim.","Existing constraints on compact dark matter that acts as wave-optics lenses are overoptimistic in the high-distortion region of lens mass and impact parameter, because they assume events are detectable when the pipeline rejects them.","A dedicated lensed template bank, especially covering the low-match, high-$\\xi^2$ region the paper maps, would be needed to recover strongly lensed events with matched filtering.","Improving detector sensitivity will not straightforwardly improve detection of strongly lensed signals: at higher SNR the distortion penalty and signal-consistency rejection become stronger, so detection efficiency actually grows as sources are placed farther away.","Follow-up lensing and parameter-estimation analyses restricted to catalog events inherit a selection bias, because the most distorted lensed events are absent from the catalogs that matched-filtering produces."],"supporting_citations":[{"why":"supplies the point-mass lens amplification factor $F(w,y)$ used to generate the lensed waveforms.","marker":"[25]"},{"why":"describes the matched-filtering search pipeline and its autocorrelation-based signal-consistency test that the study runs.","marker":"[63]"},{"why":"the optimal-SNR proxy method whose predictions for detectability the paper tests and rejects.","marker":"[72–74]"},{"why":"previous optimal-SNR-based wave-optics lensing detectability analyses that the injection results contradict.","marker":"[75, 76]"},{"why":"semianalytic sensitivity estimation that motivates why full injection campaigns are needed to capture ranking-statistic effects.","marker":"[77]"},{"why":"provides the relativistic binary-black-hole waveform model used to generate injections and templates.","marker":"[87]"},{"why":"defines the data-quality, glitch-mitigation, and template-bank conventions from the third transient catalog that the analysis follows.","marker":"[9]"},{"why":"an example compact-dark-matter lensing constraint computed with optimal SNR that the paper argues needs reassessment.","marker":"[31]"}],"fun_headline_variants":["Lensed gravitational waves: louder but harder to detect","Loudest lensed waves fail matched-filter tests","Louder isn't easier: lensed GWs evade detection","Optimal SNR is a trap for lensed GW detection","Lensing cuts detection efficiency from 90% to under 1%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central assumption is that the simplified setup — one search pipeline, a 14,606-template bank, point-mass lenses, and one week of detector noise — is representative enough that the same drop in detection efficiency would occur with the full production template bank, other pipelines, and real astrophysical event populations.","fun_headline_variants_meta":{"raw":{"variants":["Lensed gravitational waves: louder but harder to detect","Loudest lensed waves fail matched-filter tests","Louder isn't easier: lensed GWs evade detection","Optimal SNR is a trap for lensed GW detection","Lensing cuts detection efficiency from 90% to under 1%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00062,"raw_usage":{"total_tokens":2904,"prompt_tokens":1003,"completion_tokens":1901,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":1818}},"tokens_in":619,"tokens_out":1901,"duration_ms":13018,"temperature":1.0,"reasoning_tokens":1818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:51:16.318792+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same $y=0.01$, $M_{Lz}=10^5\\,M_\\odot$ injection set at 125 Mpc through the full production template bank and through an independent matched-filtering pipeline; if the detection fraction stays near the unlensed level rather than falling below 1%, the claimed pipeline rejection does not generalize. A complementary check is to search the same week of data with an unmodeled burst algorithm: a strongly magnified lensed event that the template search misses but the burst search finds would directly confirm that the loss is caused by template mismatch rather than by the signal being absent.","supporting_citations":[],"review_version":1}