{"id":"bce9fe8d-b94b-4cc9-8319-b9f4e35b9f71","arxiv_id":"2509.06901","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Normalizing flow based non-Gaussian consistency tests in a compressed detector basis select GW170104-GW170814 as the only promising lensed pair in GWTC-3.","lead":"Gravitational waves passing near a massive object can arrive at Earth as repeated copies, and this paper builds machine-learning tools to find such lensed pairs in public catalogs. The new normalizing-flow tests recover the one previously known GWTC-3 candidate, suggesting a cheaper route for future lensing searches.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claim that GW170104-GW170814 is the only significant candidate rests on a likelihood-ratio significance approximation the authors show can overestimate rejection confidence; excluding GW190828-GW200129 at 4.5σ is not robust.","rationale":"The reader's verdict is CONDITIONAL, and the reader's rationale mentions the likelihood-ratio approximation as a possible source of overestimated rejection confidence. My concern sharpens this: it is not merely a possible flaw but the load-bearing element for the paper's headline result. The NV calibration already shows the LR statistic is the least well-calibrated estimator under the null (KS significance 5%), and the authors' own analytic derivation in App. A.2 states the approximation can overestimate significance. For the critical pair GW190828-GW200129, the parameter-shift test alone would classify it as ambiguous, and only the LR test excludes it; if the LR significance is inflated, the claim that GW170104-GW170814 is the only significant candidate fails. This is an internal-consistency concern about a specific approximation, not a disagreement with external consensus, and it is testable using the paper's own formalism. The reader's weakest_assumption focused more on flow-miscalibration transferring to GWTC-3; my concern is orthogonal but complementary, so 'partial' agreement is appropriate. The verdict should remain CONDITIONAL because the concern is concrete but not yet demonstrated to invalidate the result; the proposed test would settle whether the 4.5σ exclusion is robust.","tokens_in":32612,"tokens_out":7022,"duration_ms":79785,"concrete_test":"Recompute the p-value for the GW190828-GW200129 QL using the exact quadratic-form distribution derived in App. A.2 (weighted sum of χ² variables with eigenvalues of AΣ estimated from the normalizing-flow covariances) instead of the χ²(Neff) approximation, and compare with an empirical null distribution of QL built from the noise-varying catalog pairs. If the resulting significance drops below ~3σ, the exclusion of this pair — and therefore the 'only significant candidate' claim — is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that GW170104-GW170814 is the only significant lensing candidate depends on the likelihood-ratio test rejecting GW190828-GW200129 at 4.5σ, since its parameter-shift significance is only ~2.5σ (Sec. VIII). That rejection uses the QL statistic with significance assigned via the χ²(Neff) approximation (Eq. 19). In App. A.2 the authors prove this approximation has variance ≥ 2Neff and therefore 'can in principle overestimate the significance of rejection' — and they note it is exact only when parameters are fully constrained or fully prior-dominated. For GW190828-GW200129 the LR rejection is explicitly attributed to multi-modal localization parameters, the regime where the GLM assumptions fail most severely. Moreover, the only empirical null-hypothesis calibration of the LR statistic, the noise-varying catalog, gives a KS significance of just 5% (Fig. 5b) — the worst of the three estimators, indicating miscalibration even in the near-Gaussian, high-SNR regime. Since the paper's uniqueness claim rests on excluding this pair (and, to a lesser extent, GW190512-GW190925), an overestimated LR significance could leave additional viable candidates, directly undermining the strongest claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a machine-learning workflow for identifying strongly lensed gravitational-wave candidates from event parameter posteriors. It extends the detector basis of Ezquiaga et al. with normalizing-flow density estimates, defines a KL-based information-content statistic, a non-Gaussian parameter-shift statistic, and a likelihood-ratio statistic whose null distribution is approximated as chi^2(Neff). Validation is carried out on two simulated catalogs: noise-varying realizations of a high-SNR event and 24 injections containing eight lensed pairs, followed by application to the 86 BBH events of GWTC-3. The central catalog claim is that only GW170104-GW170814 passes all criteria, while GW190512-GW190925 and GW190828-GW200129 are discarded mainly through the likelihood-ratio test. The normalizing-flow code is released as part of the tensiometer package.","tokens_in":32939,"tokens_out":6625,"duration_ms":66937,"significance":"If the calibration holds, this is a valuable methodological step: it replaces kernel-density and evidence-ratio bottlenecks with tractable flow-based statistics, demonstrates that the detector basis is more informative than the overlap basis, publicly releases the code, and independently recovers a previously discussed lensing candidate. The validation is honest about false negatives and about the limitations of the chi^2 approximation. However, the strongest catalog conclusion depends on a likelihood-ratio significance that the authors themselves show can overestimate rejection confidence, and whose null calibration is the weakest of the three estimators they consider. The manuscript is therefore a promising contribution whose headline GWTC-3 claim needs either empirical recalibration or careful softening.","major_comments":[{"comment":"The rejection of GW190828-GW200129 at 4.5 sigma is load-bearing for the claim that GW170104-GW170814 is the only significant candidate. That rejection uses the likelihood-ratio statistic QL with significance assigned via the chi^2(Neff) approximation. Appendix A.2 proves that Var(QL) >= 2 Neff and states that the approximation 'can in principle overestimate the significance of rejection', being exact only when parameters are fully data- or prior-dominated. The text attributes the 4.5 sigma rejection to multimodal localization parameters, precisely the regime where the GLM assumptions underlying the approximation fail. Because the NV null calibration for QL is also the weakest of the three estimators (KS significance 5%, Fig. 5b), the 4.5 sigma number cannot be taken at face value. Please provide an empirical null calibration of QL on realistic low-SNR and multimodal pairs, or explicitly","section":"Sec. VIII and App. A.2, Eq. (19)"},{"comment":"The likelihood-ratio p-values on the noise-varying catalog have a KS significance of only 5%. This is above the 1% threshold customarily used to flag problems, but it is marginal, and the NV events are high-SNR and nearly Gaussian. The GWTC-3 pairs are lower SNR and markedly more non-Gaussian, as the paper itself documents in Appendix B. No null calibration is provided in that regime. The paper should either calibrate QL on simulated unlensed pairs with realistic non-Gaussianity or attach a systematic uncertainty to the reported n_sigma values for the likelihood-ratio test, particularly for the pairs that determine the final candidate list.","section":"Sec. VII A, Fig. 5(b)"},{"comment":"Only 90% of trained flows pass the KS test at the 5% level, compared with the 95% expected for perfect modeling; the text calls this 'very close to the ideal case'. For the specific events that drive the catalog conclusion (GW190828, GW200129, GW190512, GW190925), an inadequately trained flow could change both the parameter-shift and likelihood-ratio values. Flow variance is shown as error bars in the injection studies but is not propagated into the reported GWTC-3 significance values or the final selection decision. Please state which GWTC-3 flows, if any, fail the KS test and quantify how the final candidate list changes under flow-model uncertainty.","section":"Sec. III and Sec. VIII"}],"minor_comments":[{"comment":"Typo in the text: 'the same gravitational wave eventh(t)' should read 'event h(t)'.","section":"Sec. II, Eq. (1)"},{"comment":"The architecture description '2 log2 Nparams spline flows' should define Nparams explicitly and indicate whether this is the number of parameters after pre-processing.","section":"Sec. III"},{"comment":"The caption describes panel (a) as the parameter-shift estimator in two bases, but the figure also compares methods in panel (b). Please expand the caption so the two panels are independently described.","section":"Fig. 5"},{"comment":"The sentence 'Wilks' theorem ensures that ... Delta theta_f = Delta theta_MAP' is misleading: for a fixed Delta theta_f this is an assumption about the null hypothesis, not a consequence of Wilks' theorem. Please rephrase.","section":"App. A.2"},{"comment":"The detector-network name is rendered as 'L VK' with an extra space in several places; likely a typographical artifact.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript relies substantially on the authors' own tensiometer/phazap framework, but the validation uses external simulated catalogs and previously reported candidates, so I do not see a circularity problem. The main issue is fixable: the catalog-level claim should be supported by an empirical calibration of the likelihood-ratio statistic in the non-Gaussian regime, or restated as conditional on the chi^2 approximation. The paper fits the journal's scope and the code release is a concrete strength."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Steph, quick take on 2509.06901. This is a genuine methodological step forward for lensed-GW searches: it replaces KDE-based overlap calculations with normalizing flows (built on the authors' tensiometer code), adds non-Gaussian parameter-shift and likelihood-ratio statistics in the detector basis, and uses a KL-divergence information plane for candidate selection. The validation on simulated catalogs is honest and reasonably thorough: p-values from the noise-varying catalog are well-calibrated for the parameter-shift statistic (KS 73%), and the injection catalog recovers six of eight lensed pairs; the two missed pairs trace to a noise-affected localization in a shared event, which is a clear explanation rather than a hand-wave. The flow-model KS tests (90% at 5% level) are close to ideal, and the code is released. That is real evidence of careful work.\n\nThe soft spots are real but not fatal. The likelihood-ratio statistic's significance is assigned via a chi^2(Neff) approximation the authors themselves show can overestimate rejection confidence (App. A.2), and the only empirical null-hypothesis calibration of the LR, the NV catalog, gives a KS significance of just 5%—the worst of the three estimators. This matters for the GWTC-3 conclusion: the rejection of GW190828-GW200129 at 4.5σ, which is what leaves GW170104-GW170814 as the only significant pair, rests mostly on the LR statistic; the parameter-shift significance for that pair is only ~2.5σ. If the LR is miscalibrated in the multimodal, lower-SNR regime of real events, that pair could remain a viable candidate, and the 'only one' claim would need weakening. The authors do flag this caveat, which is to their credit, but it means the exclusion conclusions should be read as provisional.\n\nThe decision boundary in the n_sigma-D_KL plane is also hand-tuned, as the paper acknowledges. Who benefits: people building practical lensing searches for the next catalogs, and anyone who wants a cheaper screening tool than evidence-ratio methods. It does not resolve any long-open physics question—the one surviving candidate was already known and astrophysically disfavored—so near-term impact is methodological. The paper should go to peer review; a good referee will push for either a better-calibrated LR significance or a more conservative treatment of the GW190828-GW200129 exclusion. I'd bring it to our reading group, and I'd cite it if I were working on lensing searches.","headline":"Solid, well-validated methodology for scalable lensed-GW searches; the GWTC-3 conclusion that only one pair survives is plausible but the LR-based exclusion of one candidate is not robust.","tokens_in":33411,"tokens_out":2556,"would_cite":true,"duration_ms":28365,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Machine-learning consistency tests leave one strong-lensing candidate in GWTC-3: the previously known pair GW170104-GW170814.","keywords":["gravitational lensing","gravitational waves","strong lensing","normalizing flows","posterior consistency","GWTC-3","parameter estimation","lensed event search"],"falsifier":"Take the GWTC-3 events that drive key conclusions, retrain their normalizing flows with independent seeds, and compare the flow-based parameter-shift probability and likelihood ratio at zero shift against direct sample-based calculations for those same pairs; if the flows fail the KS test on those events or the shift probabilities move by more than about one sigma, the reported candidate ranking is not robust.","tokens_in":32530,"feed_emoji":"🔭","tokens_out":3400,"duration_ms":36656,"temperature":0.7,"pith_summary":"The paper develops a set of machine-learning statistical tests to find pairs of gravitational-wave events that could be repeated images of the same merger caused by strong gravitational lensing. It uses normalizing flows to approximate the full 15-dimensional posterior of each event, then computes three consistency statistics in a compressed six-parameter detector basis: a non-Gaussian parameter-shift probability, a likelihood-ratio at maximum posterior, and the information content (KL divergence) of the pair. The method is validated on simulated catalogs with known lensing status and then applied to the GWTC-3 catalog of 86 binary black holes, where it finds exactly one pair, GW170104-GW170814, that passes all tests. If the calibration transfers to real data, the approach offers a computationally feasible, calibration-free way to rank lensing candidates as catalogs grow, without the costly joint parameter estimation used in earlier searches.","feed_headline":"One lensing candidate survives GWTC-3 search","feed_subtitle":"Normalizing flows rank pairs by parameter shift and information content; only GW170104-GW170814 passes all tests.","key_machinery":"The detector parameter basis: six parameters consisting of per-detector phases, inter-detector time delays, and a cycle count that replaces chirp mass. Normalizing flows, trained on each event's posterior and prior, make it computationally feasible to evaluate the KL divergence, the parameter-shift probability, and the likelihood ratio; the likelihood-ratio statistic is assigned a chi-squared distribution with the effective number of constrained parameters, extending Wilks' theorem to the maximum-posterior point.","core_discovery":"The central claim is that strong-lensing consistency between two gravitational-wave events can be tested without joint parameter estimation or large simulations of lensed and unlensed events. Instead, normalizing-flow approximations of each event's posterior are used to evaluate three non-Gaussian statistics directly: the probability that the parameter difference exceeds zero shift, a likelihood ratio comparing zero shift to the maximum-posterior shift, and the Kullback-Leibler divergence between prior and posterior as a measure of information content. Applied to GWTC-3, the method identifies GW170104-GW170814 as the only significant lensing candidate, a pair previously flagged by evidence-r","pith_inferences":["A natural testable extension is to calibrate the n_sigma-D_KL decision boundary on injection catalogs that reproduce GWTC-3's lower signal-to-noise ratios and strongly multimodal posteriors, since the current boundary is deduced from a small simulated sample.","The reported significance of individual pairs, especially rejections like GW190828-GW200129, depends on the accuracy of flows for those specific events; a flow that fails the KS test could shift a pair's n_sigma by enough to change the candidate list.","The likelihood-ratio statistic's chi-squared approximation is conservative in the sense that its true variance exceeds that of the approximating distribution, so reported rejection significances may be overestimated when parameters are only partially constrained.","The same machinery, with the detector basis and information plane, could be applied to future catalogs with thousands of events, where the dominant bottleneck shifts from computational cost to the reliability of flow calibration across heterogeneous event morphologies."],"forward_implications":["Pairwise strong-lensing searches can be run at scale: a fast Gaussian preselection trims the O(n^2) pair set, and full non-Gaussian statistics are then evaluated with flow evaluations rather than joint parameter estimation.","The non-Gaussian parameter shift in the detector basis is better calibrated and more informative than the overlap basis, and it avoids false rejections caused by mass-spin degeneracy and multimodal localization.","The information-content plane (parameter-shift significance vs. KL divergence) provides a practical ranking tool for candidate selection that separates promising, ambiguous, and unlikely lensing pairs.","The only surviving candidate in GWTC-3, GW170104-GW170814, matches a previously identified candidate, providing independent validation of the method against more costly techniques.","The method is directly extendable to other event classes and to sub-threshold searches, with the caveat that the decision boundary in the information-content plane still needs calibration on larger simulated catalogs."],"supporting_citations":[{"why":"Introduces the detector parameter basis and the Gaussian-shift preselection approach (phazap) that the paper extends and uses as the fast preselection step.","marker":"[36]"},{"why":"Supplies the normalizing-flow framework and the parameter-shift methodology for non-Gaussian posterior comparisons that the paper adapts to gravitational-wave lensing searches.","marker":"[35]"},{"why":"Defines the overlap basis and previously identified GW170104-GW170814 as a promising lensing candidate, providing the benchmark comparison for the new method.","marker":"[14]"},{"why":"Represent the evidence-ratio approach to lensing consistency that the new method avoids, serving as the reference techniques the paper contrasts with.","marker":"[24, 25]"},{"why":"Provides the Gaussian Linear Model formalism used to define the effective number of constrained parameters and to derive the KL-divergence/volume relation and the likelihood-ratio distribution.","marker":"[46]"},{"why":"Gives Wilks' theorem, which the paper extends to the maximum-posterior point to justify the chi-squared approximation for the likelihood-ratio statistic.","marker":"[54]"},{"why":"Is the GWTC-3 catalog dataset used for the real-data application of the method.","marker":"[57]"},{"why":"Defines the overlap basis parameters used as the comparison basis throughout the analysis.","marker":"[39]"},{"why":"Is the parameter-estimation pipeline used to produce the posterior samples on which all the flows and statistics are built.","marker":"[55]"}],"fun_headline_variants":["Machine learning flags one lensed GW pair in GWTC-3","Normalizing flows find lone lensing candidate in GWTC-3","Fast ML search for lensed gravitational waves yields one candidate","GW lensing search: one pair passes all ML tests","Only one GW event pair survives lensing ML screen"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The calibration established on high-signal-to-noise, nearly Gaussian simulated events transfers to the real GWTC-3 events, which are lower signal-to-noise and significantly more non-Gaussian, and every normalizing flow is accurate enough that the reported significances—including rejections such as GW190828-GW200129—are not corrupted by flow error.","fun_headline_variants_meta":{"raw":{"variants":["Machine learning flags one lensed GW pair in GWTC-3","Normalizing flows find lone lensing candidate in GWTC-3","Fast ML search for lensed gravitational waves yields one candidate","GW lensing search: one pair passes all ML tests","Only one GW event pair survives lensing ML screen"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000725,"raw_usage":{"total_tokens":3056,"prompt_tokens":680,"completion_tokens":2376,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":2292}},"tokens_in":424,"tokens_out":2376,"duration_ms":18075,"temperature":1.0,"reasoning_tokens":2292,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T22:51:16.018339+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the GWTC-3 events that drive key conclusions, retrain their normalizing flows with independent seeds, and compare the flow-based parameter-shift probability and likelihood ratio at zero shift against direct sample-based calculations for those same pairs; if the flows fail the KS test on those events or the shift probabilities move by more than about one sigma, the reported candidate ranking is not robust.","supporting_citations":[{"cited_title":"lenscat: a Public and Community-Contributed Catalog of Known Strong Gravitational Lenses","cited_arxiv_id":"2406.04398","evidence_quote":"Introduces the detector parameter basis and the Gaussian-shift preselection approach (phazap) that the paper extends and uses as the fast preselection step."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the overlap basis parameters used as the comparison basis throughout the analysis."}],"review_version":1}