{"id":"84853f46-5242-43ad-8556-31b6375c77da","arxiv_id":"2412.01278","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"PO2.0 reweights single-event posteriors to compute a lensing Bayes factor that includes population priors and lensing-biased parameters, detecting 65% of simulated galaxy-lensed BBH pairs at a pairwise false-alarm probability of about 2e-6.","lead":"The authors present PO2.0, a faster Bayesian method to tell whether two gravitational-wave signals are copies of the same event bent by a galaxy, rather than two unrelated events. It combines all measured binary parameters with population and lensing priors, and they report it can catch 65% of lensed pairs at very low false-alarm rates.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 65% efficiency and 13% detection probability are in-sample with respect to the simulated population used to build PO2.0's informative priors; Fig. 2 shows strong prior sensitivity, so real-catalog performance is not quantified.","rationale":"The derivation of Eq. 5 from Eq. 4 is clean; under the stated conditional-independence assumptions, reweighting individual posteriors with a detectable-population prior is algebraically equivalent to joint PE. I credit this as a genuine methodological contribution. The implemented statistic departs from the exact expression (delta-function arrival times, factorized sky and phase integrals, dropped PE priors), and the authors explicitly defer a head-to-head joint-PE comparison. That is an acknowledged open item rather than a hidden flaw, so I do not make it the primary objection. The load-bearing issue is that every quantitative headline—65%, 2e-6, 13%—is produced inside a single simulated world. The priors are built from the same Madau-Dickinson / power-law+peak / SIE / Collett / O4-selection pipeline (Appendix D) that generates the 1000 foreground lensed pairs and 1000 background unlensed events. Figure 2 shows the statistic is sensitive to the lensed-population prior: using HD_U in place of HD_L degrades the ROC. But the authors test only one wrong prior; they do not vary the BBH population, the lens population, or the selection function, and they do not quote systematic uncertainties on the ROC or on Eq. 15. The 13% detection probability is linearly proportional to the assumed lensing fraction u and rate R, both uncertain. The paper is transparent about assumptions in the conclusion, but the abstract's unqualified '65%' and '13%' overstate what can be concluded. This does not invalidate the method; it means the performance claims should be re-expressed as conditional on the fiducial population model, which is exactly the CONDITIONAL verdict the reader reached. A cross-population injection study is a concrete, bounded computation that would settle whether the efficiency is robust or model-locked.","tokens_in":27290,"tokens_out":15130,"duration_ms":142710,"concrete_test":"Run a cross-population validation: keep PO2.0's priors fixed to the fiducial simulation, but generate new test sets of ~1000 unlensed and ~1000 lensed pairs from a deliberately different population—e.g., a constant comoving merger rate instead of Madau-Dickinson, and/or a single SIS lens with fixed velocity dispersion instead of the Collett distribution—then recompute the ROC and detection probability. If efficiency at FAP~2e-6 stays within the ~3% shot-noise band, the in-sample concern is weak; if it drops by more than ~10-20%, the 65%/13% claims should be qualified as model-conditional.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claims—65% detection efficiency at pairwise FAP ~2e-6 and 13% catalog detection probability—are computed with foreground and background injections drawn from the same population model used to construct the informative priors: Madau-Dickinson merger rate, power-law+peak masses and spins (Abbott et al. 2023), SIE galaxy lenses from Collett (2015), and O4 selection cuts (Appendix D). This is an in-sample benchmark. The paper's own Fig. 2 shows that swapping the lensed-population prior for the unlensed detectable prior degrades the ROC, demonstrating that efficiency is prior-sensitive. However, that test varies only the lensed population; the BBH population, lens number density and velocity dispersion, and selection model remain identical for prior construction and injection. The forecast detection probability (Eq. 15) also scales linearly with the assumed lensing fraction u=1.5e-3 and detection rate R=100/yr. The abstract presents 65% and 13% without an uncertainty budget for population-model misspecification, although the conclusion acknowledges model dependence. Since PO2.0's gain over Haris et al. (2018) is largely purchased by these priors, the headline numbers are conditional on one fiducial model, not expected real-catalog performance.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PO2.0, a posterior-overlap statistic for identifying strongly lensed gravitational-wave events. The authors derive a Bayes factor that reweights individual-event posteriors with an informative lensed/unlensed population prior, claim that this statistic is mathematically equivalent to joint parameter estimation, and show that under various simplifications it reduces to existing statistics such as Beq, R_L/U, Mgal, and phazap-like measures. They implement PO2.0 using Gaussian KDEs and importance sampling on posterior samples produced by the cogwheel PE code, and benchmark it on simulated galaxy-lensed BBH pairs against an unlensed background. They report 65% detection efficiency at a pairwise false-alarm probability of ~2e-6, and a 13% probability of detecting a lensed event above 2.25σ in an 18-month O4-like catalog. They also present a reweighting method for estimating joint source/lens parameters (e.g., relative magnification, Morse phase) from individual posteriors, with caveats about numerical biases.","tokens_in":27525,"tokens_out":4977,"duration_ms":45930,"significance":"If the central claims hold, PO2.0 would be a substantial practical advance: it offers joint-PE-like sensitivity at a fraction of the computational cost, making deep background studies feasible for lensing searches. The derivation of the reweighting identity (Eq. 5, Appendix A) is correct and is a genuine contribution, and the paper gives a useful taxonomy showing how existing statistics arise as limiting cases. The authors also perform a careful validation of their PE pipeline with p-p plots, and they openly acknowledge numerical limitations in the parameter-estimation extension. The main risk is not internal inconsistency but that the headline efficiency and detection-probability numbers are in-sample with respect to the population models used to build the informative priors, and that the implemented statistic is an approximation to the exact Bayes factor despite the paper's 'equivalent to joint PE' claim. These issues are fixable but require a revised presentation and additional robustness analysis.","major_comments":[{"comment":"The exact Bayes factor in Eq. (5) is derived correctly, but the statistic actually evaluated is not Eq. (5). Appendix B drops the PE priors (Section B.1), approximates arrival-time posteriors as delta functions (Section B.4), and factorizes sky and phase terms, leading to Eq. (B11). The abstract and Section 2 claim that PO2.0 is 'mathematically equivalent to joint parameter estimation'; that equivalence is only true for the unsimplified Eq. (5) under exact marginalization. As written, the claim overstates what is tested. The paper should explicitly distinguish the exact reweighting identity from the implemented approximation, and ideally quantify the bias introduced by these approximations, e.g., by comparing PO2.0 with joint PE on a small subset of pairs, which the authors note is underway.","section":"§2.2, Appendix B, Eq. (5) vs. Eq. (B11)"},{"comment":"The 65% efficiency at FAP ~2e-6 is an in-sample benchmark: the foreground and background injections are drawn from the same population models used to construct the informative priors (Madau-Dickinson merger rate, power-law+peak masses/spins from Abbott et al. 2023, SIE lenses from Collett 2015, O4 selection cuts). The authors' own Fig. 2 shows that replacing the lensed prior with the unlensed detectable prior H_D^U substantially degrades the ROC, confirming that efficiency is prior-sensitive. Consequently, the abstract's 65% number is conditional on one fiducial population model and does not carry an uncertainty budget for model misspecification. The manuscript should include out-of-sample or stress-test variations (e.g., different mass distributions, lens velocity-dispersion functions, or optical depths) or at least a quantitative sensitivity analysis, to support claims about expected real-catalog performance.","section":"§3.1, Fig. 2, Appendix D"},{"comment":"The forecast detection probability of 13% is directly proportional to the assumed lensing fraction u = 1.5e-3 and the detection rate R = 100/yr, as shown in Eq. (15). Neither parameter has an uncertainty budget in the main text, nor is the dependence stated in the abstract, where the 13% appears as a headline number. Since the lensing fraction is especially uncertain, the paper should state the parametric dependence explicitly and provide a plausible range for the detection probability, e.g., varying u by a factor of 2, to avoid overstating the certainty of the forecast.","section":"§4, Eq. (15)"}],"minor_comments":[{"comment":"The sentence 'The probability of confidently identifying that lensed event using our method is 65% lower' appears to be a wording error; the intended meaning is that the identification probability is 65%, not 65% lower, since 0.13/0.20 = 0.65. Please rephrase.","section":"§6, Conclusion"},{"comment":"The p-p plot in Fig. 12 shows a KS p-value of 5.6e-25, indicating that the relative-magnification posteriors are systematically overestimated. The authors acknowledge this bias, but the conclusion still lists 'measure the relative magnification better than naive' as an achievement. Consider moving the magnification-measurement claim to future work or adding a caveat in the conclusion.","section":"§5.2, Fig. 12"},{"comment":"The comparison with phazap (Ezquiaga et al. 2023) is not entirely apples-to-apples, since phazap is described by its authors as a veto pipeline rather than a detection pipeline. The text notes this, but the ROC comparison may still be misinterpreted; a brief statement of the intended roles would help.","section":"§3.4"},{"comment":"The factorization that makes the Bayes factor independent of the Morse phase (by writing B_n as independent of Δφ) is described as 'empirically motivated.' Since this is a modeling assumption that could affect the reported efficiency, it would be helpful to state how sensitive the ROC curves are to this choice.","section":"Appendix B, Eq. (B4)"},{"comment":"The manuscript does not state whether the code and simulation datasets are publicly available. Providing a code repository would improve reproducibility, given that the method depends on several numerical choices (KDE bandwidths, importance sampling proposals, etc.).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central derivation is sound, but the paper currently sells the implemented statistic as exactly equivalent to joint PE when in fact several approximations are made (dropping PE priors, delta-function arrival times, factorization). The in-sample nature of the ROC is the main risk for the published claims: the 65% and 13% figures are conditional on one fiducial population model, and Fig. 2 shows strong prior sensitivity. A major revision that (i) states the approximation clearly, (ii) provides an out-of-sample/robustness test or a quantitative uncertainty budget, and (iii) tempers the abstract accordingly would bring the manuscript in line with its actual contributions. The parameter-estimation section is interesting but is currently limited by the acknowledged 10% bias in relative-magnitude estimation; this should be de-emphasized."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe thing to know: this paper presents a posterior-overlap statistic (PO2.0) that generalizes Haris et al. (2018) and Cheung et al. (2023) by folding in informative population priors with selection effects and all the lensing-biased parameters (luminosity distance, phase, arrival time). The Bayesian derivation in eq. 5 is correct and clean; the way it reduces to earlier statistics (Haris, Cheung, More & More, and even the time-delay ratio) is clearly explained. The paper is also honest about its limitations, which is refreshing.\n\nWhat is new: the systematic inclusion of the joint detectable lensed population prior and the biased parameters, and the demonstration that this improves efficiency over previous fast methods. The ROC comparisons are genuinely informative, and even the negative result with Mgal (delta-function approximation of distance/phase posteriors hurts) is a useful lesson. The paper explicitly states that the implemented code drops PE priors, uses delta-function arrival-time posteriors, and factorizes sky/phase terms, so the \"equivalent to joint PE\" claim holds only for the exact expression and has not yet been tested against actual joint PE. They say that comparison is underway.\n\nThe soft spots are real but addressable. The 65% efficiency at pairwise FAP ~2e-6 and the 13% catalog detection probability are computed using injections drawn from the same population model that defines the priors: Madau-Dickinson merger rate, power-law+peak masses and spins, SIE lenses from Collett 2015, and O4 selection cuts. Fig. 2 shows that swapping the lensed prior for the unlensed detectable prior degrades the ROC, which confirms the sensitivity. The forecast also scales linearly with the assumed lensing fraction u=1.5e-3 and rate R=100/yr. So the headline numbers are a model-dependent forecast, not a measurement of real-catalog performance. The paper acknowledges model dependence in the conclusion, but the abstract presents 65% and 13% without an uncertainty budget for population-model misspecification. No code or data are shipped, which is common but limits reproducibility.\n\nOn the question of the stress-test note: yes, it holds. The in-sample nature of the benchmark is the main weakness. But the Bayesian derivation itself is not circular; the priors come from external astrophysical simulations, not from fitting the statistic to the data. The circularity is only in the evaluation.\n\nWho this is for: anyone doing strong lensing searches in LVK, or working on fast Bayesian model selection with population priors. It deserves a serious referee. My main referee asks would be (1) to provide an out-of-sample test, e.g., injecting a population with different mass or lens models, and (2) to report the sensitivity of the 65% and 13% numbers to reasonable variations in the prior assumptions.\n\nRecommendation: send to peer review. The paper is methodologically sound, clearly written, and the limitations are mostly stated. The caveats are important but fixable, and the community needs this kind of work.","headline":"PO2.0 is a genuine advance in lensed-GW search methodology with a clean derivation, but the headline numbers are in-sample with respect to one fiducial population model, so they are a model-dependent forecast rather than a verified real-catalog performance.","tokens_in":28098,"tokens_out":2119,"would_cite":true,"duration_ms":19019,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PO2.0, a posterior-overlap statistic for strongly lensed gravitational waves, is mathematically equivalent to joint Bayesian parameter estimation yet fast enough for deep background studies, and identifies 65% of simulated galaxy-lensed…","keywords":["strong lensing","gravitational waves","Bayesian model selection","posterior overlap","binary black holes","parameter estimation","Bayes factor","selection effects"],"falsifier":"A concrete check is to compute PO2.0 Bayes factors and full joint-parameter-estimation Bayes factors on the same set of simulated lensed and unlensed pairs and compare their rankings and ROC curves: if the two statistics disagree substantially, the claimed equivalence is broken by the kernel-density and importance-sampling reconstruction. A complementary check is to run PO2.0 on a real O4 catalog and see whether the predicted $>2.25\\sigma$ candidate or the predicted detection rate actually appears.","tokens_in":27029,"feed_emoji":"🔭","tokens_out":7154,"duration_ms":64585,"temperature":0.7,"pith_summary":"Strongly lensed gravitational waves arrive as nearly identical pairs of signals, and the challenge is telling real lensed pairs from unlensed pairs that look alike by chance. This paper claims that a reweighted posterior-overlap statistic, PO2.0, is mathematically equivalent to the slow gold-standard joint Bayesian analysis but runs in minutes instead of about an hour. The improvement comes from feeding in population information: detectable-binary priors, galaxy-lens priors, selection effects, and the lensing-biased parameters (distance, phase, arrival time) that earlier overlap methods ignored. On simulated galaxy-lensed binary-black-hole pairs, PO2.0 identifies 65% of lensed pairs at a pairwise false-alarm probability near $2\\times10^{-6}$, translating to a 13% chance of a $>2.25\\sigma$ lensed-event detection in an 18-month, O4-like catalog. If the claim is right, the community can screen catalogs for lensed pairs without paying the computational cost of joint parameter estimation.","feed_headline":"PO2.0 catches 65% of lensed gravitational-wave pairs","feed_subtitle":"A fast overlap statistic matches expensive joint analysis, giving a 13 percent chance of a confident O4 lensing detection.","key_machinery":"The central object is the PO2.0 Bayes factor $B^L_U$, a ratio of lensed to unlensed evidences built from individual-event posteriors: posteriors $P(\\theta_{\\rm eq},\\theta_b|d_j)$ are divided by their parameter-estimation priors, multiplied by the population priors under the two hypotheses, and integrated over equal, biased, and bias parameters. Factorization into a sky-overlap factor $S^L_U$, a phase-overlap factor $P^L_{U,n}$, the time-delay ratio $R^L_U$, and a remaining overlap $B'$ reduces 14- and 16-dimensional integrals to manageable low-dimensional pieces. The numerical implementation uses kernel-density estimates of posteriors and importance sampling, with fast single-event parameter estimation to generate the hundreds of thousands of unlensed pairs needed for a deep false-alarm background.","core_discovery":"The paper's central claim is that the Bayes factor between the lensed and unlensed hypotheses, written as an integral over reweighted products of the two signals' individual posterior distributions, is mathematically equivalent to the joint-parameter-estimation Bayes factor. The key expression is Eq. (5): each individual posterior is divided by the prior used in that signal's parameter estimation, the product is weighted by the lensed population prior, and the integral runs over equal parameters ($\\theta_{\\rm eq}$) as well as the lensing-biased parameters ($\\theta_{b}$) and the biases themselves ($\\Delta\\theta_b$). This equivalence follows from splitting the joint likelihood into individual likelihoods, which is justified when the two signals do not overlap in time and their noise realizations are uncorrelated. With population priors that include selection effects and all these parameters, PO2.0 detects 65% of simulated galaxy-lensed binary-black-hole pairs at a pairwise false-alarm probability near $2\\times10^{-6}$. For an 18-month catalog at current detector sensitivity, the paper forecasts a 13% probability of detecting at least one lensed event above $2.25\\sigma$ significance, roughly double the chance offered by the earlier posterior-overlap method.","pith_inferences":["Our inference: the 13% detection forecast inherits the assumed lensing fraction $u=1.5\\times10^{-3}$, a detection rate of 100 events per year, and the O4 sensitivity curve; a different lensing fraction or high-redshift merger rate would directly change that number.","Our inference: because the Bayes factor is sensitive to the lens population model, the same machinery could be turned around to rank competing lens models on observed candidate pairs, an extension the paper flags as future work.","Our inference: a direct comparison of PO2.0 rankings with joint-parameter-estimation rankings on identical simulated pairs, which the paper says is underway, would quantify how much efficiency is lost to kernel-density reconstruction noise.","Our inference: extending the reweighted-overlap product to triplets and quadruplets of images, mentioned as future work, would directly attack the catalog false-alarm problem that grows quadratically with the number of detected signals."],"forward_implications":["A pair of detected signals can be ranked for lensing in minutes, so false-alarm backgrounds of hundreds of thousands of pairs become feasible; with about $5\\times10^5$ background pairs the paper probes pairwise false-alarm probabilities down to about $2\\times10^{-6}$.","If PO2.0's claimed equivalence to joint parameter estimation holds in practice, lensing searches no longer need to run hourly joint samplings over every candidate pair.","The forecast is a 13% probability of detecting at least one lensed pair above $2.25\\sigma$ catalog significance during 18 months of O4-like observation, compared with 7.5% for the earlier $B_{\\rm eq}\\times R^L_U$ statistic.","PO2.0 also produces joint posteriors for source and lens parameters by reweighting individual posteriors, yielding tighter intrinsic-parameter constraints and Morse-phase identification better than 80% efficiency at a false-alarm probability of 0.05.","All earlier fast statistics, including posterior overlap, the time-delay ratio, and the magnification-ratio statistic, are recovered as approximations of PO2.0; approximating distance and phase posteriors as delta functions degrades efficiency, which PO2.0 avoids by marginalizing over their uncertainties."],"supporting_citations":[{"why":"Supplies the original posterior-overlap Bayes factor and time-delay ratio that PO2.0 generalizes, plus the lensing simulation recipe used for injections.","marker":"Haris et al. (2018)"},{"why":"Provides the argument that population priors should describe detectable events rather than parameter-estimation priors, which PO2.0 adopts and benchmarks.","marker":"Cheung et al. (2023)"},{"why":"One joint-parameter-estimation formalism whose lensed-hypothesis evidence PO2.0 claims to reproduce by reweighting individual posteriors.","marker":"Liu et al. (2021b)"},{"why":"The other joint-parameter-estimation formalism compared against; PO2.0's expression matches it up to parameterization and detectable-population prior choices.","marker":"Lo & Magaña Hernandez (2023)"},{"why":"Defines the Morse phase and the geometric-optics lensed-waveform relation used to model image copies in Eq. (1).","marker":"Dai & Venumadhav (2017)"},{"why":"Provides the galaxy velocity-dispersion distribution used to build the lensed population prior.","marker":"Collett (2015)"},{"why":"Supplies the star-formation-rate-based merger-rate density that sets the high-redshift binary-black-hole population.","marker":"Madau & Dickinson (2014)"},{"why":"Gives the power-law-plus-peak mass and aligned-spin distributions used to inject the binary-black-hole population.","marker":"Abbott et al. (2023)"},{"why":"The fast parameter-estimation method used to produce individual posteriors and the roughly $10^3$-event background and foreground sets.","marker":"Roulet et al. (2022)"}],"fun_headline_variants":["Fast Bayesian search finds 65% of lensed GW pairs","PO2.0: fast lens detection matches joint analysis, 65% catch","Lensed GW pairs: efficient method detects 65% at low false alarm","New overlap method boosts lensed GW detection to 65%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated populations used to build the informative priors, namely the high-redshift merger rate, mass and spin distributions, galaxy lens population, and detector selection effects, faithfully represent the real detectable unlensed and lensed signals; PO2.0's reported efficiencies and the 13% forecast rest on those priors.","fun_headline_variants_meta":{"raw":{"variants":["Fast Bayesian search finds 65% of lensed GW pairs","PO2.0: fast lens detection matches joint analysis, 65% catch","Lensed GW pairs: efficient method detects 65% at low false alarm","New overlap method boosts lensed GW detection to 65%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2446,"prompt_tokens":1078,"completion_tokens":1368,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":694,"completion_tokens_details":{"reasoning_tokens":1290}},"tokens_in":694,"tokens_out":1368,"duration_ms":9887,"temperature":1.0,"reasoning_tokens":1290,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:30:40.030844+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check is to compute PO2.0 Bayes factors and full joint-parameter-estimation Bayes factors on the same set of simulated lensed and unlensed pairs and compare their rankings and ROC curves: if the two statistics disagree substantially, the claimed equivalence is broken by the kernel-density and importance-sampling reconstruction. A complementary check is to run PO2.0 on a real O4 catalog and see whether the predicted $>2.25\\sigma$ candidate or the predicted detection rate actually appears.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the galaxy velocity-dispersion distribution used to build the lensed population prior."},{"cited_title":"2022, Physical Review D, 106, 123015","cited_arxiv_id":null,"evidence_quote":"The fast parameter-estimation method used to produce individual posteriors and the roughly $10^3$-event background and foreground sets."}],"review_version":1}