{"id":"768bcd07-7ca9-4043-a447-5825b842855c","arxiv_id":"2512.16916","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"GW231123's apparent gravitational-lensing signal has a false-alarm probability around 4σ, so the event cannot be claimed as lensed under the two-image wave-optics model.","lead":"A deep-learning gravitational-wave analysis re-examines the leading lensing candidate GW231123 and finds its lensing evidence is only at the 4-sigma false-alarm level, not enough to claim a detection. The method cuts analysis time from CPU days to minutes, making it feasible to run the large simulation sets needed to vet rare signals.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Background Bayes factors from DINGO-lensing are unvalidated; reported 4σ false-alarm rate may be biased.","rationale":"The reader's weakest_assumption focuses on model dependence and the injection population. I agree those are limitations, but the most load-bearing assumption for the actual number '4σ' is that DINGO-lensing computes log10 B_lens accurately for every one of the 70,000 background events. The paper validates the network against bilby only on the single event, not on the distribution that defines the false-alarm rate. The low importance-sampling efficiency (0.58%) makes systematic bias plausible, and the abstract/body discrepancy adds ambiguity about which simulations were actually used. These concerns do not overturn the paper's direction—GW231123 is unlikely to be a 5σ lensing detection under the stated model—but they make the exact significance estimate conditional on verification. Since the reader already returned CONDITIONAL, I recommend keeping that verdict rather than moving to ACCEPT or REJECT.","tokens_in":10178,"tokens_out":7414,"duration_ms":77067,"concrete_test":"Randomly select ~200 of the 70,000 nonlensed injections, stratified by SNR and spin, and recompute their log10 B_lens with bilby using the same priors and waveform. Compare the complementary CDF with the DINGO-lensing curve in Fig. 4: check whether the fraction with log10 B_lens > 0 is consistent with 8% within Poisson errors and whether the tail past log10 B_lens = 4.0 contains a similar number of events. If bilby yields a significantly different tail, rerun the false-alarm estimate with reference values and update the significance. Separately, verify whether the abstract's '200,000 simulations / 3 waveform models' corresponds to an extended simulation set; if not, correct the abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numerical claim — a ~4σ false-alarm probability for GW231123's lensing support — is computed by evaluating log10 B_lens with DINGO-lensing on 70,000 synthetic nonlensed events (body text, 'GW231123: reanalyzing...'). The only cross-check against a conventional sampler (bilby) is performed for the real event itself, with a log10 B difference of 0.14. No validation is shown for the background population, which spans a wide range of SNRs, spins, and mass ratios. The importance-sampling efficiency for the event is only 0.58%, suggesting the proposal distribution is a mediocre fit; for other signals the evidence estimate could be biased, shifting the tail count that defines the p-value. If the network systematically overstates lensing support for weak or high-spin signals, the 8% false-positive fraction and 4σ value would overestimate the true significance (making the conclusion spuriously strong); if it understates support, the bound would be too conservative. A separate but also load-bearing issue is the inconsistency between the abstract ('more than 200,000 simulations with 3 different waveform models') and the body (70,000 nonlensed + 1,000 lensed, all with NRSur7dq4); this must be resolved to know what was actually computed. Both issues undermine confidence in the exact numerical claim, even if the direction (sub-5σ) might survive.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents DINGO-lensing, a simulation-based inference pipeline for parameter estimation of gravitationally lensed binary-black-hole signals under the stationary-phase, two-image wave-optics transfer function F(f)=1+sqrt(mu_rel) exp[i(2*pi*f*Delta t - pi/2)]. The authors reanalyze GW231123, the currently most promising lensing candidate, and obtain log10 B_lens = 4.0 in favor of lensing. By computing lensing Bayes factors for 70,000 nonlensed GW231123-like injections, they find that 8% of such signals have log10 B_lens > 0 and that the false-alarm probability at the observed Bayes factor is around 4 sigma. They therefore conclude that, under this lensing model, GW231123 cannot be claimed as a lensed event and that waveform self-similarity of short-duration signals can mimic lensing. They also report that 58% of 1,000 lensed injections have larger lensing support than GW231123, indicating that higher detection statistics are possible.","tokens_in":10498,"tokens_out":9254,"duration_ms":101451,"significance":"If the pipeline is validated, this work removes a major computational barrier to computing lensing false-alarm rates and provides an important cautionary result for GW231123. The paper has several genuine strengths: the false-alarm rate is obtained by direct simulation rather than fitted to the event; the event-level posterior and Bayes factor are checked against bilby (log10 B difference 0.14); the code is public; and the lensing model and its limitations are explicitly stated. However, the central numerical claim rests on the reliability of a neural evidence estimator over a broad background population, and there is an unresolved discrepancy between the abstract's stated simulation campaign and the body of the paper. A careful revision that addresses the validation and presentation issues could make the 4-sigma bound credible.","major_comments":[{"comment":"The abstract claims 'more than 200,000 simulations with 3 different waveform models,' but the body reports 70,000 nonlensed and 1,000 lensed simulations, all with NRSur7dq4. No multi-waveform-model analysis is presented anywhere in the body. This is not a typo: the abstract's quantitative conclusion is linked to a simulation campaign that the paper does not describe. Please either add the missing multi-waveform analysis with exact counts and model names, or correct the abstract to report the actual 71,000 simulations with one waveform model.","section":"Abstract vs. 'GW231123: reanalyzing...' section"},{"comment":"The 4-sigma false-alarm probability is computed using DINGO-lensing to evaluate log10 B_lens for 70,000 nonlensed injections. The only validation against a conventional sampler is performed on the real event itself, with a log10 B difference of 0.14. No calibration is shown for the background population, which spans SNR>8, spins up to 0.99, mass ratios 0.2-1, and distances 0.6-8 Gpc. The quoted importance-sampling efficiency of 0.58% for the event also indicates that evidence estimates can be noisy. Because the tail count at log10 B_lens=4.0 sets the p-value, a modest systematic bias in the neural estimator could change the false-alarm rate materially. I request a validation on a subsample of nonlensed background injections using bilby or an independent estimator, with a bias/variance check as a function of SNR and spin.","section":"'GW231123: reanalyzing...' (Fig. 4)"},{"comment":"With 70,000 background simulations, a one-sided p-value of roughly 3.2e-5 (4 sigma) corresponds to only about two events above the observed threshold. The Poisson error on that tail is large, and the achievable resolution of the false-alarm probability is limited to about 1/70,000 ~ 4.2 sigma. The paper should report the observed tail count at log10 B_lens=4.0, the p-value with a confidence interval, and state the resolution limit. The current phrasing '4-sigma false-alarm probability' and 'cannot exceed 4-sigma' is more precise than the simulation size supports. The result is also conditional on the hand-selected GW231123-like injection population and should be labeled as such.","section":"Fig. 4 and 'GW231123: reanalyzing...'"},{"comment":"The central claim's scope is limited to the two-image stationary-phase lens model of Eq. (2), with lower magnification, opposite parity, and a fixed pi/2 phase. The paper itself states that more complex cases, such as same-parity images or mu_rel > 1, are excluded. The abstract's statement 'the event cannot be claimed as lensed' therefore holds only under this restrictive model. Either the authors should justify that this two-image SPA model covers the relevant lensing configurations for GW231123, or the abstract and title should explicitly qualify the conclusion as conditional on this model. As written, the broad claim overreaches the computed quantity.","section":"After Eq. (2)"}],"minor_comments":[{"comment":"The x-axis labels '1-sigma, 2-sigma, 3-sigma, 4-sigma' should be tied to a one-sided Gaussian conversion from a tail probability, or replaced by p-values. The figure also uses a dashed vertical line for GW231123 and small vertical lines for examples; a legend would remove ambiguity.","section":"Fig. 4"},{"comment":"The interpretation that multimodal Delta t posteriors correlate with the waveform period is stated qualitatively. A simple quantitative measure over the 70,000 background simulations, such as the fraction of multimodalities aligned with the instantaneous period, would strengthen this claim.","section":"Self-similarity discussion"},{"comment":"The paper reports a log10 B difference of 0.14 but not the actual Bayes factors from both codes or the statistical uncertainty on this difference. Reporting the values would make the validation check more reproducible.","section":"DINGO-lensing vs. bilby comparison"},{"comment":"The statement that the waveform starting frequency is set to 0 Hz and that waveforms are generated as far back in time from merger as allowed is non-standard and should be justified, as it affects the effective signal duration and the self-similarity inference.","section":"Footnote [67]"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The abstract/body simulation-count discrepancy is the most serious issue and must be fixed before acceptance. The background validation request is not a request for perfection, but it is essential because the 4-sigma claim is the paper's central quantitative contribution. No scope concerns; the result is of interest to the gravitational-wave community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful paper, but the headline significance number is less robust than it looks. The new content is the large injection campaign—70,000 nonlensed and 1,000 lensed GW231123-like events—run through a neural posterior estimator fast enough to make the background tractable. They find that 8% of nonlensed heavy binaries produce log B > 0, and that the event's apparent support (log B ≈ 4) corresponds to roughly a 4σ false-alarm rate. That answers a question the LVC left open, and the proposed two-step pipeline (fast screening, then focused network) is sensible. Credit also for validating against bilby on the event itself (log B difference 0.14), using the numerical-relativity surrogate best suited to this event, and being explicit that waveform systematics and non-stationary noise would push significance down.\n\nThe soft spots are real, though not fatal. The abstract says 'more than 200,000 simulations with 3 different waveform models,' but the body describes 70,000 nonlensed + 1,000 lensed, all with NRSur7dq4. That is a direct inconsistency that needs fixing; it is not clear what was actually computed. More substantively, the significance estimate is conditioned on the two-image stationary-phase model of Eq. (2). The authors say this openly and defer more complex lenses, but the 4σ is a statement about that model, not lensing in general. Same-parity images or different phase structure would invalidate the numbers. The stress-test concern about background validation is also worth taking seriously: the only bilby cross-check is on the single real event, and the 0.58% importance-sampling efficiency suggests a mediocre proposal fit somewhere in the parameter space. If the network biases Bayes factors for weak or high-spin signals in the tail, the 4σ could shift. I don't think that changes the qualitative conclusion—this event is not claimable at 5σ under this model—but the exact number should be treated with caution.\n\nWho this is for: people working on lensed-GW searches and on simulation-based inference in GW astronomy. It deserves a serious referee. My recommendation: send it to review, but require the authors to reconcile the simulation counts, show a few background validation checks against bilby (even a handful of tail events), and explicitly scope the claim to the Eq. (2) model.","headline":"Solid and useful, but the headline 4σ is conditional on a specific lens model and the abstract/body simulation counts don't match; still worth a real referee.","tokens_in":11052,"tokens_out":2370,"would_cite":true,"duration_ms":25533,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GW231123, the strongest gravitational-wave lensing candidate to date, shows apparent lensing support at about 4σ, below the 5σ threshold, and the paper argues this can be explained by the self-similarity of the signal.","keywords":["gravitational lensing","gravitational waves","GW231123","wave-optics diffraction","simulation-based inference","neural posterior estimation","false alarm probability","self-similarity"],"falsifier":"Take a random subset of the 70,000 GW231123-like nonlensed simulations and recompute their lensing Bayes factors with a conventional nested-sampling code instead of the neural estimator; if GW231123's log10 Blens = 4.0 becomes a 5σ outlier in that distribution, the paper's 4σ bound would be an artifact of the neural approximation.","tokens_in":10003,"feed_emoji":"🔭","tokens_out":8205,"duration_ms":78647,"temperature":0.7,"pith_summary":"This paper tries to determine whether GW231123, the most promising gravitational-wave lensing candidate found so far, can be claimed as a lensed event. The authors use a deep-learning-accelerated parameter estimation pipeline to run 70,000 nonlensed simulations that mimic GW231123's properties, and find that 8% of them favor lensing with Bayes factors above zero. This places the event's apparent support at a false-alarm probability of about 4σ, below the 5σ standard usually required for a discovery claim. They conclude the apparent lensing signal is consistent with the self-similarity of short-duration, high-mass waveforms, where the inferred time delay matches the waveform's own period. If true, the first lensed gravitational wave has not yet been found, and future claims must be backed by event-specific false-alarm calculations, which this method makes computationally feasible.","feed_headline":"GW231123 lensing support caps at 4 sigma","feed_subtitle":"Deep-learning reanalysis of 70,000 look-alike signals shows 8% mimic lensing; the apparent chirp is self-similarity.","key_machinery":"The load-bearing object is the two-image stationary-phase transfer function, F(f>0)=1+√μrel exp(i(2πfΔt−π/2)), which models lensing as a single additional chirp with time delay Δt>0 and relative magnification μrel<1. This parametrization covers point and fold singularities and many near-cusp configurations, and defines what the paper counts as a 'lensed event'. On top of it, the authors build a neural posterior estimator called DINGO-lensing, which compresses detector strain into latent features and uses a normalizing flow to approximate the lensed and nonlensed posteriors, then computes Bayes factors via importance sampling. The speedup — minutes instead of tens of CPU days — is what makes","core_discovery":"The paper reanalyzes GW231123 assuming a two-image lensing model in which the observed signal is the original chirp plus a fainter, delayed copy with a fixed π/2 phase shift. At face value the event shows strong support for lensing, with log10 Blens = 4.0, larger than previous point-mass lens analyses. However, when compared to 70,000 GW231123-like nonlensed simulations, 8% of those noise-only events also favor lensing, giving GW231123 a false-alarm probability of about 4σ, not the 5σ needed for a claim. The recovered time delay of about 22 ms matches the instantaneous period of the waveform near merger, indicating that the apparent second chirp can be explained by self-similarity of the sho","pith_inferences":["If the self-similarity explanation holds, then very high-mass binaries are intrinsically poor targets for standalone lensing searches; combining the two-image statistic with additional discriminators — such as higher-order modes, phase-coherence tests, or external information about a putative lens — could restore sensitivity.","The fixed π/2 phase and μrel<1 assumption is a narrow slice of lensing phenomenology: same-parity images or μrel>1 would produce different waveforms, and the 4σ bound does not apply to those cases. A natural next test is a free-phase two-image model on the same event.","The same background-simulation machinery could be applied to other waveform distortions that produce apparent extra structure in short signals (e.g., eccentricity or environmental dephasing); elevated false-alarm rates from self-similarity are likely a general feature of few-cycle gravitational-wave signals.","If 8% of nonlensed events in a catalog look lensed, then some already-detected events may be misclassified as lensed or overlapping; re-running archived candidates with this fast method could produce a uniform false-alarm census."],"forward_implications":["GW231123 cannot be claimed as the first detected lensed gravitational wave: its apparent lensing support is within the noise floor of similar-mass nonlensed events.","Future lensing claims must report event-specific false-alarm rates estimated from simulations of nonlensed events with matching source properties, rather than quoting a raw Bayes factor.","Because 8% of GW231123-like nonlensed events mimic lensing, isolated high-mass short-duration binaries are prone to false lensing interpretations; searches using this statistic alone will be noisy for that population.","The finding that 58% of injected lensed simulations have stronger support than GW231123 implies that genuinely lensed events of this type should be detectable with high significance when the signal is favorable.","Projecting about 40% of lensed simulations above 5σ indicates that a first lensed gravitational wave discovery is feasible in upcoming observing runs using this accelerated analysis strategy."],"fun_headline_variants":["GW231123 lensing odds drop to 4 sigma","AI reanalysis: GW231123 lensing not proven","Self-similar chirps mimic GW231123 lensing","GW231123 lensing support falls short of 5 sigma","Lensing mirage: GW231123 likely not lensed"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire 4σ bound assumes lensing of GW231123 would look like exactly two images — a fainter, later chirp with a fixed π/2 phase shift — and that the simulated nonlensed population (chirp masses 90–160 solar masses, mass ratios 0.2–1, distances 0.6–8 Gpc, SNR>8) fairly represents the event's noise background.","fun_headline_variants_meta":{"raw":{"variants":["GW231123 lensing odds drop to 4 sigma","AI reanalysis: GW231123 lensing not proven","Self-similar chirps mimic GW231123 lensing","GW231123 lensing support falls short of 5 sigma","Lensing mirage: GW231123 likely not lensed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1229,"prompt_tokens":816,"completion_tokens":413,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":330}},"tokens_in":560,"tokens_out":413,"duration_ms":4727,"temperature":1.0,"reasoning_tokens":330,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T15:21:55.302527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random subset of the 70,000 GW231123-like nonlensed simulations and recompute their lensing Bayes factors with a conventional nested-sampling code instead of the neural estimator; if GW231123's log10 Blens = 4.0 becomes a 5σ outlier in that distribution, the paper's 4σ bound would be an artifact of the neural approximation.","supporting_citations":[],"review_version":1}