{"id":"affbb5af-2c6e-466d-a928-ccef6991df4d","arxiv_id":"2607.03885","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Fusing a simulation-trained common-source mass posterior with waveform features raises lensed-event detection efficiency from 20.8% to 35.2% at 1% false-positive rate and lowers the SNR for 50% efficiency from 45.3 to 33.5.","lead":"A machine-learning ranker that checks whether multiple gravitational-wave signals share one source mass evolution can better separate true lensed images from chance overlaps. That ranking step is the bottleneck for using lensed black-hole mergers as probes of dark matter and cosmology.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 20.8%→35.2% and 45.3→33.5 claims rest on a threshold whose nominal 1% FPR is not preserved under the same population used for the efficiency numbers.","rationale":"The reader already identified the calibration fragility of the unrelated multiple-merger population as the weakest assumption and correctly kept the verdict CONDITIONAL pending public data/code and real-catalog calibration. The present stress test simply makes that assumption load-bearing for the specific numerical claim: the 20.8%→35.2% and 45.3→33.5 figures are not guaranteed to survive a true 1%-FPR re-thresholding on the same population used to quote them. Held-out lens-family ROC-AUC gains and intermediate-SNR improvements remain supportive of a real ranking benefit, so the concern does not push the verdict to REJECT; it keeps CONDITIONAL and raises the bar for any future ACCEPT. No other internal inconsistency (architecture, geometric-optics premise, or held-out lens-family design) appears stronger than this operating-point mismatch.","tokens_in":25947,"tokens_out":667,"duration_ms":5967,"concrete_test":"On the same observationally motivated BBH evaluation set used for Fig. 9, recompute the unrelated-score distribution, choose the threshold that yields empirical FPR = 0.01 on that set, and re-evaluate both direct and fusion detection efficiency and the ρ_net at 50% efficiency. If the fusion-minus-direct efficiency gain falls below ~5 percentage points or the D50 factor falls below ~1.1, the abstract claim as written is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim (abstract / §III.C / Fig. 9) reports fusion detection efficiency 20.8%→35.2% and ρ_net for 50% efficiency 45.3→33.5 at a “1% reference FPR threshold calibrated on the corresponding unrelated multiple-merger sample.” In the same section the authors show that when the evaluation population is switched to the observationally motivated BBH mass distribution, thresholds calibrated on the original reference unrelated sample yield empirical false-alarm probabilities of 0.125 (direct) and 0.260 (fusion) at the nominal 1% point, and 0.018 / 0.140 at the nominal 0.1% point. Thus the numerical gains are measured at an operating point that is no longer a 1% FPR under the population for which the efficiencies are quoted. The paper correctly flags that a practical search must re-calibrate, but the headline numbers themselves are still presented as 1%-FPR results for that population. The load-bearing condition is therefore that the reported efficiency and D50-scale improvements remain after the threshold is re-set so that the empirical FPR on the observationally motivated unrelated sample is truly 1%. If the gain collapses once that re-calibration is performed, the central quantitative claim does not hold as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a physics-informed ranking method for distinguishing lensed multi-image gravitational-wave signals from unrelated multiple-merger events in the same analysis window. It trains a neural posterior estimator for the common-source phase-consistency parameters (log Mc,z, η), then fuses finite posterior samples with a residual waveform classifier via attention gating. Training uses generic multi-image injections without lens-family labels; evaluation is held out on PM, SIS, SIE, and SIE+Shear lenses with O4 H1/L1 backgrounds. Averaged over held-out families the fusion model improves ROC-AUC from 0.785 to 0.844. For an observationally motivated BBH population the authors report detection efficiency rising from 20.8% to 35.2% at a nominal 1% FPR threshold, and the SNR for 50% efficiency falling from 45.3 to 33.5 (1.35× SNR-equivalent distance scale).","tokens_in":26397,"tokens_out":1161,"duration_ms":8840,"significance":"If the ranking gains survive proper threshold recalibration, the work is a useful methodological contribution: it encodes geometric-optics source consistency as a learnable intermediate representation and demonstrates held-out lens-family generalization without training on physical lens labels. Strengths include the explicit train/test lens-family separation, HPD/PIT posterior diagnostics, controlled mass-plane rejection maps, condition-bootstrap intervals, and an honest population-shift test that flags the calibration problem. The approach is a concrete step toward physics-guided ML searches in dense GW catalogs and is of clear interest to the lensing and multi-messenger communities.","major_comments":[{"comment":"Abstract and §III.C / Fig. 9: the headline claim that fusion raises efficiency from 20.8% to 35.2% and lowers ρ_net(50%) from 45.3 to 33.5 at a “1% reference FPR threshold” for the observationally motivated BBH population is not supported as stated. In the same section the authors show that thresholds calibrated on the original reference unrelated sample yield empirical FARs of 0.125 (direct) and 0.260 (fusion) on that population. The reported efficiencies are therefore measured at operating points that are no longer 1% FPR under the population for which they are quoted. The load-bearing quantitative claim requires re-setting the threshold so that the empirical FPR on the observationally motivated unrelated sample is truly 1% (and likewise for 0.1%), then re-reporting efficiency and D50. If the gain collapses after recalibration, the central numbers must be revised or removed from the ab","section":null},{"comment":"§II.B and Table I: the unrelated multiple-merger class is constructed by superposing 2–5 independent BBHs with coalescence times drawn from the same ~0.6 s window and SNR rescaled into 8–48. The paper itself notes that for a realistic detected rate R the expected number of additional mergers in such a window is only RΔt ≪ 1. The difficult tail that controls low-FPR performance is therefore an artificial high-multiplicity, high-SNR construction. The authors should either (i) reweight or re-sample the unrelated class to a rate-consistent prior before quoting search-like FPR operating points, or (ii) clearly demote all fixed-FPR recalls to diagnostic status and lead with ROC-AUC / population-recalibrated metrics only.","section":null}],"minor_comments":[{"comment":"Table II vs. abstract: averaged R@1% FPR is 0.136 (direct) / 0.155 (fusion), while the abstract quotes 20.8% / 35.2% for the observationally motivated population. A short clarifying sentence in the abstract or §III.A would prevent readers from conflating the two numbers.","section":null},{"comment":"Fig. 2 and Fig. 7: error bands or condition-bootstrap intervals on the SNR-dependent efficiency curves would make the intermediate-SNR gain easier to assess visually; the text already reports bootstrap intervals for the averages.","section":null},{"comment":"§II.A: the Morse-phase convention (nj = 0, 1/2, 1 vs. integer Nj) is stated twice; a single consistent notation would reduce clutter.","section":null},{"comment":"Data availability: the statement that processed data “will be made publicly available upon acceptance” is fine, but providing a temporary review link or DOI for the evaluation tables would aid reproducibility checks.","section":null}],"recommendation":"major_revision","confidential_remarks":"The methodological idea is sound and the held-out lens-family design is careful; the paper is close to publishable once the population-recalibrated efficiencies are reported honestly. I would not reject on novelty grounds—the fusion of NPE source-consistency with waveform morphology is a genuine contribution—but the abstract numbers as currently written overstate what the population-shift test itself shows. A major-revision cycle focused on recalibration should be sufficient."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this paper is a clean intermediate representation: train an NPE only on the common-source (log Mc,z, η) plane, treat the samples as a set, and fuse them with a frozen waveform ResNet via attention and a gate. Training is on generic multi-image injections; PM/SIS/SIE/SIE+Shear are held out. That separation is real, not cosmetic.\n\nWhat they do well is the diagnostic stack. Averaged held-out ROC-AUC moves from 0.785 to 0.844 with bootstrap intervals that do not cross zero. They show the posterior is compact on lensed events and split/broad on unrelated overlaps, map the mass-plane rejection surface, characterize the loud/compact/partly-consistent false-alarm tail, and run HPD/PIT checks. The posterior-alone branch is weak (as expected once morphology is stripped) and fusion helps most at intermediate SNR and on the SIE-family tests. That is honest engineering, not a black-box score bump.\n\nThe soft spot is exactly the one the stress-test flags, and the authors already put the numbers on the page. In §III.C / Fig. 9 they quote fusion efficiency 20.8%→35.2% and ρ50 45.3→33.5 at a “1% reference FPR” for the observationally motivated BBH population, then immediately show that thresholds calibrated on the original unrelated sample give empirical FARs of 0.125 (direct) and 0.260 (fusion) under that same population. So the headline efficiencies are not at true 1% FPR for the population they are attached to. They correctly say a real search must re-calibrate; they still lead with the unrecalibrated numbers. Whether the gain survives a true 1% cut on the observational sample is not shown. Scope is also limited (nonspinning, geometric optics, H1/L1, SNR-rescaled injections), and code/data are not public yet.\n\nThis is for people building GW lensing search pipelines and for anyone who cares about physics-informed intermediate representations in dense catalogs. Math and citations look solid; self-cites are in-line with their prior NPE work. I would send it to referees. I would cite the method and the held-out diagnostics; I would not quote the abstract efficiency pair without the recalibration caveat. Worth a reading-group slot if anyone is actively working lensing searches.","headline":"Solid physics-informed fusion ranker with real held-out gains; the abstract’s 20.8%→35.2% / 45.3→33.5 numbers are measured at a nominal 1% threshold that the same section shows is not 1% under the quoted population.","tokens_in":26979,"tokens_out":628,"would_cite":true,"duration_ms":6910,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Fusing common-source mass posteriors with waveform features raises lensed gravitational-wave detection efficiency from 20.8% to 35.2% against unrelated multiple-merger events.","keywords":["gravitational-wave lensing","binary black holes","neural posterior estimation","source consistency","multi-image signals","unrelated mergers","machine learning classification"],"falsifier":"Inject the same fusion and direct classifiers into a catalog-matched set whose unrelated multiple-merger population follows the observed mass and rate distribution, calibrate a true 1% empirical false-positive threshold on that population, and check whether fusion still lowers the 50%-efficiency network SNR from roughly 45 to 33.","tokens_in":26835,"feed_emoji":"🌌","tokens_out":1041,"duration_ms":18111,"temperature":0.7,"pith_summary":"Gravitational lensing of gravitational waves can probe compact lenses, dark-matter substructure, and cosmological distances, but the first practical barrier is ranking true multi-image lensed signals against chance overlaps of unrelated binary mergers in the same short analysis window. This paper builds a ranking method around a geometric-optics fact: lensing may change amplitudes, arrival times, and Morse phase offsets, yet the images must still share one intrinsic source phase evolution. It trains a neural posterior estimator on the two mass parameters that control that phase evolution, then fuses finite posterior samples with direct waveform features in an attention-gated classifier. Trained only on generic multi-image simulations and tested on held-out physical lens families, the fusion ranker, for an observationally motivated binary-black-hole population, raises detection efficiency at a 1% reference false-positive threshold from 20.8% to 35.2% and lowers the network signal-to-noise ratio needed for 50% efficiency from 45.3 to 33.5—a 1.35-fold larger SNR-equivalent distance scale. A sympathetic reader cares because source consistency can become an explicit guiding principle for machine-learning searches in the denser catalogs next-generation detectors will produce.","feed_headline":"Lensed GW finder lifts efficiency 20.8% to 35.2%","feed_subtitle":"Mass-plane consistency fused with waveforms reaches 50% detection at lower SNR, expanding searchable volume.","key_machinery":"Physics-informed posterior fusion: a neural posterior estimator supplies unordered samples of the shared phase-consistency parameters (log Mc,z, η); those samples are attention-gated with frozen waveform features so source consistency can re-rank events that morphology alone leaves ambiguous.","core_discovery":"A fusion classifier that combines a simulation-trained approximate posterior over the common detector-frame chirp mass and symmetric mass ratio with direct waveform features improves ranking of lensed multi-image signals against unrelated multiple-merger events relative to waveform morphology alone, even though training never sees the physical lens families used at test time.","pith_inferences":["Catalog pipelines that already emit mass posteriors could graft a lighter fusion head onto existing features without retraining a full residual waveform network from scratch.","If wave-optics or substructure effects break pure geometric-optics phase preservation, the posterior target would need diffraction-aware parameters rather than mass-plane consistency alone.","The 1.35 distance-scale gain implies that intermediate-SNR events now discarded as ambiguous overlaps may become usable for lensing cosmology once consistency ranking is applied.","Spinning or neutron-star binaries would require enlarging the phase-consistency target to the next leading coefficients that control their frequency evolution."],"forward_implications":["At a 1% reference false-positive threshold calibrated on the matching unrelated-merger sample, detection efficiency for an observationally motivated BBH population rises from 20.8% to 35.2%.","The network SNR needed for 50% detection efficiency falls from 45.3 to 33.5, corresponding to a 1.35 times larger SNR-equivalent distance scale.","The ranking gain holds on held-out point-mass, SIS, SIE, and SIE-plus-shear lenses never used as training labels.","Low false-positive performance remains controlled by loud, temporally compact, partly source-consistent unrelated overlaps, so any real search must rebuild the unrelated-merger population for its own catalog.","Extending the same common-source idea to sky response becomes more informative as detector networks grow beyond two sites."],"fun_headline_variants":["Physics-informed posteriors rank lensed GW multi-images higher","Mass-ratio consistency fusion lifts lensed GW efficiency to 35.2%","Chirp-mass posterior fusion cuts SNR need from 45.3 to 33.5","Waveform-plus-mass posterior flags lensed GWs over chance overlaps","Geometric-optics prior lifts multi-image detection over unrelated mergers"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"A false-positive threshold that is 1% on the paper’s simulated unrelated-merger sample will still mean roughly 1% once the real catalog’s mass distribution, overlap rate, and noise are used.","fun_headline_variants_meta":{"raw":{"variants":["Physics-informed posteriors rank lensed GW multi-images higher","Mass-ratio consistency fusion lifts lensed GW efficiency to 35.2%","Chirp-mass posterior fusion cuts SNR need from 45.3 to 33.5","Waveform-plus-mass posterior flags lensed GWs over chance overlaps","Geometric-optics prior lifts multi-image detection over unrelated mergers"]},"model":"grok-4.5","effort":"low","cost_usd":0.005194,"raw_usage":{"total_tokens":1487,"prompt_tokens":836,"num_sources_used":0,"completion_tokens":103,"cost_in_usd_ticks":51940000,"prompt_tokens_details":{"text_tokens":836,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":548,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":836,"tokens_out":103,"duration_ms":4753,"temperature":1.0,"reasoning_tokens":548,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T23:15:00.923452+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Inject the same fusion and direct classifiers into a catalog-matched set whose unrelated multiple-merger population follows the observed mass and rate distribution, calibrate a true 1% empirical false-positive threshold on that population, and check whether fusion still lowers the 50%-efficiency network SNR from roughly 45 to 33.","supporting_citations":[],"review_version":1}