{"id":"631bde4d-e075-4cfd-a14d-5479e6b2b0f1","arxiv_id":"2505.09507","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A forecast from simulated strongly lensed gravitational wave data finds that Einstein Telescope could constrain H0 to 0.4-0.7% and PPN gamma to 0.5-3.3% (68% CI) under idealized assumptions.","lead":"The paper simulates a 10-year Einstein Telescope run to show that the number of strongly lensed gravitational wave events and their time-delay distribution could constrain the Hubble constant to about 0.4-0.7% and the post-Newtonian parameter gamma to about 0.5-3.3% (68% credible intervals). It offers a potential new probe of cosmic expansion and of general relativity that does not rely on electromagnetic counterparts or waveform modeling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The forecast assumes exact time delays and perfect detection of every lensed event (Sec. III, Eq. 9), but realistic SNR thresholds and identification completeness could cut the effective sample by an order of magnitude, inflating the claimed H0 and gamma precision.","rationale":"The reader's verdict is CONDITIONAL, and I agree with that verdict, but I want to sharpen the load-bearing assumption. The reader's weakest_assumption emphasizes exact time delays; for GW, time-delay measurements are likely to be far more precise than the predicted week-to-month delays, so that assumption is actually mild. The more serious gap is that the likelihood treats Lambda as the number of detected lensed events. Eq. (11) has no factor for the probability that both images exceed the detection threshold or that the pair is identified as lensed. The forecast's precision is essentially a statement about Poisson fluctuations in a sample of tens of thousands of lensed events; if the usable sample is an order of magnitude smaller, the forecast's advantage over existing joint constraints disappears. This is a concrete, quantitative issue that the authors themselves acknowledge in Section V as future work. The internal inconsistency between the abstract (0.60–0.99%) and the text (0.42–0.69%) for H0 precision is a separate reporting problem, but it does not change the scientific verdict. Because the paper is an idealized forecast and the framework is internally consistent, the appropriate outcome is still conditional acceptance, with the required revisions being a selection-function treatment and a caveat in the abstract and conclusion. Hence UNCHANGED.","tokens_in":10108,"tokens_out":13948,"duration_ms":131604,"concrete_test":"Run the forecast with a detection threshold: using the ET sensitivity curve, assign each lensed pair a network SNR for the two images based on the SIS magnification ratio and source redshift, and retain only pairs with both image SNRs above a threshold (e.g., 8). Recompute Lambda and the time-delay distribution entering Eq. (9), then rerun the Bayesian analysis of Figs. 4–5. If the 68% intervals for h and gamma widen by more than a factor of about 2 relative to the quoted 0.60–0.99% and 0.53–3.3%, the central precision claim is not robust to selection effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV reports 68% constraints of 0.60–0.99% on h and 0.53–3.3% on gamma (Figs. 4–5), with the abstract and conclusion presenting these as the method's expected precision. These numbers are driven by the likelihood in Eqs. (9)–(13): in Eq. (11), the expected number of lensed events Lambda is the total lensing rate, containing no detection or identification probability. The simulation then draws N from Poisson(Lambda) and treats the time delays as exact. Figure 1 shows Lambda of order 2e4 to 7e4 for the fiducial ET setup, so the Poisson error on N is about 0.4%; this is what yields sub-percent constraints. In a real ET catalog, a lensed event is usable only if both images are detected above the SNR threshold and the pair is identified as lensed. For the SIS model, the fainter image magnification is |mu_-| = 1/y - 1, so a large fraction of pairs have a fainter image that can fall below threshold, especially for high-redshift sources near the detector horizon. The paper's (Tobs - Delta t) factor in Eq. (11) accounts only for the temporal window, not for SNR selection. If the identified fraction f is 0.1, N shrinks to about 5e3 and the statistical errors grow by roughly sqrt(10) or about 3, pushing h precision to a few percent and gamma to order 10%, no longer clearly outperforming existing joint probes (1.5% h, 8.7% gamma). For f = 0.01, the claimed superiority is lost entirely. The paper lists selection effects as future work in Section V, but the headline numbers are quoted without that caveat. The exact-time-delay assumption is secondary for GW, since arrival-time differences can be measured far more precisely than the weeks-to-months delays, so the selection and identification problem is the load-bearing assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a Bayesian framework for jointly constraining the Hubble constant (parameterized by h) and the post-Newtonian parameter γ using the population statistics of strongly lensed gravitational wave (GW) events from binary black hole mergers. The authors derive a PPN-modified time delay and optical depth for the singular isothermal sphere lens model, construct a likelihood from the total number of lensed events and the time-delay distribution, and simulate an Einstein Telescope observation to produce forecasts. Under a flat ΛCDM cosmology with various priors on the matter density, the paper reports 68% credible-level constraints of 0.60%–0.99% on H0 (abstract) or 0.42%–0.69% (main text) and 0.53%–3.3% on γ, claiming that these significantly outperform existing joint constraints from electro-magnetic lensing.","tokens_in":10481,"tokens_out":8325,"duration_ms":81731,"significance":"The method is conceptually novel: it avoids EM counterparts, waveform modeling, and resolved stellar kinematics, instead extracting cosmological and gravitational information from the population-level statistics of lensed GW events. The derivation of the PPN-modified time delay and optical depth is explicit, and the Bayesian likelihood is clearly laid out. If the idealized assumptions (exact time delays, perfect detection and identification, fixed merger rate and velocity dispersion function) hold, the forecast would establish lensed GW population statistics as a high-precision probe. The paper also benefits from a transparent discussion of limitations in the final section, listing several extensions for future work. However, because the headline precision depends on these idealizations, the significance claim must be qualified by a realistic treatment of detection efficiency and nuisance-parameter uncertainties.","major_comments":[{"comment":"The expected number of lensed events Λ in Eq. (11) contains no detection or identification probability: it is the total lensing rate integrated over the source and time-delay populations. The simulation then draws N from Poisson(Λ) with Λ on the order of 2×10^4–7×10^4 (Fig. 1), so the count term alone yields roughly 0.4% statistical errors. In a realistic Einstein Telescope catalog, only a fraction f of lensed pairs will have both images above the SNR threshold and be identified as lensed; for f=0.1 the uncertainties grow by roughly a factor of 3, pushing the h and γ constraints to a few percent and to order 10%, comparable to or worse than existing joint constraints. This is load-bearing for the claim of significantly outperforming existing probes. The forecast should either include a detection/identification efficiency in Eq. (11) or present the constraints as an explicit function of the completeness fraction f.","section":"Section III, Eq. (11); Section IV, Fig. 1"},{"comment":"The paper explicitly assumes that N lensed events have been detected 'with exact time delays' and treats the measured Δt_i as noiseless in the likelihood. Real time-delay measurements for GW images will carry uncertainties, and the identification process itself may require waveform-based selection that is not modeled. While GW time-delay metrology may be precise, the assumption of exactness is an idealization that should be stated as such and, ideally, relaxed with a realistic error model to demonstrate that the quoted precision is not inflated by this idealization.","section":"Section III, Eq. (9); Section IV, Eqs. (12)–(13)"},{"comment":"There is a direct internal inconsistency in the headline numbers. The abstract reports H0 precision of 0.60%–0.99%, while the main text (Section IV, after Fig. 5) reports 0.42%–0.69%. From the stated 68% intervals (σ_h = 0.0042 and 0.0069 at h = 0.7), the correct relative uncertainties are 0.6% and 0.99%, so the main text's 0.42% appears to be a numerical error. The comparison baselines also differ: the abstract says existing joint constraints 'typically achieve 2% precision on H0 and 20% precision on γ', whereas Section IV quotes '1.5% and 8.7%' from references [25–27]. These discrepancies must be reconciled before the quantitative claims can be evaluated.","section":"Abstract and Section IV"},{"comment":"The forecast fixes the BBH merger rate R = 5×10^5 yr^-1 and uses velocity dispersion function parameters from Ref. [39] without marginalizing over their uncertainties. Because Λ in Eq. (11) is directly proportional to R and is sensitive to the VDF parameters, the sub-percent constraints on h and γ are conditional on these astrophysical inputs being known exactly. A robust forecast should either include priors on R and the VDF parameters or explicitly state that the quoted precision is statistical only, conditional on fixed values of these nuisance parameters.","section":"Section IV, Eq. (11) and VDF parameters"}],"minor_comments":[{"comment":"The axis labels in Figure 2 appear transposed or mislabeled: the x-axis is labeled 'log10 Tobs' but the plotted quantity is the time-delay distribution, and the y-axis label 'log10 (Δt [hrs])' is also unclear. Please clarify which variable is on each axis.","section":"Figure 2"},{"comment":"The dimensionless impact parameter y is introduced without an explicit definition; please state its normalization (e.g., y = β/θ_E) and the range assumed in the simulations.","section":"Section II, Eq. (2)"},{"comment":"The text states that zmax must be rescaled for different h and Ωm to preserve the same maximum detectable luminosity distance, but the implementation of this rescaling is not shown. A short description or formula would improve reproducibility.","section":"Section IV, zmax rescaling"},{"comment":"The sentence 'no need the EM counterparts and GW waveform knowledge' is grammatically awkward; suggest rewording to 'does not require EM counterparts or GW waveform knowledge'.","section":"Section V"}],"recommendation":"major_revision","confidential_remarks":"The internal inconsistency between the abstract's and the main text's precision numbers, as well as the differing comparison baselines, suggests that the headline numbers were not carefully checked. The authors should also verify that the quoted existing constraints (1.5%/8.7% vs 2%/20%) are correctly attributed to the cited works. The paper's core idea is promising for the journal's scope, but the idealized forecast needs to be made more realistic or much more strongly caveated before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The new thing here is small but real: this paper takes the lensed-GW population approach from Jana et al. and adds the PPN parameter gamma to the joint likelihood, forecasting simultaneous h and gamma constraints from the event count and time-delay distribution. That specific combination has not appeared before. The derivation of the PPN-modified time delay and optical depth for the SIS model is plausible, and the Bayesian machinery is internally consistent. The paper is clearly written about what it does and what it assumes.\n\nWhat it does not do is produce a measurement. The precision numbers come from a forecast where simulated data are drawn from the same model used for inference. That is fine for a forecast, but comparing those numbers to real measured errors from EM lensing (1.5% on h, 8.7% on gamma) is apples-to-oranges. The abstract and the text even disagree on the H0 precision range: the abstract says 0.60–0.99%, the text says 0.4–0.7%. That should be fixed.\n\nThe bigger soft spot, which the stress-test note gets right, is the selection function. Equation (11) counts all lensed events with no SNR threshold and no identification probability. Real ET catalogs will only include events where both images cross detection threshold and the pair is recognized as lensed. For SIS, the fainter image magnification falls off inversely with impact parameter, so a large fraction of lensed pairs will be unusable. If only 10% of the predicted events survive selection, the Poisson error on N grows by sqrt(10), pushing the h constraint toward a few percent and gamma toward order ten percent. That erodes the claimed advantage over existing joint probes. This is not a fatal flaw in a forecast, but it is the load-bearing assumption, and it should be modeled or at least honestly caveated. The exact-time-delay assumption is secondary; GW arrival times are measured very precisely.\n\nThe paper also ignores uncertainties in the velocity dispersion function, which enters the optical depth directly, and no code or data are provided. The authors list these as future work, which is honest, but the headline numbers are quoted without those caveats inside the abstract.\n\nWho is this for? People working on third-generation detector science cases and lensed GW forecasts. It is a modest extension of an existing framework, but the question it asks — can the population jointly constrain H0 and gamma — is worth a serious look. I would send it to peer review, but the referee should push on selection effects and the comparison metric. If those are addressed, this could be a useful reference for ET planning, though I would not build my own analysis on its numbers.","headline":"A clean, internally consistent forecast that adds PPN gamma to the lensed-GW population method, but its headline precision rests on optimistic assumptions about detection completeness and selection effects.","tokens_in":11077,"tokens_out":984,"would_cite":false,"duration_ms":12178,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Bayesian forecast using only the number of strongly lensed gravitational-wave events and their time delays can jointly measure the Hubble constant to sub-percent precision and the post-Newtonian parameter $\\gamma$ to a few percent.","keywords":["strong gravitational lensing","gravitational waves","Hubble constant","post-Newtonian parameter","Einstein Telescope","Bayesian population analysis","time-delay distribution","cosmology"],"falsifier":"Run the same Bayesian analysis on simulated Einstein Telescope data that includes detector noise, waveform-based lensing identification, and realistic time-delay measurement errors, and compare the resulting 68% credible intervals on $h$ and $\\gamma$ with the paper's quoted ranges. If the intervals widen beyond roughly $1\\%$ on $h$ and approach the $8.7\\%$ baseline on $\\gamma$, the central claim would be falsified.","tokens_in":9844,"feed_emoji":"🔭","tokens_out":6948,"duration_ms":63208,"temperature":0.7,"pith_summary":"The paper argues that the population statistics of strongly lensed binary black hole gravitational-wave events—specifically the total number detected and the distribution of time delays between lensed images—carry enough information to jointly measure the Hubble constant $H_0$ and the post-Newtonian parameter $\\gamma$ far more precisely than existing jointly-constrained probes. Under a flat $\\Lambda$CDM cosmology, a 10-year observation with the Einstein Telescope, and a Planck-informed prior on $\\Omega_m$, the Bayesian forecast yields 68% credible intervals of roughly $0.5\\%$ to $1\\%$ on $H_0$ and $0.5\\%$ to $3.3\\%$ on $\\gamma$. The method needs no electromagnetic counterpart, no waveform model, and no resolved stellar kinematics of the lens galaxy; it uses only event counts and time delays. If the forecast holds, lensed gravitational-wave statistics become a clean and independent probe of both cosmic expansion and gravity.","feed_headline":"Lensed gravitational waves could pin H0 to under one percent","feed_subtitle":"Population counts plus time delays from lensed black-hole mergers would measure H0 and gravity's gamma without electromagnetic counterparts.","key_machinery":"The central object is the joint population likelihood $\\mathcal{L}(N,\\{\\Delta t_i\\}|\\Omega,T_{\\rm obs}) = \\mathrm{Poisson}(N;\\Lambda(\\Omega,T_{\\rm obs})) \\times \\prod_{i=1}^N p(\\Delta t_i|\\Omega,T_{\\rm obs})$, where $\\Omega = (h,\\Omega_m,\\gamma)$. The expected count $\\Lambda$ is obtained from the PPN-corrected strong-lensing optical depth, and the time-delay distribution $p(\\Delta t|\\Omega)$ is generated by marginalizing over lens parameters. The PPN modification enters through the scaling $\\psi_{\\rm PPN} = (1+\\gamma)\\psi_{\\rm GR}/2$, which rescales both image time delays and the lensing cross-section. This machinery works because the total event count and the shape of the time-delay distribution respond differently to $h$, $\\Omega_m$, and $\\gamma$, breaking the degeneracies that would limit a single statistic.","core_discovery":"The paper's central claim is that a combined Poisson likelihood for the total number $N$ of lensed events and a product likelihood for the observed time delays $\\{\\Delta t_i\\}$, built from a PPN-corrected singular-isothermal-sphere lens model, can jointly constrain the Hubble constant and the post-Newtonian $\\gamma$ to sub-percent and few-percent precision respectively in the Einstein Telescope era. The quoted 68% credible intervals are roughly $0.4\\%$--$1\\%$ on $H_0$ (the abstract states $0.60\\%$--$0.99\\%$, while the body text varies between $0.42\\%$ and $0.69\\%$) and $0.53\\%$--$3.3\\%$ on $\\gamma$, significantly better than previous joint constraints. The forecast assumes a flat $\\Lambda$CDM cosmology, a known velocity dispersion function, and exact time-delay measurements; it requires no electromagnetic counterpart, no waveform modeling, and no kinematic modeling of the lens galaxy.","pith_inferences":["The exact-time-delay assumption is likely optimistic; if realistic time-delay measurement errors (from detector noise and waveform-based identification) are comparable to the predicted delays of weeks to months, the quoted sub-percent precision on $H_0$ will degrade, so the next natural test is to rerun the forecast with injected measurement noise.","The separation between $h$ and $\\gamma$ leans heavily on the strong Planck prior on $\\Omega_m$; without it, the degeneracies visible in the paper's own Figures 1 and 2 would likely inflate the $\\gamma$ error, meaning the method's power is coupled to external cosmological information.","In practice, the bottleneck may not be timing precision but identifying lensed events at all: when millions of unlensed binary-black-hole signals are present, waveform-based lensing identification becomes a necessary selection step, and the paper's clean statistical statement will need to be folded with that selection function."],"forward_implications":["If the forecast is correct, lensed gravitational-wave population statistics will provide a joint $H_0$--$\\gamma$ probe that beats current electromagnetic time-delay lensing constraints by a large margin, with no reliance on electromagnetic counterparts or waveform modeling.","The method is directly applicable to the binary-black-hole-dominated gravitational-wave catalog, which is exactly the source population third-generation detectors will record in large numbers.","The approach converts a previously nuisance-like feature of strongly lensed events—their time delays—into a precision cosmological and gravitational observable.","The framework is modular: it can be extended to more realistic lens models, other source redshift distributions, and alternative cosmologies such as $w$CDM, as the paper notes in its outlook."],"supporting_citations":[{"why":"Supplies the population-likelihood framework for lensed gravitational-wave cosmography that this paper extends with the PPN parameter.","marker":"[15]"},{"why":"Extends the population approach to third-generation detectors, providing the event-rate and observable assumptions used in the forecast.","marker":"[16]"},{"why":"Provides the strong-lensing optical depth and cross-section formalism used in Eq. (3) of the paper.","marker":"[37]"},{"why":"Gives the PPN scaling of the lensing potential, $\\psi_{\\rm PPN} = (1+\\gamma)\\psi_{\\rm GR}/2$, that carries the gravity modification.","marker":"[34]"},{"why":"Supplies the SIS time-delay expression used in the PPN-corrected delay formula of Eq. (2).","marker":"[35]"},{"why":"Provides the velocity dispersion function parameters that set the lens population in the optical depth calculation.","marker":"[39]"},{"why":"Establishes the previous joint $H_0$--$\\gamma$ constraint via time-delay strong lensing that serves as the comparison baseline.","marker":"[25]"},{"why":"Supplies the Gaussian prior on $\\Omega_m$ used in the more realistic of the two constraint strategies.","marker":"[12]"}],"fun_headline_variants":["Lensed GW counts and delays jointly pin H0 to under 1%","Joint lensed-GW forecast: H0 <1%, gamma up to 3.3%","Sub-percent H0 and few-percent gamma from lensed GW populations","Einstein Telescope lensed GWs could measure H0 and gamma to 1% and 3%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the $N$ lensed events are identified with exact time delays $\\{\\Delta t_i\\}$; if real time-delay measurement errors or lensed-event identification losses are comparable to the weeks-to-months delays, the quoted precision is optimistic.","fun_headline_variants_meta":{"raw":{"variants":["Lensed GW counts and delays jointly pin H0 to under 1%","Joint lensed-GW forecast: H0 <1%, gamma up to 3.3%","Sub-percent H0 and few-percent gamma from lensed GW populations","Einstein Telescope lensed GWs could measure H0 and gamma to 1% and 3%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000815,"raw_usage":{"total_tokens":3592,"prompt_tokens":988,"completion_tokens":2604,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":2510}},"tokens_in":604,"tokens_out":2604,"duration_ms":20254,"temperature":1.0,"reasoning_tokens":2510,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:30:16.864672+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same Bayesian analysis on simulated Einstein Telescope data that includes detector noise, waveform-based lensing identification, and realistic time-delay measurement errors, and compare the resulting 68% credible intervals on $h$ and $\\gamma$ with the paper's quoted ranges. If the intervals widen beyond roughly $1\\%$ on $h$ and approach the $8.7\\%$ baseline on $\\gamma$, the central claim would be falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the strong-lensing optical depth and cross-section formalism used in Eq. (3) of the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SIS time-delay expression used in the PPN-corrected delay formula of Eq. (2)."},{"cited_title":"Cosmic Time Slip: Testing Gravity on Supergalactic Scales with Strong-Lensing Time Delays","cited_arxiv_id":"1906.06324","evidence_quote":"Establishes the previous joint $H_0$--$\\gamma$ constraint via time-delay strong lensing that serves as the comparison baseline."}],"review_version":1}