{"id":"240a4732-318d-41a8-9d1b-ea5c530a0ff8","arxiv_id":"2501.05560","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Simulated bright siren events show LVK-era detections will not competitively constrain Horndeski gravity parameters, while one year of Einstein Telescope observations could detect αM ≠ 0 at over 3σ.","lead":"This paper uses simulated gravitational wave events with electromagnetic counterparts to forecast how well future detectors can test modified gravity. It finds that the next ten bright sirens will not strongly constrain cosmological gravity, but one year of the Einstein Telescope could detect a deviation from general relativity at more than 3 sigma.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ET 3σ forecast is self-consistent under the ΩΛ ansatz, but the ansatz dependence is unquantified; a steeper low-z αM(z) could reduce the integrated signal below 3σ.","rationale":"The reader's weakest_assumption correctly identifies the ΩΛ(z) parametrisation as the fragile input to the ET 3σ detection claim. I agree that this is the most load-bearing element of the central claim. My partial rather than full agreement comes from a quantitative check: among the three forms explicitly listed in eq 2.7, ΩΛ is actually the least aggressive for z<0.5, so the listed alternatives would not weaken the claim. The danger is from unlisted steeper low-z ansätze or from treating the injection and recovery with the same functional form as evidence of robustness. The paper's own Section 6 caveat is asserted, not quantified, and the ET forecast is precisely the place where an order-unity change in the integrated effect decides whether the 3σ headline survives. The LVK conclusion that ten bright sirens alone are not competitive is robust to this concern. The αT0 posterior interpretation error in Section 5.1 does not bear on the central ET claim, but it is a concrete mistake worth fixing. Overall the conditional verdict remains appropriate; no change is needed beyond the reader's existing conditionality.","tokens_in":26051,"tokens_out":15210,"duration_ms":155565,"concrete_test":"Repeat the ET analysis (Sections 3.2/5.2) with the same 150 mock events and injected αM0=1, but generate the dGW and Δta data under alternative low-z-concentrated ansätze—e.g. αM(z)=αM0[ΩΛ(z)/ΩΛ0]^2, αM(z)=αM0(1+z)^{-4}, and αM(z)=αM0 exp(-z/0.15)—while recovering with the ΩΛ ansatz. Compute the recovered αM0 posterior and its significance relative to GR. If all alternatives still exceed 3σ, the concern is resolved; if any plausible alternative falls below 3σ, the abstract's forecast should be explicitly qualified as ansatz-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim is the Section 5.2/Figure 5 result that 150 bright-siren ET events detect αM0=1 at >3σ. This detection is driven by the integrated modification in eq 2.8, whose amplitude over the GRB-limited range z<0.5 is set by the ΩΛ(z) ansatz of eq 2.7. The same ansatz is used to generate the mock data and to build the recovery likelihood, so the forecast tests self-consistency, not robustness to the functional form of αM(z). For z≤0.5 the integrated kernel I_Ω=∫0^z (ΩΛ/ΩΛ0)(1+z')^{-1}dz' is ≈0.22 at z=0.5; the other forms in eq 2.7 bracket it (α∝a gives ≈0.33, α∝a^3 gives ≈0.24), so among the listed ansätze ΩΛ is not the most aggressive. But a steeper low-z concentration, e.g. αM∝[ΩΛ/ΩΛ0]^2 or an exponential cutoff at z~0.1-0.2, can reduce I by more than a factor of two and take the recovered αM0 significance from ~5σ toward or below 3σ. Section 6 acknowledges the ansatz could change results but only asserts 'not expected to exceed order unity', which is exactly the margin on which the headline 3σ claim sits. This does not affect the LVK qualitative conclusion (Section 5.1). The separate text error in Section 5.1, where the αT0 posterior from Figure 4 is described as one order of magnitude tighter than GW170817 when it is actually wider, is secondary and should be corrected.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents forecasts for joint inference of the Hubble constant H0 and the Horndeski parameters αM and αT using bright siren gravitational-wave events, i.e. binary neutron star mergers with electromagnetic counterparts. The analysis covers two detector eras: a LIGO-Virgo-KAGRA O4/O5-like scenario with 10 mock events (one with an associated GRB) and an Einstein Telescope scenario with 150 GRB-associated mock events. The authors build a Bayesian likelihood including selection effects, marginalize over emission time delays, and use the ΩΛ(z) parametrisation of αM(z) and αT(z). Their main results are that ten LVK bright sirens alone leave αM0 essentially unconstrained and are not competitive with current dark siren or DESI constraints, while 150 ET events with GRB information recover the injected αM0=1 at better than 3σ, and also improve H0 and αT0 constraints. The paper emphasizes the role of GRB detection in breaking the inclination-distance degeneracy and thereby improving distance estimation.","tokens_in":26416,"tokens_out":13366,"duration_ms":134061,"significance":"If the ET forecast is robust, the paper gives a useful quantitative answer to a frequently asked question about the near-term and long-term value of bright sirens for cosmological tests of gravity. The LVK conclusion is well supported: with only ten nearby events, the integrated modification in Eq. (2.8) is small and the recovered αM0 posterior is far wider than current non-GW constraints; this part of the paper is robust to the modelling simplifications. The paper also has clear strengths: the selection-effect treatment in Eqs. (4.4)-(4.9), the arrival-time-delay derivation in Appendix B, the use of bilby-based simulated distance posteriors for the LVK scenario, and the honest discussion of caveats in Section 6. The main weakness is that the headline Einstein Telescope 3σ claim is conditional on the chosen ΩΛ(z) ansatz and on the assumed 150 GRB-associated events per year. These dependencies are acknowledged but not quantified, and the current caveat in Section 6 that changes are 'not expected to exceed order unity' is asserted rather than demonstrated.","major_comments":[{"comment":"The headline claim that one year of Einstein Telescope observations detects αM0=1 at greater than 3σ is sensitive to the assumed redshift parametrisation of αM(z). The mock data and the recovery likelihood both use the ΩΛ ansatz of Eq. (2.7), so the forecast demonstrates self-consistency under that ansatz rather than robustness to the functional form of αM(z). Because the GRB flux cut restricts the ET sample to z≲0.5, the integrated kernel in Eq. (2.8) is only about 0.22 at z=0.5 under this ansatz, and a more low-z-concentrated form such as αM∝[ΩΛ/ΩΛ0]^2 reduces that kernel by more than a factor of two. This would move the recovered αM0 significance from roughly 5σ toward or below the 3σ threshold. The Section 6 caveat that modifying the ansatz is 'not expected to exceed order unity' is precisely the margin on which the 3σ claim rests, yet no calculation is provided. Please quantify the robustness by repeating the ET forecast for the three ansätze in Eq. (2.7) and for at least one steeper low-z ansatz, and report the minimum significance over this set.","section":"§2.3, §5.2, §6; Eqs. (2.7)-(2.8); Fig. 5"},{"comment":"The 'one year' ET claim also depends on the assumed yield of 150 GRB-associated events. This number is built from an ET BNS rate at the high end of forecasts (up to 6×10^4 events per year), a 6% inclination selection, a Fermi/Swift flux threshold, and a GRB luminosity distribution centred at 5×10^49 erg/s with only 10% dispersion. Each of these choices is defensible, but the combination is optimistic: a factor of two or three reduction in the joint detection rate, e.g. from a lower BNS rate within current bounds or a broader short-GRB luminosity function, would reduce the αM0 significance roughly as sqrt(N) and could push the 150-event result toward or below 3σ. Since Fig. 7 already shows results for N=50, 150, 300 and 500, please make the N-dependence explicit in terms of detection significance and state the minimum joint rate required for a 3σ one-year claim. This would make the forecast robust to the main rate and selection uncertainties.","section":"§3.2, §5.2; Figs. 5 and 7"}],"minor_comments":[{"comment":"The statement that the αT0 posterior is 'one order of magnitude tighter than GW170817' is incorrect: the quoted result αT0 = -22.53+77.10/-76.98 in units of ×10^-16 is wider than the GW170817-based bound |αT0|≲ few×10^-15. This comparison should be corrected because it is used to characterize the LVK result.","section":"§5.1, Fig. 4"},{"comment":"The sentence beginning 'Although the αi are largest at low redshifts...' ends with the incomplete clause 'starts making a more marked at higher redshifts.' Please complete the sentence and clarify the intended comparison between ansätze.","section":"§2.3"},{"comment":"There are several typographical and grammatical errors, e.g. 'cosider' should be 'consider', and 'This, along with our results later on suggest' should be 'This, along with our results later on, suggests'. A careful proofreading pass is needed.","section":"§3.1"},{"comment":"The relationship between Figure 3 and Figure 4 is confusing: Figure 3 does not show αT0, while the text says the purple contours 'represent the scenario where we wish to constrain both GR and cosmology' with flat priors on H0, αM0 and αT0. Please clarify whether αT0 is marginalized or fixed in each panel and harmonize the quoted posterior numbers between the text and the captions.","section":"Fig. 3 caption and §5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid forecasting paper whose LVK conclusion should survive further scrutiny; the main risk is the ET 3σ headline, which rests on the ΩΛ ansatz and on the assumed yield of 150 GRB-associated events. The authors' own caveat in Section 6 is in the right direction but is not quantified. I would be comfortable with acceptance after a robustness section that reports significance across ansätze and event rates. No concerns about scope, attribution, or citation practice."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on GW cosmology. The paper builds a full Bayesian forecast for bright sirens in Horndeski gravity, with selection effects and GRB time-delay information, and lands on two results: ten LVK-era bright sirens will not competitively constrain αM, and one year of Einstein Telescope with 150 GRB-associated events could detect αM0=1 at >3σ. The first conclusion is solid and the more useful one. The second is real but rests on the ΩΛ(z) ansatz for the α parameters, and the paper's own Section 6 admits that a different ansatz could change the results, with differences 'not expected to exceed order unity.' As the stress-test note says, that order-unity margin is exactly where the 3σ claim sits; a steeper low-z αM(z) can cut the integrated signal by more than a factor of two and drop the significance toward or below 3σ. So the ET headline should be read as a self-consistency check of the parametrisation, not a robust prediction.\n\nWhat the paper does well: the setup is careful. The likelihoods include the inclination–distance degeneracy, the GRB selection via flux threshold, and the emission-time-delay marginalisation. The comparison between GR+cosmology and GR-only scenarios is clean. The qualitative message – that near-term bright sirens alone won't settle modified gravity, so dark sirens and cross-correlations remain essential – is well supported and worth stating plainly. The αT0 text error in Section 5.1 is minor and fixable: the posterior shown in Figure 4 is wider than the GW170817 bound, not one order of magnitude tighter.\n\nThe main limitation is reproducibility: no code or data provided, so the specific numbers can't be checked. That is a common gap for forecast papers, but here the headline depends on the ansatz and the mock generation details, so it matters a bit more.\n\nWho it's for: anyone planning LVK O5 or ET observing strategy, and people working on modified gravity forecasts. It deserves a serious referee; the LVK conclusion is useful and the ET claim, though fragile, is an honest forecast under stated assumptions. I'd recommend sending to peer review and pushing for the ansatz dependence to be quantified, e.g. by rerunning the ET case under the other ansätze in eq. 2.7, and for the αT0 statement to be corrected.","headline":"Careful forecast with a robust LVK punchline and a fragile ET headline that leans on the α(z) ansatz; definitely worth a serious referee.","tokens_in":26953,"tokens_out":1813,"would_cite":true,"duration_ms":17224,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["83C35","83D05","83F05"],"pacs":["04.30.-w","04.80.Nn","95.36.+x","98.80.-k"],"model":"deepseek-v4-flash","headline":"The paper forecasts that the next ten bright sirens will not competitively constrain modified gravity, while one year of third-generation detector data could detect a Horndeski-style departure from GR at greater than 3σ.","keywords":["bright sirens","gravitational wave cosmology","modified gravity","Horndeski theory","standard sirens","third-generation detectors","gamma-ray bursts","Hubble constant"],"falsifier":"Re-run the third-generation forecast with the alternative ansätze $\\alpha_i(z)=\\alpha_{i0} a$ and $\\alpha_i(z)=\\alpha_{i0} a^p$ used in the paper, and with gamma-ray burst luminosities drawn from the brighter end of the observed distribution; if the $\\alpha_{M0}$ posterior no longer excludes GR at $3\\sigma$, the headline forecast fails. The first real year of third-generation bright-siren data would settle it directly.","tokens_in":1968,"feed_emoji":"🔭","tokens_out":2294,"duration_ms":100722,"temperature":0.7,"pith_summary":"This paper asks whether the next bright sirens—gravitational wave events with electromagnetic counterparts, which pin down the source redshift—can by themselves test whether gravity differs from general relativity on cosmological scales. It argues they cannot in the near term: ten simulated bright sirens from the upcoming current-generation network leave the Horndeski parameter $\\alpha_M$, which controls how the effective Planck mass runs with time, too poorly constrained to compete with existing dark-siren or galaxy-survey results. The same analysis finds that one year of third-generation detector data, with gamma-ray burst counterparts for roughly 150 events, could detect a mild departure such as $\\alpha_M \\neq 0$ at more than $3\\sigma$. The point matters for planning: it says bright-siren follow-up alone will not settle modified gravity in the 2020s, and that a portfolio of dark siren, bright siren, and large-scale-structure methods is needed.","feed_headline":"One year of third-gen GW data could detect modified gravity","feed_subtitle":"Forecast: bright sirens alone won't settle cosmology, but gamma-ray-linked events at next detectors could.","key_machinery":"The argument runs on the ratio between the gravitational-wave luminosity distance $d_{\\rm GW}$ and the electromagnetic luminosity distance $d_L$. In Horndeski gravity, $d_{\\rm GW} = d_L \\exp\\!\\left(\\int_0^z \\frac{\\alpha_M(z')}{2(1+z')}\\,dz'\\right)$, so any departure from GR accumulates with source redshift, while $\\alpha_T$ changes the gravitational wave speed and produces an arrival-time delay between the GW and gamma-ray burst signals. The paper adopts the $\\Omega_\\Lambda$ redshift ansatz $\\alpha_i(z) = \\alpha_{i0}\\,\\Omega_\\Lambda(z)/\\Omega_{\\Lambda 0}$ and links the distance and speed effects through the same Horndeski scalar–tensor framework rather than treating them independently. The resolving power comes from gamma-ray burst detection: restricting binary inclination to less than $20^\\circ$ breaks the distance–inclination degeneracy and tightens distance posteriors, while the full Bayesian likelihood marginalizes over the emission time delay and includes selection effects.","core_discovery":"The paper's central claim is a quantitative forecast. For a fiducial Horndeski universe with $\\alpha_{M0}=1$, the next ten bright sirens (only one with a detected gamma-ray burst) give $H_0 = 69.44^{+6.50}_{-5.55}\\,\\mathrm{km\\,s^{-1}\\,Mpc^{-1}}$ and an $\\alpha_{M0}$ posterior so wide that GR and $\\alpha_{M0}=1$ are both consistent; these events are not competitive with existing probes. If the Hubble tension is resolved by an external prior on $H_0$, the $\\alpha_{M0}$ error bars shrink by 65–70 percent, but the result remains comparable to current constraints. In the third-generation scenario, 150 bright sirens with gamma-ray burst information restrict the binary inclination range, breaking the inclination–distance degeneracy, and yield $\\alpha_{M0} = 0.98^{+0.22}_{-0.18}$, $\\alpha_{T0} = 2.91^{+0.22}_{-0.20} \\times 10^{-16}$, and $H_0 = 70.17^{+1.98}_{-1.66}\\,\\mathrm{km\\,s^{-1}\\,Mpc^{-1}}$, excluding GR at more than $3\\sigma$ for the injected $\\alpha_{M0}=1$. The paper also shows that without the gamma-ray-burst-restricted inclination, the third-generation error bars roughly double, and that wrongly assuming GR when the universe is non-GR biases $H_0$ badly.","pith_inferences":["A testable extension: applying the same likelihood to space-based detectors at higher redshift should strengthen the $\\alpha_M$ signal, since the effect in the distance ratio accumulates with propagation distance; the paper only simulates ground-based detectors.","The results imply that the gamma-ray burst detection horizon, not the gravitational wave horizon, is the binding constraint for bright-siren cosmology; pushing GRB sensitivity beyond the simulated cutoff would directly convert more third-generation events into useful probes.","The strong $H_0$\\u2013$\\alpha_M$ degeneracy shown here suggests that any independent improvement in the Hubble constant measurement will propagate directly into sharper modified-gravity constraints, a lever arm the authors quantify in their testing-GR-only scenario."],"forward_implications":["With only the next ten bright sirens, $\\alpha_{M0}$ will remain consistent with both GR and $\\alpha_{M0}=1$, so near-term bright-siren campaigns cannot by themselves rule cosmological modified gravity in or out.","A one-year third-generation campaign with roughly 150 gamma-ray-linked events should detect $\\alpha_{M0}=1$ and the injected $\\alpha_{T0}$ at more than $3\\sigma$, making modified gravity detectable if it is present at that level.","Gamma-ray burst information is worth roughly a factor of two in error bars: restricting inclination with GRB data halves the $H_0$ and $\\alpha_{M0}$ uncertainties in the third-generation scenario.","If the Hubble tension is resolved first, ten bright sirens become 65–70 percent more powerful for $\\alpha_{M0}$, so external cosmology priors and modified-gravity tests are not independent.","Bright and dark sirens should be analysed jointly; the forecasts support a portfolio strategy rather than relying on a few exceptional events."],"supporting_citations":[{"why":"Supplies the GW170817/GRB170817A joint detection that sets the time-delay measurement precision and the $\\alpha_T$ prior used in the forecasts.","marker":"[15]"},{"why":"Introduces the $\\alpha_M$ and $\\alpha_T$ parameters and the $\\Omega_\\Lambda$ redshift parametrisation the analysis adopts.","marker":"[61]"},{"why":"Gives the modified gravitational-wave luminosity distance relation that converts $\\alpha_M$ into a distance-ratio effect.","marker":"[74]"},{"why":"Provides the current dark-siren constraint $\\alpha_{M0}=1.5^{+2.2}_{-2.1}$ that the bright-siren forecasts are compared against.","marker":"[57]"},{"why":"Supplies the selection-effect term in the likelihood via joint cosmological and population inference.","marker":"[9]"},{"why":"Provides the general method for computing selection biases from multiple uncertain observations, used for both detector scenarios.","marker":"[103]"},{"why":"Supplies the gamma-ray burst flux thresholds and detection model that set which simulated bright sirens have counterparts.","marker":"[100]"},{"why":"Provides the third-generation detector sensitivity and event-rate assumptions behind the 150-event scenario.","marker":"[79]"}],"fun_headline_variants":["Third-gen GWs: one year to spot modified gravity at 3σ","Bright sirens alone won't crack cosmology; third-gen might","Gamma-ray bursts could unlock gravity tests in 3rd-gen detectors","Next 10 bright sirens won't settle gravity; one year of 3G might"],"cache_read_input_tokens":28928,"weakest_assumption_plain":"The forecast that third-generation data will detect $\\alpha_M \\neq 0$ at $3\\sigma$ rests on the assumed redshift ansatz $\\alpha_i(z) \\propto \\Omega_\\Lambda(z)$; a different but still viable redshift dependence for $\\alpha_M$ and $\\alpha_T$ could shift the integrated distance and time-delay effects enough to make the claimed detection marginal.","fun_headline_variants_meta":{"raw":{"variants":["Third-gen GWs: one year to spot modified gravity at 3σ","Bright sirens alone won't crack cosmology; third-gen might","Gamma-ray bursts could unlock gravity tests in 3rd-gen detectors","Next 10 bright sirens won't settle gravity; one year of 3G might"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000741,"raw_usage":{"total_tokens":3382,"prompt_tokens":1097,"completion_tokens":2285,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":2204}},"tokens_in":713,"tokens_out":2285,"duration_ms":17274,"temperature":1.0,"reasoning_tokens":2204,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:13:46.288694+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the third-generation forecast with the alternative ansätze $\\alpha_i(z)=\\alpha_{i0} a$ and $\\alpha_i(z)=\\alpha_{i0} a^p$ used in the paper, and with gamma-ray burst luminosities drawn from the brighter end of the observed distribution; if the $\\alpha_{M0}$ posterior no longer excludes GR at $3\\sigma$, the headline forecast fails. The first real year of third-generation bright-siren data would settle it directly.","supporting_citations":[{"cited_title":"Joint gravitational wave-short GRB detection of Binary Neutron Star mergers with existing and future facilities","cited_arxiv_id":"2401.13636","evidence_quote":"Supplies the gamma-ray burst flux thresholds and detection model that set which simulated bright sirens have counterparts."}],"review_version":1}