{"id":"bb3c47a7-d9a5-473e-91fa-4df9b76a9185","arxiv_id":"2502.03463","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Future gravitational wave detectors should measure most neutron star radii to within about 5%, while current detectors give biased and imprecise estimates.","lead":"This simulation study projects how precisely future gravitational wave detectors can measure neutron star radii. It finds that the Einstein Telescope plus Cosmic Explorer network should reach about 5 percent radius accuracy across most of the mass range, strengthening the case for next-generation observatories.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5% accuracy claim is only tested for EoSs inside the GW170817-derived 2300-model set; a true EoS in a sparse or excluded region could bias R beyond 5% even for ECC.","rationale":"The reader's weakest_assumption identifies the same non-uniform EoS prior, and I agree that it is the most load-bearing concern. The methodology is Bayesian model averaging over a fixed discrete set; with an incomplete set, the posterior approximates the closest available models rather than the true radius, and increasing SNR sharpens the likelihood onto the nearest model, which can carry a bias set by the discretization gap. The paper's own Sec. IVB disclosure supports this. The ECC results are promising because high SNR sharpens EoS discrimination and suppresses the bias for the three tested EoSs, but all three are inside the 2300-model set, so the out-of-set case is untested. The proposed check directly injects an EoS outside the set and measures the resulting bias. The zero-noise choice is a secondary concern: it is a standard device for isolating systematic bias, and the reported credible intervals still encode the expected statistical width, so it does not invalidate the projection. Code unavailability is a reproducibility issue rather than a correctness issue. The verdict remains CONDITIONAL: the claim is plausible but hinges on EoS-set coverage, which is testable and currently unverified.","tokens_in":20899,"tokens_out":6303,"duration_ms":62164,"concrete_test":"Generate 100 zero-noise BNS signals for the ECC network using a spectral EoS that is deliberately not contained in the public P1800115 set, e.g., a stiff EoS with R_1.4 ≈ 15.5 km and M_max > 2.1 M_sun, constructed with the same Lindblom spectral parameterization within the allowed coefficient ranges. Run the full Bilby+BEOMS evidence-weighted radius inference (Eqs. 14-17) on these injections and compute the fractional bias ΔR = (R_inf - R_true)/R_true for each event. If the median |ΔR| exceeds 5%, the accuracy claim in the abstract does not extend to EoSs outside the GW170817-constrained ensemble; if it stays below 5%, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central accuracy claim rests on Eq. (17): p(R|d) is a weighted mixture over the ~2300 discrete EoSs from the GW170817 re-analysis [63]. If the true EoS is absent or sparsely represented in this set, the posterior cannot converge to the true radius regardless of SNR; it converges to the nearest represented models. The authors themselves flag this in Sec. IVB and Fig. 2: the density of EoSs peaks near APR4 and is low above DD2, and they state that this non-uniformity 'introduces biases in the estimation of R.' However, the validation in Figs. 4-5 only injects APR4, SLy, and DD2, all of which are represented in the set, so it does not probe the failure mode where the true EoS lies in an empty or sparse region, e.g., R_1.4 > 15 km or < 10 km. Equal prior weight over the discrete models also implicitly weights the EoS space by the local density of the GW170817 posterior samples, which is not a physically motivated prior. Thus the abstract's 'accurate ... for both soft and stiff equations of state' is conditional on the unverified assumption that the 2300-model set covers the full plausible EoS space densely enough that the discretization gap is below the 5% threshold.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a simulation study of neutron-star radius measurement with current and next-generation gravitational-wave detector networks. Using Bilby with the dynesty sampler and relative binning, the authors perform zero-noise injections of 100 binary neutron star signals (IMRPhenomPv2_NRTidalv2) drawn from a double-Gaussian galactic mass population, for three injected equations of state (APR4, SLy, DD2) and three networks (O5 at A+ sensitivity, the A# upgrade, and an Einstein Telescope plus Cosmic Explorer combination labeled ECC). Radius inference proceeds by computing Bayesian evidences for roughly 2300 equations of state taken from the GW170817 spectral-decomposition public release of Ref. [63] and forming an evidence-weighted mixture posterior for R (Eq. 17). The main claims are that O5 and A# radius estimates are biased and imprecise, that ECC achieves fractional radius errors of about 5% or better across most of the mass range for both soft and stiff equations of state, and that replacing a flat neutron-star mass prior by an astrophysically motivated double-Gaussian prior does not significantly change individual-event radius posteriors.","tokens_in":21132,"tokens_out":18275,"duration_ms":161546,"significance":"If the headline projection holds, the paper delivers a concrete, mass-dependent forecast that an Einstein Telescope plus Cosmic Explorer network will measure neutron-star radii to roughly 5% across most of the mass range, complementing NICER and multimessenger constraints. The study's strengths are its standard, reproducible analysis stack (Bilby, dynesty, relative binning, zero-noise injections), its use of a publicly available 2300-equation-of-state ensemble, and its explicit, candid discussion in Sec. IVB of the bias introduced by the non-uniform ensemble. The mass-dependent treatment of radius uncertainty is a real improvement over projections that quote a single R_1.4 value. The central accuracy claim, however, is validated only for three equations of state that are well represented in the adopted ensemble; quantifying the discretization bias for equations of state in sparsely sampled regions would materially increase the paper's value.","major_comments":[{"comment":"The headline accuracy claim (Abstract: 'resolved ... accurately, across most of the mass range to within ≲5% for both soft and stiff equations of state'; Sec. V) is established only for the three injected equations of state APR4, SLy, and DD2, all of which lie in well-populated regions of the ~2300-model ensemble of Ref. [63]. The paper itself states in Sec. IVB that the non-uniform density of this ensemble 'introduce[s] biases in the estimation of R', and that the sparse representation of equations of state predicting radii above DD2 'will cause a significant underestimation of R' for DD2-like signals. Because the radius posterior in Eq. (17) is an evidence-weighted mixture over this discrete ensemble, a true equation of state in a sparse or excluded region (for example R_1.4 ≈ 10 km or ≈ 15 km in Fig. 2) would cause the posterior to converge to the nearest represented models rather than to the true radius, regardless of detector sensitivity. The current validation does not probe this failure mode. I request either (a) additional ECC injections using held-out equations of state drawn from the sparsely populated regions of Fig. 2 to bound the discretization bias, or (b) an explicit restriction of the 5% accuracy claim to equations of state within the support of the adopted ensemble. Without one of these, the abstract's unqualified accuracy statement goes beyond what Figs. 4 and 5 demonstrate.","section":"Sec. IVB; Eq. (17); Fig. 2"},{"comment":"The equal model prior assumed in Eqs. (15)-(16) assigns identical prior weight to each of the ~2300 discrete equations of state. Since those models are posterior samples from the GW170817 spectral-parameter analysis of Ref. [63], the effective prior over the physical equation-of-state space is proportional to the local sampling density of that posterior, not to any physically motivated measure. The paper acknowledges the resulting problem in Sec. IVB ('Addressing these biases will require either constructing a uniformly distributed EoS set or assigning a probability to each EoS'), but neither solution is implemented and the impact on the reported radii is not quantified. In particular, the ECC results in Figs. 4-5 use the same unweighted mixture, so the claimed 5% accuracy for DD2 cannot be separated from the compensating effects of the ensemble's sparsity above DD2 and of the sampling-density prior. I ask the authors to quantify how strongly the mixture weights in Eq. (17) depend on the sampling density of the Ref. [63] posterior (for example, by reweighting the ensemble uniformly over R_1.4, or by jackknifing the ensemble) and to report whether the ECC 5% claim survives such reweighting.","section":"Eqs. (14)-(17); Sec. IVB"}],"minor_comments":[{"comment":"The sentence stating that zero-noise injection 'ensur[es] that our results represent the ensemble average over many Gaussian noise realizations' overstates what a single noiseless realization provides: a zero-noise run measures the systematic bias of the posterior median under the assumed priors and noise spectral density, whereas realization-to-realization scatter would require averaging over noise draws. Since the zero-noise convention is standard for injection-recovery studies, I suggest rewording rather than changing the analysis.","section":"Sec. IIIB"},{"comment":"The evidence computation fixes the chirp mass to its posterior mean in Eq. (12), but the accuracy of this reduction for the evidence ratios, and hence for the mixture weights in Eq. (17), is not demonstrated. Since M is the best-measured parameter this is likely a small effect; a brief validation (for example, recomputing evidences with M marginalized for a few events) would make the model-selection step watertight.","section":"Sec. IVA, Eqs. (9)-(13)"},{"comment":"The phrase 'to within ≲5%' is ambiguous between the fractional bias of the posterior median (the ΔR statistic of Fig. 5) and the width of the 90% credible interval, and at high masses the 90% intervals in Fig. 5 exceed 5% even when the median does not. Please state explicitly which quantity the headline refers to, and define precisely how the '5% bands' of Fig. 4 ('radius variations expected for a uniform 5% change in the EoS') were computed.","section":"Abstract; Sec. V; Fig. 5"},{"comment":"The Fig. 2 caption states that 'the density of EoSs is highest for radii larger than those predicted by APR4', while Sec. IVB states that 'more EoSs [are] clustered around APR4' and Sec. V says the density 'peaks around the APR4 model'. These statements should be reconciled and the location of the histogram peak in the right panel stated quantitatively.","section":"Fig. 2 caption; Sec. IVB"},{"comment":"The horizontal axis label '% Radius Error [km]' mixes dimensions (a percentage expressed in km); please clarify whether the cumulative quantity is a fractional error in percent or an absolute error in km.","section":"Fig. 6"},{"comment":"BEOMS, the model-selection pipeline that computes the evidences feeding Eq. (17), is cited as 'in prep' and is not publicly described or released. Given that the paper's central results rest on this pipeline, a description of the algorithm or a public code release would substantially improve reproducibility.","section":"Ref. [108]; Sec. IV"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope as a detector-sensitivity projection for equation-of-state science. Two points for the editor's awareness: (1) The equation-of-state ensemble used as the model prior derives from GW170817 itself, so the projected 5% accuracy claim is conditioned on the GW170817-supported region of equation-of-state space; this is discussed in the major comments, but it is also worth attention when judging the headline for a broad readership. (2) The BEOMS pipeline is cited as 'in prep'; if the authors decline to release it, the editor may wish to weigh that against the paper's reproducibility standards. Neither point changes my recommendation of major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. It's a straightforward, competently done simulation forecast: they use Bayesian model selection over ~2300 EoSs from the GW170817 spectral-decomposition set to project neutron star radius measurements for O5, A#, and ET+CE. The headline, that the XG network gets radii to within ~5% across most of the mass range for both soft and stiff EoSs, matches what Huxford et al. and Walker et al. already found with simpler methods. What's new here is the evidence-weighted mixture over the full EoS ensemble, the mass-dependent presentation of uncertainties instead of just quoting R1.4, and the mass-prior robustness check. The mass-prior result - that an astrophysically motivated prior doesn't change single-event radius posteriors - is a useful practical point.\n\nThe paper is transparent about its main weakness. Section IVB and Fig. 2 show the EoS ensemble is non-uniform: dense around APR4 and sparse above DD2, and they state this 'introduces biases in the estimation of R.' That's the correct caveat, but the abstract's 'precisely and accurately ... for both soft and stiff equations of state' doesn't carry it. The validation only injects APR4, SLy, and DD2, all represented in the set. If the true EoS sat in a sparsely populated region (say R1.4 > 15 km), the posterior would be pulled toward the nearest models and the bias could easily exceed 5% regardless of SNR. The stress-test note you sent lands on this point. It doesn't sink the paper, because the authors flag it, but it does mean the 5% accuracy claim is conditional on the EoS prior being a fair discretization of the true EoS space.\n\nMinor issues: zero-noise injections are standard for forecasts, so I don't hold that against them. No code or data release, and the BEOMS pipeline is 'in prep,' which makes independent verification harder than it should be. The KDE and fixed-chirp-mass approximations in the evidence calculation are reasonable but worth a check.\n\nWho should read this: anyone writing the science case for ET or CE, and people working on EoS inference from GWs. It's a solid, if incremental, contribution. Recommendation: send it to peer review. A good referee should ask them to qualify the abstract and, ideally, to run one injection from outside the EoS set, or reweight the set, to show the accuracy claim isn't an artifact of the prior.","headline":"Solid, transparent forecast of XG neutron star radius measurements; the 5% accuracy claim is real but conditional on a non-uniform EoS prior the authors themselves flag.","tokens_in":21700,"tokens_out":3397,"would_cite":true,"duration_ms":28947,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Next-generation gravitational-wave detectors could measure neutron star radii to within about 5 percent across most of the mass range, for both soft and stiff equations of state.","keywords":["neutron star radius","gravitational waves","equation of state","next-generation detectors","Bayesian inference","tidal deformability","Einstein Telescope","Cosmic Explorer"],"falsifier":"Inject a simulated binary neutron star signal whose true equation of state lies in a region the 2,300-model set covers sparsely, for example a very stiff equation of state with a 1.4 solar-mass radius at or above the DD2 value, and check whether the evidence-weighted posterior still recovers the injected radius within 5 percent; if the recovered value is biased low by more than 5 percent despite the high signal-to-noise ratios of the next-generation network, the accuracy claim fails for that equation of state.","tokens_in":20691,"feed_emoji":"🔭","tokens_out":5316,"duration_ms":46483,"temperature":0.7,"pith_summary":"The paper argues that a network of next-generation gravitational-wave observatories, combining the Einstein Telescope with two Cosmic Explorer detectors, can measure the radius of a neutron star to within about 5 percent across most of the allowed mass range, for both soft and stiff equations of state. It reaches this conclusion through a Bayesian model-selection analysis of simulated binary neutron star mergers, using roughly 2,300 equations of state derived from GW170817 as a physically motivated set of hypotheses. The same analysis shows that current-generation detectors (O5 and the A-sharp upgrade) produce radius estimates that are both imprecise and systematically biased, overestimating radii for soft equations of state and underestimating them for stiff ones. The paper also finds that choosing an astrophysically motivated mass prior rather than a flat one does not noticeably change single-event radius inference.","feed_headline":"Next-gen detectors could measure neutron star radii to ~5%","feed_subtitle":"Simulated mergers show Einstein Telescope plus Cosmic Explorer pin down radius across most masses, soft or stiff EoS.","key_machinery":"The load-bearing machinery is the evidence-weighted radius posterior. For each of roughly 2,300 spectral-decomposition equations of state from the GW170817 analysis, the pipeline computes a Bayesian evidence for the hypothesis that that equation of state generated the signal, then draws radius samples from each equation of state's mass-radius relation in proportion to its normalized evidence. The reduced parameter space uses the effective tidal deformability, chirp mass, and symmetric mass ratio, with the chirp mass fixed to its mean to simplify the integrals. Posterior samples for the tidal parameters come from a Bayesian inference library with a nested sampling algorithm, using an inspiral-merger-ringdown waveform model with tidal corrections and relative binning for speed. The evidence weighting is what turns a discrete model set into a continuous-looking radius posterior, and the density of that model set is what ultimately controls the accuracy.","core_discovery":"On its own terms, the central claim is that next-generation detectors will act as cosmic calipers: with a network of one Einstein Telescope and two Cosmic Explorer interferometers, the fractional error in inferred neutron star radius stays below about 5 percent for nearly the entire mass range allowed by the equation of state, regardless of whether the true equation of state is soft (APR4), intermediate (SLy), or stiff (DD2). The radius uncertainty is mass-dependent, so the paper deliberately avoids quoting a single number for the radius of a 1.4 solar mass star and instead presents the uncertainty as a function of injected mass. Using that reference mass alone, the authors argue, misses that the uncertainty grows toward the maximum mass. The accuracy of the result relies on weighting each candidate equation of state by its Bayesian evidence rather than selecting a single best model, and the paper identifies a non-uniform density in the candidate equation-of-state set as the main source of the residual biases.","pith_inferences":["A natural extension the paper leaves implicit is to replace the discrete, non-uniform equation-of-state set with a continuously parameterized or uniformly sampled prior; the authors acknowledge this is needed to remove residual biases, and doing so could make the 5 percent accuracy claim more robust.","The mass-prior result likely does not generalize to low signal-to-noise events or to priors with very different support, since the paper only tests a flat prior against a double-Gaussian prior with the same mass range.","If the 5 percent radius precision is realized, combining individual-event radius posteriors hierarchically across a population of mergers could push the effective constraint on the equation of state well below what a single event achieves, possibly resolving the radius to a few hundred meters.","The same evidence-weighting machinery could be applied to measure radius from post-merger or multi-messenger signals, where the tidal imprint is not the only radius-sensitive feature."],"forward_implications":["If the claim holds, a single Einstein Telescope-Cosmic Explorer network could deliver mass-dependent neutron star radius measurements to better than 5 percent for most of the allowed mass range, without needing to assume the radius is constant across masses.","Radius measurements would become precise enough to distinguish soft, intermediate, and stiff equations of state from gravitational-wave data alone, complementing electromagnetic and multi-messenger constraints.","Current-generation detectors (O5 and the A-sharp upgrade) would be unable to provide unbiased radius estimates on their own, so radius science would have to wait for the next-generation network.","The mass-prior insensitivity result implies that individual-event radius inference is dominated by signal loudness and the equation-of-state model set, not by the choice of mass prior, simplifying population-level pipelines."],"supporting_citations":[{"why":"Supplies the roughly 2,300 spectral-decomposition equations of state that form the model set for evidence calculation.","marker":"[63]"},{"why":"Provides the GW170817 observation whose re-analysis produced the equation-of-state samples used as the prior set.","marker":"[19]"},{"why":"Provides the expected binary neutron star detection rates and the detector network configurations used in the simulation.","marker":"[42]"},{"why":"Supplies the waveform model with tidal corrections used to generate and analyze the simulated signals.","marker":"[79]"},{"why":"Provides the Bayesian parameter-estimation framework and prior settings used for the single-event posteriors.","marker":"[84]"},{"why":"Supplies the nested sampling algorithm used to compute the posterior samples and evidence values.","marker":"[88]"},{"why":"Earlier Fisher-matrix study constraining radius across the mass range, which this paper extends with a Bayesian evidence-weighted approach.","marker":"[60]"},{"why":"Describes the Bayesian model-selection pipeline used to compute equation-of-state evidences.","marker":"[108]"}],"fun_headline_variants":["Gravitational wave calipers measure neutron star radii within 5%","Next-gen detectors shrink neutron star radius uncertainty to 5%","Precision neutron star radii from future gravitational wave observatories","Cosmic Explorer and Einstein Telescope: 5% radius precision for neutron stars"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 5 percent accuracy claim presumes that the discrete set of roughly 2,300 spectral-decomposition equations of state from GW170817 is a fair and sufficiently dense covering of the true equation of state, even though the paper itself shows the set is non-uniform, clustering near APR4 and thinly covering higher-radius stiffer regions.","fun_headline_variants_meta":{"raw":{"variants":["Gravitational wave calipers measure neutron star radii within 5%","Next-gen detectors shrink neutron star radius uncertainty to 5%","Precision neutron star radii from future gravitational wave observatories","Cosmic Explorer and Einstein Telescope: 5% radius precision for neutron stars"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000827,"raw_usage":{"total_tokens":3601,"prompt_tokens":919,"completion_tokens":2682,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":2607}},"tokens_in":535,"tokens_out":2682,"duration_ms":18525,"temperature":1.0,"reasoning_tokens":2607,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T04:39:43.682919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject a simulated binary neutron star signal whose true equation of state lies in a region the 2,300-model set covers sparsely, for example a very stiff equation of state with a 1.4 solar-mass radius at or above the DD2 value, and check whether the evidence-weighted posterior still recovers the injected radius within 5 percent; if the recovered value is biased low by more than 5 percent despite the high signal-to-noise ratios of the next-generation network, the accuracy claim fails for that equation of state.","supporting_citations":[],"review_version":1}