{"id":"f3691264-0b88-4cad-a02b-36f026621f60","arxiv_id":"2501.11585","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 3,000-injection simulation of A+ era binary neutron star observations finds systematic tidal deformability biases and projects only marginal equation of state discrimination from about 25 detections.","lead":"This simulation study injects 3,000 binary neutron star merger signals into simulated A+ detector noise and recovers them with full parameter estimation, finding that tidal deformability estimates are systematically biased by mass. The authors use the recovered posteriors to project how well the neutron star equation of state can be constrained in the next observing run, concluding that only marginal EOS discrimination is possible and that corrections will be needed.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of marginal EOS distinguishability is asserted but never quantified; no Bayes factor, posterior-probability, or separation metric is reported.","rationale":"Reader's CONDITIONAL verdict is appropriate; the work is a substantial simulation study but the headline requires a condition. However, I disagree with prioritizing the same-waveform-model injection issue as the weakest assumption. That is a valid external-validity caveat, since real waveforms may differ from IMRPhenomPv2 NRTidalv2, but it does not speak to whether the paper's own analysis demonstrates the stated conclusion. The prior-shrinkage mechanism is also a real interpretive flaw: the reported Lambda bias pattern (underestimate at low mass, overestimate at high mass) is exactly what a wide uniform Lambda prior with weak likelihoods produces, so the 'systematic biases' are at least partly prior effects. But even that does not settle whether the corrected EOS posteriors actually separate the three EOSs. The missing quantitative metric is the single load-bearing gap because it is necessary for the central claim, it is straightforward to supply from the existing samples, and it determines whether the abstract is accurate. I therefore recommend keeping the conditional verdict, with the condition being a quantitative distinguishability analysis.","tokens_in":16448,"tokens_out":11583,"duration_ms":132173,"concrete_test":"Re-analyze the existing event groups (10, 20, 25, and 30 events per group) with an explicit discrimination metric: for each group, compute the LWP posterior probability of each of the three GP-EOS hypotheses (or a Bayes factor between each pair), using the same PSR-conditioned prior, and report the fraction of realizations in which the injected EOS receives the highest posterior probability and the median log Bayes factor for the true versus closest alternative. If the true EOS is favored in only a bare majority of realizations, or the log Bayes factor is below a stated threshold, the 'marginally distinguished' claim needs to be qualified; if it is favored in nearly all realizations, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is that the paper's central scientific result is not actually demonstrated. The abstract and introduction assert that the three equations of state can be marginally distinguished with necessary corrections, but the results section never defines or computes a distinguishability statistic. Figures 3-5 show recovered median EOS curves and 90% credible intervals, and the text describes where the medians are biased, but it does not report whether the posteriors for hqc18, sly230a, and mpa1 are separated at a stated confidence level, nor the frequency with which the true EOS is recovered. A 'marginal distinction' could range from barely non-overlapping credible bands to clearly separated medians, and the paper provides no criterion. This matters because the conclusion that the A+ era will provide only marginal EOS discrimination is the paper's headline; without a quantitative measure, the claim is supported only by visual inspection of selected groups.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a large-scale simulation study of binary neutron star (BNS) gravitational-wave detections at the upcoming A+ sensitivity of the LIGO-Virgo-KAGRA network. The authors perform full parameter estimation on 3,000 injected signals (1,000 for each of three equations of state: hqc18, sly230a, and mpa1) using reduced-order quadrature in Bilby, then carry out hierarchical Bayesian EOS inference with the LWP package. They report systematic tidal deformability biases (overestimation at higher masses, underestimation at lower masses) and claim that, with 'necessary corrections,' the three EOSs can be marginally distinguished in the A+ era with roughly 25 detections over three years. They conclude that precision EOS constraints must wait for next-generation detectors.","tokens_in":16666,"tokens_out":6421,"duration_ms":69512,"significance":"If the central claim were demonstrated, this would be an important quantitative projection for the O5 observing era, showing that EOS discrimination remains marginal and that systematic biases must be corrected even before next-generation detectors. The computational effort is substantial and the bias trends are of community interest. The paper leverages validated public tools (Bilby, LWP, ROQ bases) and provides a realistic population simulation. However, the headline result is not actually quantified, and the proposed corrections are never shown to improve the inference, which severely limits the paper's current contribution.","major_comments":[{"comment":"The central claim that the three EOSs 'can be marginally distinguished with necessary corrections' is never quantified. No Bayes factor, posterior probability, overlap fraction, or classification accuracy is reported. The paper should compute a measure of separation between the three recovered EOS posteriors (e.g., pairwise Bayes factors, the posterior weight of the true EOS, or the fraction of groups where the 90% credible interval excludes the median of another EOS) and state the criterion for 'marginal.' Without this, the headline claim is supported only by visual inspection of Figures 3-5.","section":"Sec. 1 and Sec. 4.3 (Fig. 4)"},{"comment":"The abstract states that the work 'quantif[ies] the needed corrections,' and Sec. 5 claims that this 'method can enhance the accuracy of EOS constraints,' but no corrected EOS inference is shown. Figure 6 reports the relative percentage bias as a function of mass, but the authors never apply a correction (e.g., a mass-dependent calibration or a reweighting of the posteriors) to the hierarchical inference and demonstrate that it improves recovery of the true EOS. The phrase 'with necessary corrections' therefore needs an explicit demonstration of the correction's effect on the final EOS constraints.","section":"Abstract and Sec. 5"},{"comment":"The systematic bias in tidal deformability is attributed to SNR and EOS steepness (Sec. 4.5), but the role of the PE prior is not analyzed. The uniform prior on Lambda_1 and Lambda_2 in [0, 5000] (Table 1) interacts with the steep Lambda(m) relation and can itself produce a shrinkage pattern resembling the reported underestimation at low mass and overestimation at high mass. The authors should test this hypothesis, for instance by comparing the recovered posteriors to the prior predictive distribution or by repeating a subset of PE runs with a different prior.","section":"Sec. 3 (Fig. 2) and Table 1"},{"comment":"The detection criteria are inconsistent between sections. Sec. 2.1 requires network SNR > 11.2 and per-detector SNR > 5, while Sec. 2.2 requires network SNR > 11.2 and single-detector SNR > 4 in at least two detectors. These different thresholds affect which of the 3,000 injections enter the analyzed sample (617, 635, and 676 events for hqc18, sly230a, and mpa1). The authors should reconcile these definitions and verify that the sample selection and the resulting bias trends are robust to the choice of threshold.","section":"Sec. 2.1 vs. Sec. 2.2"},{"comment":"The bias estimates and derived corrections are conditional on injecting and recovering with the same waveform model, IMRPhenomPv2 NRTidalv2. The pipeline therefore cannot detect waveform-model systematic errors, yet the paper proposes these corrections for real O5 data. The authors should either test with an alternative waveform model (e.g., a different tidal approximant or a model with dynamical tides) or clearly state that the corrections are only valid if the injected model is an accurate description of real BNS signals.","section":"Secs. 2.1 and 2.2"}],"minor_comments":[{"comment":"There are numerous typos and grammatical errors, including 'repreent' (Fig. 3 caption), 'geoup' (Fig. 5 caption), 'the underlying biases is still present' (Sec. 4.2), and inconsistent formatting of 'L WP' and 'L VK' throughout.","section":"Throughout"},{"comment":"The reference 'Shoemaker et al. 2024' appears only as a URL in the bibliography; it should be formatted consistently with the other references.","section":"Bibliography"},{"comment":"The caption says 'each geoup' and mentions 'gray shaded regions' while the figure panels use colored solid lines; ensure the caption matches the figure content.","section":"Fig. 5 caption"},{"comment":"The title 'Effect of Chirp Mass and EOS Softness' is confusing because mpa1 is described as stiff in Sec. 2.1 but called 'the steepest one' here; define 'softness' in this context to avoid ambiguity.","section":"Sec. 4.4"},{"comment":"The paper does not explicitly state that the Gaussian-process EOS prior from Legred et al. (2021) is conditioned only on pulsar mass measurements and not on GW data; adding this statement would preempt concerns about circular inference.","section":"Sec. 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper's computational pipeline is a significant resource, and the bias characterization is potentially useful. However, the central claim of 'marginal distinguishability' is not supported by any quantitative metric, and the proposed corrections are never validated. I recommend major revision rather than rejection because the missing analysis is within the scope of a revision. One additional point for the editor: the authorship overlap with the developers of LWP and the EOS prior (Landry is a coauthor) is normal in this field, but the paper would benefit from an explicit statement that the prior is externally validated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nIf you are planning to use the A+ era for EOS projection, this is the paper to start from. It is the largest end-to-end simulation I know of at A+ sensitivity: 3,000 full ROQ-accelerated PE runs across three EOSs, followed by hierarchical EOS inference with a Gaussian-process prior. The scale is real, and the execution is credible: the population is realistic, the detection cuts are stated, and the LWP-based inference handles selection effects in the standard way. The mass-dependent tidal bias they report — overestimation at high mass, underestimation at low mass — is a useful result on its own.\n\nThe soft spots are in the conclusions. The headline claim that the three EOSs can be 'marginally distinguished' is never quantified. No Bayes factor, no posterior separation metric, no probability that the true EOS is preferred. Figures 3–5 show medians and credible bands, but visual inspection is doing the work. That is the load-bearing claim, and it needs a number.\n\nThe bias attribution also skips the obvious mechanism. The PE prior on tidal deformability is uniform over [0,5000]. That prior pulls low-Lambda events upward and high-Lambda events downward, which is exactly the pattern in Figure 6. The paper instead floats waveform modeling limitations. Prior shrinkage is not the whole story, but it should be controlled for before invoking other systematics.\n\nTwo more caveats. Injection and recovery use the same waveform model, IMRPhenomPv2 NRTidalv2, so the 'necessary corrections' derived from these simulations may not transfer to real data if the real waveforms differ. And there is a minor internal inconsistency: the intro says a detection requires SNR > 5 in each detector, while Section 2.2 says SNR > 4 in at least two detectors.\n\nThe circularity concern raised in some passes is not well-founded: the Legred et al. prior is conditioned on pulsar masses, not on the simulated GW data, and the inference is not circular in the main pipeline.\n\nWho this is for: anyone writing an O5-era EOS projection or planning a BNS population analysis. It deserves a serious referee, and I would send it with the request that the authors add a proper distinguishability statistic and an honest discussion of prior shrinkage. The missing pieces are fixable, and the simulated dataset is worth preserving.\n\nCandidly,\n[You]","headline":"Large-scale A+ era EOS projection with real computational value, but the headline 'marginal distinguishability' claim is asserted, not demonstrated, and the bias interpretation glosses over prior shrinkage.","tokens_in":17141,"tokens_out":5023,"would_cite":true,"duration_ms":48173,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper projects that A+ era gravitational-wave detections will distinguish neutron-star equations of state only marginally, and only after correcting a systematic tidal deformability bias.","keywords":["neutron star equation of state","tidal deformability","binary neutron star mergers","gravitational wave parameter estimation","A+ detector era","Bayesian EOS inference","systematic bias","reduced order quadrature"],"falsifier":"Run the same 3,000-signal pipeline using a different waveform model for recovery, or with waveforms that add effects the model omits, and check whether the bias map changes; a significantly different map would show that the corrections are not robust. Alternatively, compare the corrected $\\Lambda(m)$ recovered from real O5 detections with independent radius measurements from X-ray and radio observations; a mismatch in the same bias pattern would falsify the calibration.","tokens_in":16303,"feed_emoji":"🔭","tokens_out":6905,"duration_ms":71357,"temperature":0.7,"pith_summary":"The paper projects how well the upgraded LIGO-Virgo-KAGRA network will constrain the neutron-star equation of state (EOS) in its next long observing run. By simulating 3,000 binary neutron-star mergers under three candidate EOS models, the authors estimate that the network will detect roughly 25 such events over three years, and that those events will separate the three models only marginally. The separation is possible only after correcting a systematic bias that makes the inferred tidal deformability too low for low-mass neutron stars and too high for high-mass ones. The paper concludes that precise EOS constraints from gravitational waves must wait for next-generation detectors.","feed_headline":"A+ detectors will barely separate neutron-star equations of state","feed_subtitle":"3,000 simulated mergers: ~25 detections will only marginally separate the models, and only after bias fixes.","key_machinery":"The load-bearing object is the tidal deformability--mass relation $\\Lambda(m)$ for each EOS, computed by solving the Tolman--Oppenheimer--Volkoff equations together with the quadrupolar tidal deformation equation. The argument runs through a three-stage pipeline: population simulation that assigns each neutron star a $\\Lambda$ from its mass and the chosen EOS; parameter estimation of each simulated signal to produce per-event posteriors on masses and tidal deformabilities; and hierarchical Bayesian inference that combines these posteriors with a Gaussian-process prior on the EOS to produce combined constraints. Reduced-order quadrature makes the large simulation tractable by speeding up likelihood evaluations by a factor of hundreds. The comparison between injected and recovered $\\Lambda(m)$ curves, expressed as relative percentage error as a function of mass, is what exposes the systematic bias and supports the correction strategy.","core_discovery":"Using full parameter estimation on 3,000 simulated binary neutron-star signals at A+ sensitivity, with tidal deformabilities assigned from three EOS models spanning soft, average, and stiff behavior, the paper finds that the three EOSs can be marginally distinguished once systematic biases are corrected. The central quantitative result is a mass-dependent bias in the recovered tidal deformability $\\Lambda$: it is underestimated by roughly 25\\--30\\% at low masses and overestimated by up to about 125\\% at high masses, with the crossover mass set by the steepness of the true EOS. Because this bias persists when more events are added and even when only the loudest events are used, the paper argues that it is systematic rather than statistical, and that simulation-based corrections are required to recover the true $\\Lambda(m)$ relation. Under the projected three-year O5 campaign of roughly 25 detections, the hierarchy among the three EOSs is recovered only marginally, and the paper treats this as evidence that precision EOS work will need next-generation detectors.","pith_inferences":["The corrections are calibrated under the assumption that the same waveform model generates and recovers the signals; if real waveforms include effects absent from that model, the bias map could shift and the corrections would need to be re-derived.","The observed underestimation at low mass and overestimation at high mass resembles a shrinkage or regression-to-the-mean pattern, so the bias may generalize beyond the three specific EOSs studied and could affect real binary neutron-star analyses in O5.","The crossover mass where the bias flips sign could itself be a measurable EOS diagnostic; locating it in real data would test whether the simulation-based corrections are transferring correctly.","The saturation at about ten loud events suggests a practical observing strategy: rather than simply accumulating detections, the collaboration could prioritize a small set of loud events and build a dedicated calibration bank from simulations."],"forward_implications":["A three-year A+ campaign will leave the soft, average, and stiff EOS models distinguishable only at the margin, not decisively separated.","Systematic tidal biases will not average away with more detections; they must be corrected using large simulated calibration sets.","Including more than about ten high-SNR events adds little to EOS recovery under A+ sensitivity, so event selection matters more than event count.","Reliable EOS slopes require very loud events with network SNR above 35, which will be rare in the A+ era.","Precision EOS constraints from gravitational waves will likely require next-generation detectors or the proposed A# and Voyager upgrades."],"supporting_citations":[{"why":"Supplies the hierarchical Bayesian method for joint EOS inference from multiple gravitational-wave observations.","marker":"Landry et al. 2020"},{"why":"Provides the Gaussian-process EOS prior used in the inference, conditioned on the two heaviest known pulsars.","marker":"Legred et al. 2021"},{"why":"Defines the IMRPhenomPv2 NRTidalv2 waveform model used for both injection and recovery.","marker":"Dietrich et al. 2019"},{"why":"Introduces the reduced-order-quadrature technique that accelerates parameter estimation.","marker":"Canizares et al. 2015"},{"why":"Extends the reduced-order-quadrature approach used to make the 3,000 parameter-estimation runs tractable.","marker":"Qi & Raymond 2021"},{"why":"Supplies the binary neutron-star merger rate used to translate detections into expected event numbers.","marker":"Abbott et al. 2023"},{"why":"Provides the binary neutron-star population model for masses and spins used in the injections.","marker":"Landry & Read 2021"},{"why":"Provides the reduced-order-quadrature bases specific to the IMRPhenomPv2 NRTidalv2 waveform model.","marker":"Morisaki et al. 2023"}],"fun_headline_variants":["Mass-dependent tidal bias skews neutron-star EOS fits","A+ era detectors barely separate neutron-star EOS models","Bias corrections needed for marginal EOS separation in A+ era","3,000 simulated mergers: ~25 detections barely tell EOS apart"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole projection assumes that the waveform model used to create and to analyze the signals is correct; if real binary neutron-star waveforms differ from it, the measured biases and the corrections built from them may not apply.","fun_headline_variants_meta":{"raw":{"variants":["Mass-dependent tidal bias skews neutron-star EOS fits","A+ era detectors barely separate neutron-star EOS models","Bias corrections needed for marginal EOS separation in A+ era","3,000 simulated mergers: ~25 detections barely tell EOS apart"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000566,"raw_usage":{"total_tokens":2663,"prompt_tokens":907,"completion_tokens":1756,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":1683}},"tokens_in":523,"tokens_out":1756,"duration_ms":15083,"temperature":1.0,"reasoning_tokens":1683,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:04:26.642386+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 3,000-signal pipeline using a different waveform model for recovery, or with waveforms that add effects the model omits, and check whether the bias map changes; a significantly different map would show that the corrections are not robust. Alternatively, compare the corrected $\\Lambda(m)$ recovered from real O5 detections with independent radius measurements from X-ray and radio observations; a mismatch in the same bias pattern would falsify the calibration.","supporting_citations":[],"review_version":1}