{"id":"e26d841f-c139-4231-823c-1f31eb55c27a","arxiv_id":"2505.21667","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"Combining BAYESTACK radius posteriors for four masses through a common piecewise-polytrope EoS forecasts ~1 km (uniform prior) or ~0.55 km (astrophysical prior) radius constraints from A+ era post-merger signals.","lead":"The paper combines gravitational-wave posteriors for four different neutron-star radii into one equation-of-state estimate, claiming future A+ detectors could measure radii to about 1 km. A generalist reader should care because it is a concrete forecast of how post-merger signals, where current detectors are weak, could sharpen neutron-star physics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eqs. (8)-(9) multiply four BAYESTACK posteriors computed from the same 357-event catalog D as if they were independent likelihoods, and do not divide out the first-stage radius prior; the quoted ~1 km and ~0.55 km intervals are therefore artificially narrow.","rationale":"The load-bearing step is the new statistical combination in Eqs. (8)-(9), because the abstract's quantitative claims (1 km uniform, 0.55 km with astrophysical priors) depend entirely on it. The reader's weakest_assumption is correct: p(R1.X|D) from Eq. (7) is a posterior on the same 357-event catalog, so the product in Eq. (9) is not a product of independent likelihoods. This is not a minor normalization issue. The same data are reused four times, and for the MM Miller/Riley priors the first-stage radius prior is counted again. The paper gives no evidence that the four BAYESTACK posteriors are independent or that the prior was removed. The concrete test above would settle whether the quoted intervals survive a valid joint treatment. I also note the paper's own bias results (Fig. 5 top-right: injected M-R curve outside the 90% band when MM priors are used; Sec. III's statement that 157 of 357 Set C events have chirp mass beyond the available SPH library) are second-order concerns even if Eq. (9) were fixed; these systematics are acknowledged in Sec. V. But the independence/non-division flaw is the load-bearing one because it determines the headline precision and coverage statements. The critique concerns the equations and the stated substitution in Sec. II.B, not the authors' intent.","tokens_in":15367,"tokens_out":5809,"duration_ms":60684,"concrete_test":"Implement the joint inference that Eq. (9) is meant to approximate: use the per-event likelihood factors from Eq. (7) with all four radius slices sharing the same event catalog D, taking L(D|R1.X) = p(R1.X|D)/p(R1.X) as the first-stage likelihood contribution (ideally by moving the R1.X-Υ mapping inside the per-event integral). Run this on Set C with the same priors, and compare the 90% credible widths of R1.2, R1.4, R1.6, R1.8 and the coverage of the injected SFHx M-R curve with Figs. 3 and 5. If the widths grow by more than roughly 20%, or the injected curve moves from inside to outside the 90% band, the advertised ~1 km / ~0.55 km constraints are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Sec. II.B, Eq. (8) writes p(R1.2,...,R1.8,Υ|D) ∝ p(Υ) ∏_{i=1}^4 p(d_i|R1.Xi) p(R1.Xi|Υ)/p(d_i), and the text states that \"The likelihoods p(d_i|R1.Xi) are the posteriors, p(R1.Xi|D), obtained from Eq. (7).\" But in Eq. (7), each p(R1.X|D) is already the posterior from the full event catalog D (e.g., Set C with 357 events). Reusing the same D four times multiplies the identical 357-event likelihood four times, quadrupling the effective sample size and systematically narrowing every reported 90% interval. It also treats the four radius slices as independent probes when they are four different summaries of the same data. Additionally, p(R1.X|D) contains the first-stage radius prior p(R1.X); substituting it for p(d_i|R1.Xi) without dividing by p(R1.X) double-counts that prior. For a constant uniform prior the double counting is absorbed into normalization, but for the MM Miller/Riley priors of Sec. III.B it is not constant, so the \"~0.55 km\" claim from astrophysical-prior runs is invalid as stated. The paper's own bias results (Fig. 5 top-right, where the injected M-R curve falls outside the 90% band under MM priors) are a separate accuracy concern; the product-likelihood error is more fundamental because it breaks the statistical interpretation of the quoted uncertainties.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper extends Criswell et al.'s hierarchical Bayesian inference for post-merger neutron-star signals from a single radius proxy R1.6 to four proxies R1.2, R1.4, R1.6, and R1.8. The authors compute BAYESTACK (BaSt) posteriors for each proxy from simulated event sets (Sets B, C, D) under uniform and astrophysical (MM Miller/Riley) radius priors, then combine the four posteriors through a piecewise-polytrope EoS model to constrain M-R and M-Lambda curves. They report that the combined analysis constrains radii to within ~1 km with uniform priors and ~0.55 km with astrophysical priors for A+ noise, and they discuss biases from the empirical f_peak(M,R1.X) relation and waveform interpolation.","tokens_in":15730,"tokens_out":7077,"duration_ms":84249,"significance":"The paper addresses an important and timely question, and the two-stage idea is a natural extension of Criswell et al. It makes use of public data and code, considers two injected EoSs and several event sets, and is candid about interpolation and empirical-relation biases. If the combination step were valid, the projected ~1 km radius constraints would be a valuable forecast for A+. However, the central statistical step in Eqs. (8)-(9) is not a valid posterior combination; the product of four posteriors from the same event catalog double-counts data and priors. The quantitative conclusions and the main claim are therefore not supported by the analysis as presented.","major_comments":[{"comment":"The combination step is not a valid posterior. In Eq. (8) the four factors p(d_i|R1.X_i) are set equal to the BaSt posteriors p(R1.X_i|D) from Eq. (7), but every one of those posteriors is computed from the same full event catalog (e.g., Set C with 357 events). The four factors are therefore not independent likelihoods, and the product in Eq. (9) reuses the same events four times, inflating the effective information and producing artificially narrow 90% intervals. Moreover, Eq. (7) already includes the first-stage radius prior p(R1.X_i); substituting it for p(d_i|R1.X_i) without dividing by p(R1.X_i) double-counts that prior. The double counting is harmless only for a constant uniform prior, but not for the nonuniform MM Miller/Riley priors used in Sec. III.B. The ~0.55 km claim in the abstract is therefore not statistically justified by the derivation as written. A correct hierarchical likelihood should start from the per-event data and the EoS-parameter mapping, not multiply four posteriors from the same catalog.","section":"II.B, Eqs. (8)-(9)"},{"comment":"The recovery tests do not independently validate the empirical relation. Eq. (5) is fitted to SPH simulated waveforms, and the injected signals in Sets B and C are generated by nearest-neighbor interpolation from the same SPH waveform library, as described in Sec. II.A and Sec. III. Thus the recovery is partly a consistency check that re-derives the fitted relation; it does not test the empirical relation against an independent waveform set. This circularity should be stated explicitly and, ideally, mitigated by testing with an independent simulation code or by quantifying the interpolation-induced bias beyond the qualitative discussion in Sec. IV.A.","section":"II.A and III (injection construction)"},{"comment":"The claim that astrophysical priors yield ~0.55 km constraints is also in tension with Fig. 5: with MM Riley and MM Miller priors the injected M-R curve lies outside the 90% band. The text attributes this to interpolation bias and empirical-relation inaccuracy, but no procedure is given to correct or debias the quoted intervals. Until the method is shown to be calibrated, the narrower credible intervals under informative priors should not be presented as a robust prospect.","section":"IV.C, Fig. 5"}],"minor_comments":[{"comment":"The notation in Eq. (8), rendered as '4Y_{i=1}', should be a product symbol, and the reuse of D for both the full event catalog in Eq. (7) and the four-element set in Eq. (8) is confusing and should be clarified.","section":"II.B, Eq. (8)"},{"comment":"There is a typo: 'GW1710817' should be 'GW170817'.","section":"III.B"},{"comment":"The text refers to both 'MM Miller' and 'MM Riley' priors, but the figures shown use only MM Riley; please clarify which priors are used in each panel and state whether the MM Miller results are omitted or deferred.","section":"III.B and IV.A"},{"comment":"Several figures contain garbled Unicode math symbols and unclear legends (e.g., Fig. 5 mixes 'red error bars', 'red dotted', and 'blue dashed' without a clean key); please regenerate the figures with standard math rendering and explicit panel labels.","section":"Figures 2-5"},{"comment":"Reference [52] is cited as '2024' but the arXiv version is 2023; please update the citation or the year.","section":"Appendix B"}],"recommendation":"reject","confidential_remarks":"The central statistical flaw in Eqs. (8)-(9) is severe: the main numerical results follow from multiplying four posteriors that are all derived from the same event catalog, and the informative-prior results also double-count a nonuniform prior. This cannot be repaired by local edits; the analysis would need to be redone with a proper joint hierarchical likelihood and then recalibrated. I would be open to reconsidering a new submission that fixes this issue and validates the method on independent injections, but the present manuscript is not publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper extends Criswell et al.'s HBI stacking of post-merger signals to four radius slices (1.2, 1.4, 1.6, 1.8 Msun) and then combines those with a piecewise-polytrope EoS to claim ~1 km radius constraints, or ~0.55 km with astrophysical priors. The first part is a legitimate extension and the injection study is careful. But the second-stage combination has a load-bearing statistical flaw: Eq. (9) multiplies four BAYESTACK posteriors that are all computed from the same 357-event catalog D. There is no split of events by mass slice, so the same data are used four times. That quadruples the effective sample size and narrows every 90% interval. The paper also substitutes the posteriors for the likelihoods in Eq. (8) without dividing out the first-stage radius prior; with the Miller/Riley non-uniform priors this double-counts the prior and corrupts the quoted 0.55 km number. The ~1 km uniform-prior claim is too narrow for the same reason, though the prior issue is less severe there.\n\nOn top of that, the empirical relation fpeak(M, R1.X) is fitted to the same SPH waveform library used to generate the injected events via nearest-neighbor interpolation, so the recovery partially re-derives the relation rather than testing it independently. The paper's own Fig. 5 shows the injected M-R curve falling outside the 90% band for both MM priors; the authors acknowledge this bias but attribute it to waveform interpolation and empirical-relation validity, not to the product-likelihood problem.\n\nWhat's genuinely good: the BAYESTACK posteriors for multiple radii are a useful extension, the study is transparent about the empirical relation's shortcomings, and the code and data are public. With a corrected hierarchical scheme—building a joint likelihood over shared events, or splitting events, and dividing out the first-stage priors—the method could be promising.\n\nWho should read it: someone working on post-merger EoS constraints, especially to see the pitfalls of stacking posteriors. I wouldn't cite it in its current form. A serious referee should see it because the idea is real and the fix is well-defined, but I agree with the REJECT verdict until the statistics are redone.\n\nRecommendation: send to review, but expect heavy revision; the central statistical step must be fixed.","headline":"The multi-radius extension is clearly described, but the combination step reuses the same events four times, so the headline radius constraints are not statistically supported.","tokens_in":16299,"tokens_out":3627,"would_cite":false,"duration_ms":38390,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["97.60.Jd","04.30.-w"],"model":"deepseek-v4-flash","headline":"This paper claims that combining post-merger gravitational-wave signals from many binary neutron star mergers through a two-stage hierarchical inference can constrain neutron star radii to within about 1 km, and to about 0.55 km when…","keywords":["neutron star equation of state","post-merger gravitational waves","hierarchical Bayesian inference","R1.X radius","piecewise polytrope","A+ detector","binary neutron star mergers","empirical relations"],"falsifier":"Inject many simulated catalogs with a known equation of state, run the two-stage inference on disjoint halves of the events for different radius slices, and check whether the true radius lies inside the 90% interval as often as claimed; a coverage deficit would show that the independence assumption is wrong.","tokens_in":15121,"feed_emoji":"🌌","tokens_out":4208,"duration_ms":41274,"temperature":0.7,"pith_summary":"This paper argues that the next generation of gravitational-wave detectors can turn a large catalog of binary neutron star post-merger signals into tight measurements of the neutron star radius. By extending a hierarchical Bayesian inference method from a single radius proxy, R1.6, to four radii at 1.2, 1.4, 1.6, and 1.8 solar masses, and then imposing the condition that all four follow the same equation of state, the authors show that injected radii can be recovered to within about 1 km with uniform priors and about 0.55 km with astrophysical priors in A+ noise. The tighter radius would directly constrain the equation of state of dense matter, which is currently poorly known.","feed_headline":"Post-merger signals could pin neutron star radii to 1 km","feed_subtitle":"Combining four radius measurements through one equation of state tightens bounds to about a kilometer.","key_machinery":"The load-bearing identity is the empirical relation $f_{\\rm peak} = b_0 + b_1 M + b_2 M^2 + b_3 R_{1.X} M + b_4 R_{1.X} M^2 + b_5 R_{1.X}^2 M$, which links the post-merger peak frequency to chirp mass and a radius proxy. This relation feeds the first-stage BAYESTACK posterior $p(R_{1.X}|D)$ for each of four masses; the second stage multiplies these four posteriors, with a delta-function prior mapping $R_{1.X}$ to the radius computed from the piecewise polytrope parameters via the Tolman-Oppenheimer-Volkhoff equations. The product is the equation-of-state posterior that is then sliced to give joint radius constraints.","core_discovery":"The central claim is that the two-stage hierarchical inference—first obtaining per-mass radius posteriors from ensembles of post-merger detections, then combining them through a piecewise polytropic equation of state—yields radius constraints of about 1 km at 90% credibility for a four-year A+ catalog of 357 events. Using priors informed by NICER and GW170817 tightens this to roughly 0.55 km. The paper also finds that the injected equation of state is recovered at the boundary of the 90% bound for the mass-tidal deformability curve, but shows a bias in the mass-radius curve that it attributes to waveform interpolation and to the limited validity of the empirical frequency-radius relation at high masses.","pith_inferences":["A fully joint inference that processes all events simultaneously in one hierarchical model, rather than combining four separately computed posteriors, would test the independence assumption and could either validate or widen the reported intervals.","The per-radius 90% intervals reported assume the product of four posteriors is a valid joint posterior; if the same events dominate all four slices, the true uncertainty may be larger than stated.","The empirical-relation bias at high mass suggests the actual constraining power of future catalogs may depend more on improving fpeak models than on detector sensitivity alone.","If third-generation detectors deliver many more events, the same two-stage scheme could probe temperature-dependent effects such as phase transitions, as the paper hints."],"forward_implications":["If the A+ era produces roughly 357 post-merger detections, radius bounds near 1 km become achievable without external priors.","With NICER/GW170817-informed priors, the same catalog yields ~0.55 km radius bounds, sharp enough to distinguish many proposed equations of state.","Because the four radius slices are tied to a single equation of state, the method also constrains the mass-radius and mass-tidal-deformability curves across the 1.2–1.8 solar mass range.","The identified biases at high chirp mass imply that improving the empirical relation would remove the main systematic limitation."],"supporting_citations":[{"why":"Supplies the original hierarchical Bayesian inference formulation for R1.6 that this work extends to four radii.","marker":"[1]"},{"why":"Provides the simulation data and empirical relation connecting fpeak with chirp mass and R1.X.","marker":"[39]"},{"why":"The BAYESTACK package used to compute the first-stage radius posteriors from Eq. (7).","marker":"[40]"},{"why":"The catalog of simulated BNS events used for the injection study.","marker":"[41]"},{"why":"Defines the piecewise polytrope equation-of-state model adopted for the second-stage EoS inference.","marker":"[50]"},{"why":"Provides the astrophysical mass-radius priors derived from NICER and GW170817 used as informed priors.","marker":"[52]"},{"why":"Public fpeak posterior datasets from BayesWave used as inputs to the BAYESTACK inference.","marker":"[49]"},{"why":"Supplies the A+ and advanced LIGO/Virgo noise curves used for the injection sets.","marker":"[38]"}],"fun_headline_variants":["Post-merger waves tie neutron star radii to ~1 km","Combining post-merger signals yields 1-km radius bounds","Hierarchical inference shrinks neutron star radius error to 1 km","Post-merger data shrinks neutron star radius constraints to ~1 km"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported bounds assume that the radius posteriors from the four mass slices are statistically independent even though they are all computed from the same catalog of events, so any shared systematic error would make the combined intervals too narrow.","fun_headline_variants_meta":{"raw":{"variants":["Post-merger waves tie neutron star radii to ~1 km","Combining post-merger signals yields 1-km radius bounds","Hierarchical inference shrinks neutron star radius error to 1 km","Post-merger data shrinks neutron star radius constraints to ~1 km"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000808,"raw_usage":{"total_tokens":3569,"prompt_tokens":988,"completion_tokens":2581,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":2506}},"tokens_in":604,"tokens_out":2581,"duration_ms":18881,"temperature":1.0,"reasoning_tokens":2506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:26:33.906589+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject many simulated catalogs with a known equation of state, run the two-stage inference on disjoint halves of the events for different radius slices, and check whether the true radius lies inside the 90% interval as often as claimed; a coverage deficit would show that the independence assumption is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the original hierarchical Bayesian inference formulation for R1.6 that this work extends to four radii."},{"cited_title":"Vretinaris, N","cited_arxiv_id":null,"evidence_quote":"Provides the simulation data and empirical relation connecting fpeak with chirp mass and R1.X."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The BAYESTACK package used to compute the first-stage radius posteriors from Eq. (7)."},{"cited_title":"Petrov, L","cited_arxiv_id":null,"evidence_quote":"The catalog of simulated BNS events used for the injection study."},{"cited_title":"Framework for Multi-messenger Inference from Neutron Stars: Combining Nuclear Theory Priors","cited_arxiv_id":"2306.04386","evidence_quote":"Provides the astrophysical mass-radius priors derived from NICER and GW170817 used as informed priors."},{"cited_title":"Douchin and P","cited_arxiv_id":null,"evidence_quote":"Public fpeak posterior datasets from BayesWave used as inputs to the BAYESTACK inference."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the A+ and advanced LIGO/Virgo noise curves used for the injection sets."}],"review_version":1}