REVIEW 3 major objections 4 minor 35 references
How Loud Must a Neutron-Star Merger Be to Reveal Its Equation of State?
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The evidence distinguishing two neutron-star equations of state grows as the square of (tidal contrast × SNR), and a three-configuration calibration predicts a new binary's decisive loudness to 0.3%.
desk verdict Real, reproducible scaling study whose headline out-of-sample prediction works because of an SNR-dependent coefficient, not a transferable prefactor; worth publishing after the claim is reframed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Occam-factor (Laplace) expansion of the Bayesian evidence, which the paper reduces to $\Delta\log Z \approx \frac{1}{2}\|\delta h_{\perp}\|^{2}$ for the zero-noise injections used here. Here $\delta h_{\perp}$ is the noise-weighted waveform mismatch between the correct-EOS template and the wrong-EOS template after the wrong arm's nuisance parameters (mass ratio, spins, coalescence time, distance, orientation) have been optimized out; the subscript $\perp$ records that the raw EOS-induced waveform difference is projected orthogonal to those nuisance directions. Because the tidal phase in the tidally corrected waveform model used in the study is, at leading post-Newtonian order, linear in the mass-weighted tidal deformability $\tilde{\Lambda}$, the mismatch $\delta\psi_{\perp}(f) \approx \Delta\tilde{\Lambda}\,g_{1}(f)$ factors out of the integral, giving $\Delta\log Z \approx C(\Delta\tilde{\Lambda}\,\mathrm{SNR})^{2}$ with $C = I_{\perp}/2$ and $I_{\perp}$ a distance-independent spectral integral over the noise-weighted template power. This single identity converts model selection---normally a nested-sampling computation---into a closed-form estimate, and it also identifies the failure mode: when the wrong arm's posterior pins against a prior boundary (e.g., mass ratio drifting to its upper limit), the projection $\delta h_{\perp}$ changes and the quadratic scaling can break.
What would settle it
Rerun the failing spot-check configuration (a Skyrme EOS injected, a relativistic mean-field EOS recovered, at component masses 1.8 and 1.2 solar masses) with the mass-ratio prior widened beyond the value the wrong arm drifts to; if $\Delta\log Z$ returns to the predicted $C(\Delta\tilde{\Lambda}\,\mathrm{SNR})^{2}$ value, the quadratic law survives and the suppression is a prior-boundary artifact, while if it remains near 44% of prediction, the leading-order scaling itself fails for that configuration.
Extended reading notes
Core claim
The paper's central claim is that, for zero-noise simulated binary neutron star signals, the Bayesian evidence difference between a recovery model whose template EOS matches the injected EOS and one whose template EOS does not follows $\Delta\log Z = A\,\mathrm{SNR}^{n}$ with $n \simeq 1.74$--$1.95$ across four configurations spanning tidal-deformability contrasts of about 51% and 28%, both EOS-role assignments, and two binary mass points. More strongly, the evidence gap is, to leading order, $\Delta\log Z \approx C(\Delta\tilde{\Lambda}\,\mathrm{SNR})^{2}$, a quadratic law derived from a Laplace/Occam-factor expansion of the evidence integral in which the waveform mismatch between correct and wrong templates is linear in the tidal-deformability contrast at leading post-Newtonian order. The prefactor $C$ is not universal---it depends on the binary masses, spins, sky position, detector network, and priors---but when calibrated on three curves it transferred to a fourth, independent mass point: the pre-registered prediction of $\mathrm{SNR} \approx 50.8$ for decisive evidence matched the measured $50.7$, a $0.3\%$ discrepancy. The paper also reports that four of five additional spot-checks across other EOS pairs (including a Skyrme functional against relativistic mean-field models) matched the analytic prediction to within about 15%, while the one outlier, coincident with strong mass-ratio compensation, measured only 44% of the predicted evidence gap, showing the simple scaling is a leading-order estimate that can be suppressed when the wrong template's nuisance parameters absorb the mismatch.
Load-bearing premise
The scaling law holds only if the phase difference between the correct and wrong templates is, to leading order, linear in the tidal-deformability contrast and if the wrong template's nuisance parameters do not absorb enough of that difference to change the projected mismatch; when mass ratio, spin, or coalescence time can compensate the EOS error---as in one of the five spot-checks---the quadratic law can overpredict the evidence gap by more than a factor of two.
Editorial extensions
If this is right
- The SNR threshold for decisive evidence ($\Delta\log Z = 5$) EOS discrimination can be predicted analytically for an unmeasured configuration once $C$ is calibrated on a few configurations; the Curve 4 test confirms this at the $0.3\%$ level.
- For the fiducial configuration with about a 51% tidal-deformability contrast, decisive discrimination occurs at network SNR $\approx 56$ (strong evidence at $\approx 38$ and substantial at $\approx 23$); reducing the contrast to about 28% pushes the decisive threshold to $\approx 89$.
- The calibrated scaling implies that 'tidal deformability doppelgänger' EOS pairs with $\Delta\tilde{\Lambda} \approx 10$--$30$ would require SNR $\approx 900$--$2800$ for decisive discrimination, reachable only for rare, very nearby events rather than typical detections.
- Because $C$ depends on the binary's masses, spins, and the detector network, the operational rule is to recalibrate $C$ for each new EOS pair and mass regime rather than to reuse a single value across an ensemble.
- The one failing spot-check (44% of predicted evidence, coincident with mass-ratio compensation) shows that parameter compensation can suppress the evidence gap by more than a factor of two; the quadratic law is a leading-order benchmark, not a universal guarantee.
Reading between the lines
- The near-quadratic law implies diminishing returns on detector sensitivity: doubling the SNR only quadruples the evidence, so pushing a marginal event from SNR 30 to 60 buys the same discrimination gain as going from 60 to 120; observational planning should therefore target events already near the threshold, not only the loudest.
- The clean transfer of $C$ across one mass change hints that $C$ may approximately factorize into a mass-dependent piece and a network-dependent piece; if that factorization holds across a denser mass grid, $C$ could be precomputed per network and stored as a lookup table, making the closed-form estimate an online planning tool.
- The failure mode identified in the spot-check suggests building a map of 'compensation valleys'---binary configurations where a wrong EOS can hide its mismatch by shifting mass ratio, spin, or time to prior boundaries---would bound the regime of validity of the quadratic law; the paper's wall-pinning analysis provides the first entries of such a map.
- The zero-noise, fixed-sky assumption is checked by a single three-seed comparison showing a shift within the expected noise scatter; a full noise-averaged, free-sky campaign across all configurations would test whether the $0.3\%$ transferability survives realistic conditions, which is the natural next step the paper leaves open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper uses nested-sampling Bayesian parameter estimation on zero-noise simulated binary neutron star signals in an ET+2CE network to compute the evidence difference ΔlogZ between recovery with the correct and with a wrong EOS template. It reports a near-quadratic power-law scaling ΔlogZ = A SNR^n with n ≈ 1.74–1.95 across four configurations spanning two tidal-deformability contrasts, an EOS true/wrong role swap, and two binary mass points. A Laplace/Occam expansion is used to derive the leading-order form ΔlogZ ≈ C(ΔΛ̃ SNR)^2; C is calibrated from the highest-SNR anchor points of Curves 1–3, and a pre-registered prediction for the decisive-evidence SNR of the unseen Curve 4 agrees with the measured 50.7 to 0.3%. The Supplemental Material reports anchor-point sensitivity, a densification of Curve 1, five additional spot-checks (one failing at 44% of the prediction), and a free-sky/real-noise robustness check.
Significance. If the claimed scaling were established, it would give a practical closed-form estimate for the SNR at which a third-generation detector network can distinguish two candidate neutron-star equations of state, which is useful for observational planning and event prioritization. The paper's strengths are its pre-registered out-of-sample test, public code and data, and unusually transparent handling of caveats: anchor leverage, the narrow SNR range without the anchor, the zero-noise/fixed-sky simplification, and the one failing spot-check are all explicitly disclosed. The prefactor-drift issue discussed below, however, means the central 'transferable prefactor' claim is not yet supported in its current form.
major comments (3)
- [Supplemental Material, Eq. (4) and Tables S6–S7] The calibrated prefactor C is not constant along a curve, so the Curve 4 prediction does not validate transfer of the high-SNR-calibrated prefactor. For Curve 1, K = ΔlogZ/SNR² falls monotonically from 1.005×10⁻³ at SNR = 117.41 to 9.012×10⁻⁴ at the d_L = 42 Mpc anchor (Table S7). For Curve 4, the anchor gives K = 12359.66/3192.43² = 1.213×10⁻³, hence C4 = K/(518.91)² ≈ 4.50×10⁻⁹, which is 37% below the calibrated mean C̄ = 7.19×10⁻⁹ (Table S1). Inserting C4 into the quadratic formula gives a decisive SNR of about 64, not the measured 50.7. The 0.3% agreement at threshold therefore comes from Curve 4's low-SNR effective coefficient being close to C̄, not from the prefactor transferring from the high-SNR calibration. The claim that 'the coefficient transfers between configurations' (main text Results) needs to be replaced by an explicit statement that C is SNR-dependent and that the prediction is for an effective low-SNR coefficient.
- [Supplemental Material, Table S8] The claimed common exponent n ≈ 1.74–1.95 is not robust without the anchor. Dropping the single highest-SNR point lowers n to 1.46 (Curve 1, 8-point), 1.20 (Curve 2), 1.75 (Curve 3), and 1.65 (Curve 4); the densification repair was performed only for Curve 1. Thus the abstract's 'common scaling' rests on one anchor point per curve for three of the four configurations. The paper's caveat in the Discussion is welcome, but the main-text and abstract statements should be qualified to say that the exponent is anchor-sensitive and that the 1.74–1.95 range is not an independently constrained measurement.
- [Supplemental Material, Table S9, third row] The spot-check failure (SLy4/FSU2 at m1/m2 = 1.8/1.2, measured ΔlogZ = 1.11 ± 0.32 vs. 2.50 predicted) is not a small statistical fluctuation but a demonstration that the quadratic scaling can fail by more than a factor of two when the wrong-EOS arm compensates through mass-ratio drift. Because the analytic derivation treats the posterior-volume term Δvol as subdominant (Supplemental Material, Eq. (2) and surrounding text), this failure shows that the neglected term is not always subdominant. The framework therefore cannot claim to predict the SNR threshold without additional information about parameter-compensation channels; at minimum this limitation should be stated in the abstract and conclusion, not only in the Discussion.
minor comments (4)
- [Abstract] The abstract's 'within 0.3%' is misleading without the prediction interval; the 95% interval for the decisive SNR is [45.4, 58.8], so the central-value agreement is much tighter than the actual uncertainty. Please report the interval alongside the 0.3% in the abstract.
- [Main text, Table I caption] The mass-ratio notation 'q1 = 1.232/1.54' is ambiguous; since q is defined as m2/m1 with q ≤ 1 elsewhere, write q1 = 0.80 and q2 = 0.857 to avoid confusion.
- [Supplemental Material, Eq. (3)] The factor of 4 in the definition of J is not explained; please state the Fourier convention or inner-product normalization that produces it.
- [Code and data availability] References for the nested-sampling pipeline and data repositories are given, but the version or commit hashes are not; for reproducibility, please provide pinned versions.
Circularity Check
No significant circularity: the prefactor C is calibrated on Curves 1–3, but the Curve 4 prediction is a genuine out-of-sample test, and the paper explicitly treats C as configuration-dependent rather than as a derived universal constant.
full rationale
The paper's central quantitative claim is an empirical scaling law, ΔlogZ ≈ C(ΔΛ̃·SNR)^2, with a prefactor C calibrated on the anchor points of Curves 1–3 and then used to predict the decisive-evidence SNR for a fourth, previously unrun mass point. This is a legitimate out-of-sample calibration test rather than a circular reduction: the predicted quantity (Curve 4's threshold) is not in the fitting set, and the paper states that the prediction was made before Curve 4 was analyzed. The analytic derivation in the Supplemental Material provides the functional form from a Laplace/Occam-style expansion and the standard mismatch criterion from Lindblom, Owen & Brown and Read et al., which are external, established results; the paper does not rely on self-citation for the load-bearing argument. The paper also explicitly acknowledges that C is not universal, that it depends on masses, spins, and detector configuration, and that the fitted exponent n is anchor-sensitive and trends toward 2 rather than being an independent confirmation of exact quadratic scaling. The observed monotonic drift in K = ΔlogZ/SNR^2 (Table S7) and the Curve 4 anchor's lower effective C are presented as caveats and higher-order corrections, not hidden. Even if the 0.3% threshold agreement is partly coincidental or the high-SNR prefactor does not transfer cleanly, that is a scientific robustness concern, not a circularity: no equation is defined in terms of the predicted quantity, no fitted parameter is renamed as a prediction of the same data, and no unique model choice is justified by a self-citation chain. The derivation chain is therefore self-contained with respect to circularity, with all quantitative input parameters either measured from independent simulations or explicitly calibrated with stated transferability assumptions.
Assumptions & free parameters
free parameters (3)
- Scaling prefactor C =
7.19e-9 (mean of Curves 1-3)
- Per-curve amplitude A =
1.46e-3 to 3.57e-3
- Per-curve exponent n =
1.74 to 1.95 (anchor-dependent; 1.20-1.75 without anchor)
assumptions (5)
- standard math The Gaussian likelihood and Laplace approximation give log Z approximately log L(theta_hat) + log pi(theta_hat) + (1/2) log |2 pi Sigma|.
- domain assumption Injected data are zero-noise, so the noise cross-term <n|delta h_perp> vanishes exactly.
- ad hoc to paper The waveform mismatch between true and wrong EOS at leading order is a phase difference linear in Lambda_tilde: delta Psi_perp(f) approximately Delta Lambda_tilde times g1(f).
- ad hoc to paper The posterior-volume term Delta vol is subdominant and can be dropped or absorbed into C scatter.
- domain assumption The coefficient C calibrated at one binary mass point transfers to another mass point (Curve 4) and to other EOS pairs (spot-checks).
Cite this review
Pith. "Pith review of How Loud Must a Neutron-Star Merger Be to Reveal Its Equation of State?." pith.science (2026). https://pith.science/paper/2DKAHUZN
@misc{pith2026260805794,
author = {Pith},
title = {Pith review of: How Loud Must a Neutron-Star Merger Be to Reveal Its Equation of State?},
year = {2026},
howpublished = {\url{https://pith.science/paper/2DKAHUZN}},
note = {Machine review of arXiv:2608.05794}
}
abstract
The tidal response of neutron stars during binary inspiral encodes the equation of state (EOS) of dense matter in the gravitational-wave signal. Quantifying the signal-to-noise ratio (SNR) required to distinguish competing EOS models with third-generation detectors is therefore essential. We perform Bayesian nested-sampling parameter estimation on simulated binary neutron star signals observed by an Einstein Telescope plus two Cosmic Explorer detector network and compute the evidence difference between correct- and incorrect-EOS recovery models over a broad range of SNR. Across two tidal-deformability contrasts, a swap of the true and recovery EOS, and two binary mass points, we find a common scaling, $\Delta\log Z = A\,\mathrm{SNR}^{n}$ with $n \simeq 1.74$--$1.95$, where the EOS contrast and binary properties determine only the prefactor $A$. This behavior follows from an Occam-factor argument, yielding $\Delta\log Z \propto (\Delta\tilde{\Lambda}\,\mathrm{SNR})^{2}$. Calibrating this relation on three configurations predicts, before the run, the SNR required for decisive EOS discrimination in the fourth to within $0.3\%$. These results establish a quantitative framework for assessing the EOS-discrimination reach of third-generation gravitational-wave detector networks.
Figures
Reference graph
Works this paper leans on
-
[1]
J. S. Read, L. Baiotti, J. D. E. Creighton, J. L. Friedman, B. Giacomazzo, K. Kyutoku, C. Markakis, L. Rezzolla, M. Shibata, and K. Taniguchi, Phys. Rev. D88, 044042 (2013), arXiv:1306.4065
arXiv 2013
-
[2]
B. P. Abbottet al.(LIGO Scientific, Virgo), Phys. Rev. Lett.119, 161101 (2017), arXiv:1710.05832 [gr-qc]
arXiv 2017
-
[3]
Punturoet al., Class
M. Punturoet al., Class. Quant. Grav.27, 194002 (2010)
2010
-
[4]
D. Reitzeet al., Bull. Am. Astron. Soc.51, 035 (2019), arXiv:1907.04833 [astro-ph.IM]
arXiv 2019
-
[5]
A. Puecher, A. Samajdar, and T. Dietrich, Phys. Rev. D 108, 023018 (2023), arXiv:2304.05349 [astro-ph.IM]
arXiv 2023
-
[6]
L. Wade, J. D. E. Creighton, E. Ochsner, B. D. Lackey, B. F. Farr, T. B. Littenberg, and V. Raymond, Phys. Rev. D89, 103012 (2014)
2014
-
[7]
B. D. Lackey and L. Wade, Phys. Rev. D91, 043002 (2015)
2015
-
[8]
R. Kashyap, I. Gupta, A. Dhani, M. Bapna, and B. Sathyaprakash, Phys. Rev. D113, 063019 (2026), arXiv:2502.03831 [gr-qc]
arXiv 2026
Show all 35 references
-
[9]
Biswas, E
B. Biswas, E. Smyrniotis, I. Liodis, and N. Stergioulas, Phys. Rev. D109, 064048 (2024), arXiv:2309.05420
2024 arXiv
-
[10]
B. P. Abbottet al.(LIGO Scientific, Virgo), Class. Quant. Grav.37, 045006 (2020), arXiv:1908.01012 [gr- qc]
2020 arXiv
-
[11]
Lindblom, B
L. Lindblom, B. J. Owen, and D. A. Brown, Phys. Rev. D78, 124020 (2008), arXiv:0809.3844 [gr-qc]
2008 arXiv
-
[12]
T. D. P. Edwards, K. W. K. Wong, K. K. H. Lam, A. Coogan, D. Foreman-Mackey, M. Isi, and A. Zimmer- man, Phys. Rev. D110, 064028 (2024), arXiv:2302.05329 [astro-ph.IM]
2024 arXiv
-
[13]
K. W. K. Wong, M. Isi, and T. D. P. Edwards, Astrophys. J.958, 129 (2023), arXiv:2302.05333 [astro-ph.IM]
2023 arXiv
-
[14]
Wouters, P
T. Wouters, P. T. H. Pang, T. Dietrich, and C. Van Den Broeck, Phys. Rev. D110, 083033 (2024), arXiv:2404.11397 [astro-ph.IM]
2024
-
[15]
Dietrich, S
T. Dietrich, S. Bernuzzi, and W. Tichy, Phys. Rev. D96, 121501 (2017), arXiv:1706.02969 [gr-qc]
2017 arXiv
-
[16]
Dietrich, A
T. Dietrich, A. Samajdar, S. Khan, N. K. Johnson- McDaniel, R. Dudi, and W. Tichy, Phys. Rev. D100, 044003 (2019), arXiv:1905.06011 [gr-qc]
2019 arXiv
-
[17]
J. R. Oppenheimer and G. M. Volkoff, Phys. Rev.55, 374 (1939)
1939
-
[18]
R. C. Tolman, Phys. Rev.55, 364 (1939)
1939
-
[19]
A. W. Steiner, M. Hempel, and T. Fischer, Astrophys. J. 774, 17 (2013)
2013
-
[20]
Typel, G
S. Typel, G. R¨ opke, T. Kl¨ ahn, D. Blaschke, and H. H. Wolter, Phys. Rev. C81, 015803 (2010)
2010
-
[21]
B. P. Abbottet al.(LIGO Scientific, Virgo), Astrophys. J. Lett.848, L12 (2017), arXiv:1710.05833 [astro-ph.HE]
2017 arXiv
-
[22]
Hotokezaka, E
K. Hotokezaka, E. Nakar, O. Gottlieb, S. Nissanke, K. Masuda, G. Hallinan, K. P. Mooley, and A. T. Deller, Nature Astron.3, 940 (2019), arXiv:1806.10596
2019 arXiv
-
[23]
Jeffreys,Theory of Probability, 3rd ed
H. Jeffreys,Theory of Probability, 3rd ed. (Oxford Uni- versity Press, Oxford, 1961)
1961
-
[24]
Chabanat, P
E. Chabanat, P. Bonche, P. Haensel, J. Meyer, and R. Schaeffer, Nucl. Phys. A635, 231 (1998), [Erratum: Nucl.Phys.A 643, 441–441 (1998)]
1998
-
[25]
C. A. Raithel and E. R. Most, Phys. Rev. Lett.130, 201403 (2023), arXiv:2208.04294 [astro-ph.HE]
2023 arXiv
-
[26]
C. A. Raithel and E. R. Most, Phys. Rev. D108, 023010 (2023), arXiv:2208.04295 [astro-ph.HE]
2023 arXiv
-
[27]
how loud must a neutron-star merger be to reveal its equation of state?
S. M. A. Imam, Jim nested sampling: Code and data for “how loud must a neutron-star merger be to reveal its equation of state?” (2026), GitHub repository. Accessed July 2026
2026
-
[28]
how loud must a neutron- star merger be to reveal its equation of state?
S. M. A. Imam, Data for “how loud must a neutron- star merger be to reveal its equation of state?”, Zenodo (2026)
2026
-
[29]
Iacovelli, M
F. Iacovelli, M. Mancarella, S. Foffa, and M. Maggiore, Astrophys. J. Suppl.263, 2 (2022), arXiv:2207.06910
2022 arXiv
-
[30]
Srivastava, D
V. Srivastava, D. Davis, K. Kuns, P. Landry, S. Ballmer, M. Evans, E. D. Hall, J. Read, and B. S. Sathyaprakash, Astrophys. J.931, 22 (2022), arXiv:2201.10668
2022 arXiv
-
[31]
Branchesiet al., Journal of Cosmology and Astropar- ticle Physics2023(07), 068
M. Branchesiet al., Journal of Cosmology and Astropar- ticle Physics2023(07), 068
-
[32]
B. P. Abbottet al., Astrophys. J. Lett.848, L13 (2017), arXiv:1710.05834
2017 arXiv
-
[33]
G. A. Lalazissis, T. Nikˇ si´ c, D. Vretenar, and P. Ring, Phys. Rev. C71, 024312 (2005)
2005
-
[34]
Chen and J
W.-C. Chen and J. Piekarewicz, Phys. Rev. C90, 044305 (2014)
2014
-
[35]
xylophone
G. A. Lalazissis, J. K¨ onig, and P. Ring, Phys. Rev. C55, 540 (1997), arXiv:nucl-th/9607039. Supplemental Material This supplement provides (i) the full derivation of the analytic scaling relation ∆ logZ≈C(∆ ˜Λ·ρ) 2 summarized in the main text, (ii) the full per-point ∆ logZd...
1997 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.