{"id":"aeac34be-47a1-4f78-b9e9-98d90b148ae1","arxiv_id":"2608.05794","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The evidence for distinguishing two neutron-star equations of state grows as SNR to roughly the 1.8 power, with a prefactor set by tidal-deformability contrast, and this scaling predicted an unseen binary configuration's discrimination threshold within 0.3%.","lead":"This paper simulates neutron-star merger signals in a future gravitational-wave detector network and measures how loud a signal must be to tell two competing equations of state apart. It finds a simple power-law rule that predicts the required loudness, which could help plan observations with next-generation detectors.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Calibrated prefactor C fails to transfer to Curve 4's high-SNR anchor (C4≈4.5e-9 vs. 7.19e-9); the 0.3% threshold match is a low-SNR coincidence, not evidence for universal quadratic scaling.","rationale":"The paper is honest and the empirical campaign is substantial, including a genuine pre-registered out-of-sample test. However, the most load-bearing assumption is not merely that the tidal phase is linear in ΔΛ̃ and that Δvol is subdominant; it is that the prefactor C is effectively constant across SNR so that calibrating at the anchor and extrapolating to the threshold is valid. The paper's own data already show C_eff drifts: monotonically for Curve 1 (Table S7) and by roughly 30–40% between low and high SNR for Curve 4. The calibrated C from Curves 1–3 fails to predict Curve 4's high-SNR anchor by 37%, so the 0.3% threshold agreement is a low-SNR near-coincidence, not evidence of a transferable prefactor. This goes beyond the reader's weakest assumption, which attributes the failure to parameter compensation and spot-checks; it directly concerns the headline out-of-sample result. The reader's CONDITIONAL verdict remains appropriate, but the condition should explicitly require demonstrating that C is SNR-independent (or characterizing its SNR dependence) before the calibrated prefactor is used for prediction. I therefore keep the verdict unchanged while flagging this more specific, quantitatively checkable concern.","tokens_in":18811,"tokens_out":16758,"duration_ms":144269,"concrete_test":"Predict the Curve 4 anchor-point evidence from the calibrated mean C and the measured anchor SNR: ΔlogZ_pred = C̄(ΔΛ̃4·SNR4)^2 ≈ 7.19e-9 × (518.91 × 3192.43)^2 ≈ 1.97e4. The measured anchor value is 12359.66 (Table S6), a 37% deficit. If the prefactor C transferred, the anchor should match within the ~30% scatter used to define C; the actual mismatch shows the calibration is SNR/mass dependent, and the 0.3% threshold agreement is not a test of the scaling law. Re-run the Curve 4 out-of-sample prediction using C4 instead of C̄; the predicted decisive SNR shifts to ≈64, outside the measured 50.7 and outside the original 95% interval.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central out-of-sample prediction calibrates C from the SNR≈3355 anchor points of Curves 1–3 and then assumes ΔlogZ = C(ΔΛ̃·SNR)^2 holds down to the decisive threshold near SNR≈50. But C is not constant along a curve: for Curve 1, K=ΔlogZ/SNR^2 decreases monotonically from 1.005e-3 at SNR=117 to 9.012e-4 at the anchor (Table S7). For Curve 4, the anchor gives K=1.213e-3; with ΔΛ̃=518.91, C4=K/ΔΛ̃^2≈4.50e-9, which is 37% below the calibrated mean C̄=7.19e-9 (Table S1). Using C4 with the quadratic formula predicts a decisive SNR of ≈64, not the measured 50.7. The threshold prediction only works because Curve 4's effective C at low SNR (≈6–8e-9) is close to C̄, not because C transfers. The pre-registered test therefore validates a low-SNR effective coefficient, not the high-SNR prefactor being calibrated. This undermines the framework claim: the scaling is not a single power law with a transferable prefactor; the prefactor drifts systematically with SNR and with mass point. The paper's own Table S7 contains the evidence for the drift but interprets it as scatter.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses nested-sampling Bayesian parameter estimation on zero-noise simulated binary neutron star signals in an ET+2CE network to compute the evidence difference ΔlogZ between recovery with the correct and with a wrong EOS template. It reports a near-quadratic power-law scaling ΔlogZ = A SNR^n with n ≈ 1.74–1.95 across four configurations spanning two tidal-deformability contrasts, an EOS true/wrong role swap, and two binary mass points. A Laplace/Occam expansion is used to derive the leading-order form ΔlogZ ≈ C(ΔΛ̃ SNR)^2; C is calibrated from the highest-SNR anchor points of Curves 1–3, and a pre-registered prediction for the decisive-evidence SNR of the unseen Curve 4 agrees with the measured 50.7 to 0.3%. The Supplemental Material reports anchor-point sensitivity, a densification of Curve 1, five additional spot-checks (one failing at 44% of the prediction), and a free-sky/real-noise robustness check.","tokens_in":19210,"tokens_out":9984,"duration_ms":85462,"significance":"If the claimed scaling were established, it would give a practical closed-form estimate for the SNR at which a third-generation detector network can distinguish two candidate neutron-star equations of state, which is useful for observational planning and event prioritization. The paper's strengths are its pre-registered out-of-sample test, public code and data, and unusually transparent handling of caveats: anchor leverage, the narrow SNR range without the anchor, the zero-noise/fixed-sky simplification, and the one failing spot-check are all explicitly disclosed. The prefactor-drift issue discussed below, however, means the central 'transferable prefactor' claim is not yet supported in its current form.","major_comments":[{"comment":"The calibrated prefactor C is not constant along a curve, so the Curve 4 prediction does not validate transfer of the high-SNR-calibrated prefactor. For Curve 1, K = ΔlogZ/SNR² falls monotonically from 1.005×10⁻³ at SNR = 117.41 to 9.012×10⁻⁴ at the d_L = 42 Mpc anchor (Table S7). For Curve 4, the anchor gives K = 12359.66/3192.43² = 1.213×10⁻³, hence C4 = K/(518.91)² ≈ 4.50×10⁻⁹, which is 37% below the calibrated mean C̄ = 7.19×10⁻⁹ (Table S1). Inserting C4 into the quadratic formula gives a decisive SNR of about 64, not the measured 50.7. The 0.3% agreement at threshold therefore comes from Curve 4's low-SNR effective coefficient being close to C̄, not from the prefactor transferring from the high-SNR calibration. The claim that 'the coefficient transfers between configurations' (main text Results) needs to be replaced by an explicit statement that C is SNR-dependent and that the prediction is for an effective low-SNR coefficient.","section":"Supplemental Material, Eq. (4) and Tables S6–S7"},{"comment":"The claimed common exponent n ≈ 1.74–1.95 is not robust without the anchor. Dropping the single highest-SNR point lowers n to 1.46 (Curve 1, 8-point), 1.20 (Curve 2), 1.75 (Curve 3), and 1.65 (Curve 4); the densification repair was performed only for Curve 1. Thus the abstract's 'common scaling' rests on one anchor point per curve for three of the four configurations. The paper's caveat in the Discussion is welcome, but the main-text and abstract statements should be qualified to say that the exponent is anchor-sensitive and that the 1.74–1.95 range is not an independently constrained measurement.","section":"Supplemental Material, Table S8"},{"comment":"The spot-check failure (SLy4/FSU2 at m1/m2 = 1.8/1.2, measured ΔlogZ = 1.11 ± 0.32 vs. 2.50 predicted) is not a small statistical fluctuation but a demonstration that the quadratic scaling can fail by more than a factor of two when the wrong-EOS arm compensates through mass-ratio drift. Because the analytic derivation treats the posterior-volume term Δvol as subdominant (Supplemental Material, Eq. (2) and surrounding text), this failure shows that the neglected term is not always subdominant. The framework therefore cannot claim to predict the SNR threshold without additional information about parameter-compensation channels; at minimum this limitation should be stated in the abstract and conclusion, not only in the Discussion.","section":"Supplemental Material, Table S9, third row"}],"minor_comments":[{"comment":"The abstract's 'within 0.3%' is misleading without the prediction interval; the 95% interval for the decisive SNR is [45.4, 58.8], so the central-value agreement is much tighter than the actual uncertainty. Please report the interval alongside the 0.3% in the abstract.","section":"Abstract"},{"comment":"The mass-ratio notation 'q1 = 1.232/1.54' is ambiguous; since q is defined as m2/m1 with q ≤ 1 elsewhere, write q1 = 0.80 and q2 = 0.857 to avoid confusion.","section":"Main text, Table I caption"},{"comment":"The factor of 4 in the definition of J is not explained; please state the Fourier convention or inner-product normalization that produces it.","section":"Supplemental Material, Eq. (3)"},{"comment":"References for the nested-sampling pipeline and data repositories are given, but the version or commit hashes are not; for reproducibility, please provide pinned versions.","section":"Code and data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually honest and the code/data availability is commendable, but the advertised 0.3% out-of-sample agreement is less compelling once the SNR dependence of the prefactor is taken into account. The author already has most of the evidence (Table S7) in hand; a revision that reframes the claim around an effective low-SNR coefficient, or that recalibrates C at the threshold SNR, would make the central claim sound. I would also encourage the editor to ask that the exponent claim be softened given Table S8."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know the main result is genuinely out-of-sample: the author calibrated a prefactor C on three Bayesian nested-sampling curves, then predicted the decisive-evidence SNR for a fourth, unseen mass point before running it, and hit 50.7 vs the predicted 50.8. That is good practice and the paper ships code and data, so the result is checkable.\n\nWhat is new here is not the quadratic scaling itself — that follows from standard mismatch arguments (Read et al., Lindblom et al.) — but the attempt to turn it into a calibrated predictive tool. The paper is also unusually candid: it discloses the anchor-point leverage on the fitted exponent, the zero-noise fixed-sky simplification, the prior-boundary wall-pinning, and one spot-check that missed by more than a factor of two. That level of transparency deserves credit.\n\nThe soft spot is real and the stress-test note has it right. The calibration uses the anchor points at SNR≈3355, but K ≡ ΔlogZ/SNR² declines monotonically along each curve. Table S7 shows K for Curve 1 falling from 1.005e-3 at SNR=117 to 9.012e-4 at the anchor — about a 10% drift. For Curve 4, the anchor gives C≈4.5e-9, which is 37% below the calibrated mean of 7.19e-9 and would predict a decisive SNR of ≈64, well outside the pre-registered 95% interval. The 0.3% match only works because Curve 4's effective C near the threshold is close to the calibrated mean. So the test validates a low-SNR effective coefficient, not the transfer of the high-SNR prefactor. The paper's own Table S7 contains the evidence for the drift, but the text does not connect it to the calibration scheme.\n\nThis does not sink the paper. The qualitative scaling — evidence grows roughly as SNR^1.8–2, with a contrast-dependent prefactor — is credible and useful for first-order planning. But the phrase \"quantitative framework\" and the transferability claim need to be softened. The author already concedes C must be recalibrated per EOS pair and mass regime; the same logic applies to SNR. The out-of-sample test should be described as testing the effective low-SNR coefficient, and the SNR dependence of C should either be modeled or explicitly measured.\n\nReader: someone planning 3G EOS-discrimination analyses or wanting a closed-form SNR estimate will find this useful, with the stated caveats. Yes, send it to peer review. The issues are fixable in revision and the empirical core is solid.","headline":"Real, reproducible scaling study whose headline out-of-sample prediction works because of an SNR-dependent coefficient, not a transferable prefactor; worth publishing after the claim is reframed.","tokens_in":19633,"tokens_out":5714,"would_cite":true,"duration_ms":52483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The evidence distinguishing two neutron-star equations of state grows as the square of (tidal contrast × SNR), and a three-configuration calibration predicts a new binary's decisive loudness to 0.3%.","keywords":["neutron star equation of state","gravitational waves","Bayesian model selection","tidal deformability","third-generation detectors","Einstein Telescope","Cosmic Explorer","Occam factor"],"falsifier":"Rerun the failing spot-check configuration (a Skyrme EOS injected, a relativistic mean-field EOS recovered, at component masses 1.8 and 1.2 solar masses) with the mass-ratio prior widened beyond the value the wrong arm drifts to; if $\\Delta\\log Z$ returns to the predicted $C(\\Delta\\tilde{\\Lambda}\\,\\mathrm{SNR})^{2}$ value, the quadratic law survives and the suppression is a prior-boundary artifact, while if it remains near 44% of prediction, the leading-order scaling itself fails for that configuration.","tokens_in":18623,"feed_emoji":"🔊","tokens_out":15489,"duration_ms":118960,"temperature":0.7,"pith_summary":"This paper asks a practical question for the next generation of gravitational-wave observatories: how loud must a single neutron-star merger be before the data can say which dense-matter equation of state (EOS) produced it? Using full Bayesian nested-sampling analysis of simulated signals in an Einstein Telescope plus two Cosmic Explorer network, the author finds that the evidence favoring the correct EOS over a wrong one grows as a single power law in the signal-to-noise ratio, $\\Delta\\log Z = A\\,\\mathrm{SNR}^{n}$ with $n \\simeq 1.74$--$1.95$, and that this nearly quadratic growth follows from a simple Occam-factor argument. The same argument collapses the scaling to $\\Delta\\log Z \\approx C(\\Delta\\tilde{\\Lambda}\\,\\mathrm{SNR})^{2}$, where $\\Delta\\tilde{\\Lambda}$ is the difference in tidal deformability between the two candidate EOS and $C$ is a prefactor that depends on the binary and the network. Calibrating $C$ on three configurations and then predicting the decisive-evidence SNR for a fourth, unseen binary mass reproduces the measured value to within $0.3\\%$. This converts what would otherwise be an expensive campaign of nested-sampling runs for every new EOS pair into a closed-form estimate of the loudness needed for EOS discrimination, giving third-generation detectors a concrete design and event-selection benchmark.","feed_headline":"A single law predicts how loud a merger must be to reveal its EOS","feed_subtitle":"Tested on a fourth, unseen binary, the formula pinned the decisive signal-to-noise ratio to within 0.3 percent.","key_machinery":"The load-bearing object is the Occam-factor (Laplace) expansion of the Bayesian evidence, which the paper reduces to $\\Delta\\log Z \\approx \\frac{1}{2}\\|\\delta h_{\\perp}\\|^{2}$ for the zero-noise injections used here. Here $\\delta h_{\\perp}$ is the noise-weighted waveform mismatch between the correct-EOS template and the wrong-EOS template after the wrong arm's nuisance parameters (mass ratio, spins, coalescence time, distance, orientation) have been optimized out; the subscript $\\perp$ records that the raw EOS-induced waveform difference is projected orthogonal to those nuisance directions. Because the tidal phase in the tidally corrected waveform model used in the study is, at leading post-Newtonian order, linear in the mass-weighted tidal deformability $\\tilde{\\Lambda}$, the mismatch $\\delta\\psi_{\\perp}(f) \\approx \\Delta\\tilde{\\Lambda}\\,g_{1}(f)$ factors out of the integral, giving $\\Delta\\log Z \\approx C(\\Delta\\tilde{\\Lambda}\\,\\mathrm{SNR})^{2}$ with $C = I_{\\perp}/2$ and $I_{\\perp}$ a distance-independent spectral integral over the noise-weighted template power. This single identity converts model selection---normally a nested-sampling computation---into a closed-form estimate, and it also identifies the failure mode: when the wrong arm's posterior pins against a prior boundary (e.g., mass ratio drifting to its upper limit), the projection $\\delta h_{\\perp}$ changes and the quadratic scaling can break.","core_discovery":"The paper's central claim is that, for zero-noise simulated binary neutron star signals, the Bayesian evidence difference between a recovery model whose template EOS matches the injected EOS and one whose template EOS does not follows $\\Delta\\log Z = A\\,\\mathrm{SNR}^{n}$ with $n \\simeq 1.74$--$1.95$ across four configurations spanning tidal-deformability contrasts of about 51% and 28%, both EOS-role assignments, and two binary mass points. More strongly, the evidence gap is, to leading order, $\\Delta\\log Z \\approx C(\\Delta\\tilde{\\Lambda}\\,\\mathrm{SNR})^{2}$, a quadratic law derived from a Laplace/Occam-factor expansion of the evidence integral in which the waveform mismatch between correct and wrong templates is linear in the tidal-deformability contrast at leading post-Newtonian order. The prefactor $C$ is not universal---it depends on the binary masses, spins, sky position, detector network, and priors---but when calibrated on three curves it transferred to a fourth, independent mass point: the pre-registered prediction of $\\mathrm{SNR} \\approx 50.8$ for decisive evidence matched the measured $50.7$, a $0.3\\%$ discrepancy. The paper also reports that four of five additional spot-checks across other EOS pairs (including a Skyrme functional against relativistic mean-field models) matched the analytic prediction to within about 15%, while the one outlier, coincident with strong mass-ratio compensation, measured only 44% of the predicted evidence gap, showing the simple scaling is a leading-order estimate that can be suppressed when the wrong template's nuisance parameters absorb the mismatch.","pith_inferences":["The near-quadratic law implies diminishing returns on detector sensitivity: doubling the SNR only quadruples the evidence, so pushing a marginal event from SNR 30 to 60 buys the same discrimination gain as going from 60 to 120; observational planning should therefore target events already near the threshold, not only the loudest.","The clean transfer of $C$ across one mass change hints that $C$ may approximately factorize into a mass-dependent piece and a network-dependent piece; if that factorization holds across a denser mass grid, $C$ could be precomputed per network and stored as a lookup table, making the closed-form estimate an online planning tool.","The failure mode identified in the spot-check suggests building a map of 'compensation valleys'---binary configurations where a wrong EOS can hide its mismatch by shifting mass ratio, spin, or time to prior boundaries---would bound the regime of validity of the quadratic law; the paper's wall-pinning analysis provides the first entries of such a map.","The zero-noise, fixed-sky assumption is checked by a single three-seed comparison showing a shift within the expected noise scatter; a full noise-averaged, free-sky campaign across all configurations would test whether the $0.3\\%$ transferability survives realistic conditions, which is the natural next step the paper leaves open."],"forward_implications":["The SNR threshold for decisive evidence ($\\Delta\\log Z = 5$) EOS discrimination can be predicted analytically for an unmeasured configuration once $C$ is calibrated on a few configurations; the Curve 4 test confirms this at the $0.3\\%$ level.","For the fiducial configuration with about a 51% tidal-deformability contrast, decisive discrimination occurs at network SNR $\\approx 56$ (strong evidence at $\\approx 38$ and substantial at $\\approx 23$); reducing the contrast to about 28% pushes the decisive threshold to $\\approx 89$.","The calibrated scaling implies that 'tidal deformability doppelgänger' EOS pairs with $\\Delta\\tilde{\\Lambda} \\approx 10$--$30$ would require SNR $\\approx 900$--$2800$ for decisive discrimination, reachable only for rare, very nearby events rather than typical detections.","Because $C$ depends on the binary's masses, spins, and the detector network, the operational rule is to recalibrate $C$ for each new EOS pair and mass regime rather than to reuse a single value across an ensemble.","The one failing spot-check (44% of predicted evidence, coincident with mass-ratio compensation) shows that parameter compensation can suppress the evidence gap by more than a factor of two; the quadratic law is a leading-order benchmark, not a universal guarantee."],"supporting_citations":[{"why":"Supplies the noise-weighted inner product form $\\|\\delta h\\|^{2}$ that underlies the SNR$^2$ evidence penalty in the derivation.","marker":"[1]"},{"why":"Establishes the waveform-mismatch distinguishability criterion used by the Occam-factor expansion.","marker":"[11]"},{"why":"Provides the nested-sampling inference engine used for every evidence computation in the study.","marker":"[12–14]"},{"why":"The tidally corrected waveform model whose phase is, to leading order, linear in $\\tilde{\\Lambda}$; carries the EOS imprint.","marker":"[15, 16]"},{"why":"Supplies the softer SFHo EOS template used in the main contrast and role-swap configurations.","marker":"[19]"},{"why":"Supplies the DD2 hadronic branch onto which the hybrid EOS variants are patched; defines the stiff template family.","marker":"[20]"},{"why":"Supplies the conventional substantial/strong/decisive evidence thresholds used to define the SNR benchmarks.","marker":"[23]"},{"why":"Documents tidal-deformability doppelgänger EOS pairs whose small contrast implies SNR $\\approx 900$--$2800$ for decisive discrimination.","marker":"[25, 26]"}],"fun_headline_variants":["A single law predicts how loud a merger must be to reveal its EOS","New formula predicts exact SNR needed to reveal neutron-star EOS","One scaling law tells when a merger's signal reveals its EOS","Quadratic law forecasts EOS-revealing SNR within 0.3%","Loudness threshold for neutron-star EOS now predictable to 0.3%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The scaling law holds only if the phase difference between the correct and wrong templates is, to leading order, linear in the tidal-deformability contrast and if the wrong template's nuisance parameters do not absorb enough of that difference to change the projected mismatch; when mass ratio, spin, or coalescence time can compensate the EOS error---as in one of the five spot-checks---the quadratic law can overpredict the evidence gap by more than a factor of two.","fun_headline_variants_meta":{"raw":{"variants":["A single law predicts how loud a merger must be to reveal its EOS","New formula predicts exact SNR needed to reveal neutron-star EOS","One scaling law tells when a merger's signal reveals its EOS","Quadratic law forecasts EOS-revealing SNR within 0.3%","Loudness threshold for neutron-star EOS now predictable to 0.3%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000676,"raw_usage":{"total_tokens":3170,"prompt_tokens":1138,"completion_tokens":2032,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":754,"completion_tokens_details":{"reasoning_tokens":1933}},"tokens_in":754,"tokens_out":2032,"duration_ms":12777,"temperature":1.0,"reasoning_tokens":1933,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:13:13.735511+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the failing spot-check configuration (a Skyrme EOS injected, a relativistic mean-field EOS recovered, at component masses 1.8 and 1.2 solar masses) with the mass-ratio prior widened beyond the value the wrong arm drifts to; if $\\Delta\\log Z$ returns to the predicted $C(\\Delta\\tilde{\\Lambda}\\,\\mathrm{SNR})^{2}$ value, the quadratic law survives and the suppression is a prior-boundary artifact, while if it remains near 44% of prediction, the leading-order scaling itself fails for that configuration.","supporting_citations":[{"cited_title":"Jeffreys,Theory of Probability, 3rd ed","cited_arxiv_id":null,"evidence_quote":"Supplies the conventional substantial/strong/decisive evidence thresholds used to define the SNR benchmarks."}],"review_version":2}