{"id":"6a2c2c5c-be2e-44e5-ac42-3a0050053248","arxiv_id":"2506.10530","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Using mismatch distances computed at best-fit parameters, per parameter, the paper predicts accurate bias SNRs for aligned-spin binary black hole measurements and derives waveform accuracy requirements for next-generation detectors.","lead":"The paper shows that the standard 'faithfulness mismatch' metric is too pessimistic for deciding when a gravitational-wave model will bias measurements, and that a corrected, per-parameter calculation accurately predicts the signal-to-noise ratio at which bias sets in. It matters because next-generation detectors will see very loud signals where waveform model errors become the dominant source of uncertainty, and this work sets practical accuracy targets.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the Pythagorean-relation concern is not load-bearing because bias distances are computed directly.","rationale":"The reader's verdict (ACCEPT, high confidence) is appropriate. The paper is a careful methodological study whose central claim, that a correctly calculated mismatch-based criterion predicts bias SNRs, is validated against full Bayesian inference with independent Fisher and PCA cross-checks. The reader identified the Pythagorean relation as the weakest assumption, but my reading of the text shows that the bias distances are computed directly and do not rely on Eq. (9) for the results. The actual load-bearing assumption is the linear-signal approximation connecting mismatch distances to credible-interval boundaries. This assumption is standard in the field and is empirically supported by the nine validation cases, with the paper's own caveats (prior railing, numerical precision above SNR ~500) clearly stated. The limited scope (simplified waveform model, small number of cases) is appropriate for a proof-of-principle and does not undermine the central methodological claim. A useful additional check would be to test the method in a larger-mismatch regime, but this is an extension rather than a correction. Hence the verdict remains UNCHANGED, and the agreement with the reader is partial because the specific weakest assumption they named is not the one that actually carries the argument.","tokens_in":30565,"tokens_out":20032,"duration_ms":241140,"concrete_test":"Run one additional validation case with a deliberately larger model error, for example a mass-ratio q=18 aligned-spin injection at a parameter-space point where PhenomD's faithfulness mismatch is an order of magnitude larger than in BAM-5, and perform full Bayesian parameter estimation at the predicted 1D bias SNR for one parameter (e.g., the primary mass). Check whether the true parameter value still falls on the 90% credible-interval boundary as the mismatch-based prediction asserts; if it does, the linear-signal-approximation mapping is supported outside the small-mismatch regime, and if it does not, the generality of the central claim is bounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption, the approximate Pythagorean relation Eq. (9), is not actually load-bearing for the paper's central claim. The bias distances used in the headline results are computed directly as mismatches between model waveforms, M(h_bf, h_s) in Eq. (8) and M(h_bf, h_bf|theta_i) in Sec. V, not by subtracting the effectualness mismatch from the faithfulness mismatch. Eq. (9) appears only as an equivalence with Ref. [22] and as a post-hoc check that the relation holds (Table I). Even if the Pythagorean decomposition failed, the directly computed distances would be unchanged. The genuinely load-bearing premise is the mapping from these mismatch distances to the 90% credible boundary of the Bayesian posterior, which rests on the linear-signal (Gaussian posterior) approximation and the chi-square scaling in Eq. (6). The paper validates this mapping against full parameter estimation for nine cases with small mismatches (d^2 <= 3.7e-3) and a single simplified aligned-spin (2,2)-mode model, and it explicitly discloses the failure modes that do appear (prior railing in BAM-3/BAM-4, and numerical unreliability above SNR ~500). Within that stated scope, the agreement between predicted and measured bias SNRs is convincing and supported by independent Fisher and PCA cross-checks. I therefore do not find a significant objection that would change the acceptance verdict.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper revisits the standard mismatch-based indistinguishability criterion for gravitational-wave signal models, which is known to be overly conservative for estimating the SNR at which model error biases parameter measurements. The authors propose two corrections: first, replace the faithfulness mismatch between the true signal and the model at the true parameters by the mismatch between the model at the best-fit parameters and the model at the true parameters (Eq. 8); second, compute per-parameter bias SNRs by fixing the parameter of interest to its true value and optimizing all others. They test these predictions against full Bayesian parameter estimation for nine aligned-spin, quadrupole-only binary black hole cases, using both in-sample (BAM) and out-of-sample (NRHybSur3dq8) signals as true signals, and cross-check with PCA-based credible-interval rescaling and Fisher analysis. They find that the N-D bias SNR correctly places the true parameters on the 90% credible boundary in the N-D posterior, and that per-parameter bias SNRs match the onset of bias in 1D marginalized posteriors within the reliable SNR regime (<500), with deviations attributable to prior railing at the chi_2z=-1 boundary of the model and to numerical precision at high SNR.","tokens_in":30745,"tokens_out":8548,"duration_ms":104811,"significance":"The paper provides a practical and computationally cheap method for setting waveform accuracy requirements for next-generation detectors, replacing overly conservative faithfulness-based estimates. The central derivation is standard (Gaussian posterior, chi-square scaling), and the paper validates its predictions against full PE for nine cases, including out-of-sample waveforms, with independent Fisher and PCA cross-checks. The method is not entirely new (the equivalence to Ref. [22] is acknowledged), but the application to binary black hole waveform models and the systematic validation against full PE is a useful contribution. The paper honestly discloses its limitations: only the (2,2) multipole of aligned-spin binaries is considered, predictions above SNR~500 are numerically unreliable, and prior-railing cases are identified. Within this stated scope, the evidence supports the central claim. The Pythagorean relation in Eq. (9) is not load-bearing because the bias distances are computed directly from Eq. (8), consistent with the skeptic's assessment.","major_comments":[],"minor_comments":[{"comment":"The name \"Baumgate-Shapiro-Shibata-Nakamura\" should be \"Baumgarte-Shapiro-Shibata-Nakamura\".","section":"Section III.A"},{"comment":"The word \"Pricipal\" should be \"Principal\".","section":"Figure 4 caption"},{"comment":"The class name \"StandardScalar\" should be \"StandardScaler\".","section":"Section IV.A"},{"comment":"The word \"verticle\" should be \"vertical\".","section":"Figures 10 and 11 captions"},{"comment":"The word \"coalesence\" should be \"coalescence\".","section":"Section II"},{"comment":"The abstract's claim of \"accurate estimates\" should be tempered by the caveat, stated later in the paper, that SNR predictions above ~500 are numerically unreliable; a short qualifier in the abstract would avoid overstatement.","section":"Abstract and Section V.A"}],"recommendation":"accept","confidential_remarks":"This is a carefully written proof-of-principle paper. The authors disclose their limitations explicitly, and the central claim is supported by full PE validation with independent cross-checks. The in-sample/out-of-sample distinction (BAM vs. surrogate signals) strengthens the test. I recommend acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know before reading: the method here is not new, and the authors say so. What the paper actually does is take the per-parameter bias-SNR idea from the existing literature and run it on aligned-spin BBH waveforms, with real checks against Bayesian parameter estimation. The main claim is that a correctly computed mismatch, meaning the distance between the model at best-fit parameters and the model at true parameters rather than the faithfulness mismatch, predicts the SNR at which individual parameters become biased. That claim holds up in the tested cases.\n\nThe validation is the strongest part. Nine injections, full PE, PCA-based CI rescaling, and a Fisher analysis that is done carefully, with the alignment procedure and both parameter sets. The agreement is convincing, and the paper is transparent about where it breaks: prior railing in BAM-3/BAM-4, numerical unreliability above SNR ~500, and the simplified (2,2)-mode aligned-spin model. The in-sample nature of some BAM cases is a real caveat, but the SUR injections are out-of-sample and behave consistently.\n\nI agree with the stress-test note that the supposed Pythagorean-relation weakness is not load-bearing. Eq. (9) is used as a check and as a link to Toubiana-Gair; the bias distances in the headline results are computed directly from the mismatches. If the Pythagorean decomposition failed, those numbers would not change. The load-bearing assumptions are the Gaussian posterior scaling and the linear-signal approximation, and the paper validates those against PE for the regime it claims. There is no general proof, but the scope is stated clearly.\n\nThe broader accuracy-requirement discussion, four orders of magnitude versus modest improvements, is necessarily tentative because it is based on one simplified model. That is the main soft spot: the headline numbers for future detectors are more illustrative than definitive. The citation pattern is fine; prior work is credited rather than buried. For waveform modelers and ET/CE planning, this is a useful framework.\n\nI would send this to a serious referee, and I expect it to be accepted with the limitations taken at face value. It is honest work, and the caveats are in the paper rather than hidden.","headline":"A carefully validated demonstration that per-parameter mismatch calculations give accurate bias SNRs for aligned-spin BBHs, with the method's novelty honestly bounded and its limitations stated up front.","tokens_in":31343,"tokens_out":2154,"would_cite":true,"duration_ms":28128,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["83C35"],"pacs":[],"model":"deepseek-v4-flash","headline":"The standard mismatch criterion is conservative because it uses the wrong distance; the paper shows that using the model's best-fit waveform and per-parameter distances makes it accurately predict the SNR at which biased measurements begin.","keywords":["gravitational waves","waveform mismatch","parameter bias","indistinguishability SNR","binary black holes","Bayesian parameter estimation","waveform accuracy","next-generation detectors"],"falsifier":"Take a deliberately deficient model whose error is known to lie along a curved direction of parameter space, compute the per-parameter bias SNRs from the Pythagorean relation, and compare with a full Bayesian posterior run at the predicted SNR; a significant offset of the true parameter from the 90% boundary would disprove the relation's scope. A cheaper version is to measure the left- and right-hand sides of $M(h_{\\mathrm{bf}},h_s) \\approx M(s,h_s)-M(s,h_{\\mathrm{bf}})$ across a grid of high-mismatch signals and identify where the approximation fails.","tokens_in":30305,"feed_emoji":"🔭","tokens_out":9333,"duration_ms":109486,"temperature":0.7,"pith_summary":"Gravitational-wave astronomers judge when model error will corrupt source measurements by computing a mismatch between a true signal and a waveform model and converting it into an indistinguishability SNR. This paper argues that the standard form of that criterion is conservative for a specific and fixable reason: it compares the true signal with the model evaluated at the true parameters, thereby counting model error that lies perpendicular to the model manifold and cannot shift the measured parameters. Once the mismatch is instead taken between the model at its best-fit parameters and the model at the true parameters, and once it is computed separately for each parameter, the resulting bias SNR correctly places the true parameter on the boundary of the 90% credible interval. The authors demonstrate this by comparing predictions with full Bayesian parameter estimation for nine aligned-spin binary-black-hole signals, where per-parameter bias SNRs can exceed the standard estimate by a wide margin and vary strongly across parameter space. If this stands, waveform accuracy requirements for next-generation detectors can be set parameter by parameter rather than by a single conservative number.","feed_headline":"Mismatch computed right predicts when GW parameters turn biased","feed_subtitle":"Per-parameter bias SNRs from waveform-model geometry match full Bayesian posteriors for aligned-spin binaries.","key_machinery":"The load-bearing object is the normalized waveform distance $\\hat d = \\sqrt{M}$, built from the noise-weighted inner product, together with a decomposition of the total model error into a component tangent to the model manifold, which shifts parameters, and a component perpendicular to it, which only reduces recovered SNR. The central identity is the approximate Pythagorean relation $\\hat d^2_{\\mathrm{bias}} = M(h_{\\mathrm{bf}},h_s) \\approx M(s,h_s)-M(s,h_{\\mathrm{bf}})$, which lets the bias-inducing distance be read off as the difference between the faithfulness and effectualness mismatches. For individual parameters the machinery is the same relation applied along a one-parameter slice: fix the parameter of interest at its true value, optimize all others, and feed the distance between the unrestricted and restricted best-fit models into the one-degree-of-freedom chi-square criterion. This machinery converts the local geometry of the waveform model into a numerical SNR prediction for every parameter, which is exactly what the paper checks against Bayesian posteriors.","core_discovery":"The central discovery is that the mismatch criterion, calculated with the right distance and the right number of degrees of freedom, is an accurate predictor rather than a conservative bound. The right distance is $\\hat d^2_{\\mathrm{bias}} = M(h_{\\mathrm{bf}}, h_s)$, the mismatch between the model at the parameters that best fit the signal and the model at the true parameters, which isolates the part of model error that can produce a systematic parameter shift. For a set of $N$ parameters held fixed, Eq. (8) then places the true parameters on the $N$-dimensional 90% credible boundary at the predicted SNR; for one parameter, holding that parameter fixed and optimizing the rest gives a per-parameter bias SNR that similarly marks the boundary of the marginalized 90% interval. The paper verifies both statements with explicit posterior computations for aligned-spin binaries, and shows that the two distances obey the approximate Pythagorean relation $M(h_{\\mathrm{bf}},h_s) \\approx M(s,h_s)-M(s,h_{\\mathrm{bf}})$. It also argues that $\\hat d=\\sqrt{M}$, not $M$, is the quantity that scales linearly with SNR and with waveform phase and amplitude uncertainty, which relaxes the apparent accuracy requirement by a square root.","pith_inferences":["An implication the paper leaves implicit is that the direction of model error, not just its magnitude, becomes a design target: a model built so that its residuals are nearly orthogonal to the principal parameter directions would systematically push its bias SNRs upward, easing accuracy requirements.","The same per-parameter Pythagorean decomposition should transfer to any smooth parametric signal family with Gaussian high-SNR likelihoods, including extreme-mass-ratio inspirals or neutron-star waveforms, provided the posterior is unimodal.","A concrete extension would be to produce full parameter-space maps of per-parameter bias SNRs for current generic-binary models, turning them into trust-region surfaces that say, for a given mass-spin location, which SNR would bias which parameter; this is exactly the direction the paper flags for future work."],"forward_implications":["The standard faithfulness SNR is a lower bound on the bias SNR and can be low by an order of magnitude or more, so model-accuracy claims based on it understate how loud signals can be before individual parameters drift.","The $N$-dimensional bias SNR predicts when the joint posterior moves off the true parameters, while per-parameter bias SNRs predict when each marginalized parameter moves, and the two can differ by large factors.","Accuracy requirements for next-generation detectors should be expressed through $\\hat d = \\sqrt{M}$ and per-parameter bias SNRs; a conservative requirement of mismatch below about $10^{-6}$ for SNR $\\sim 1000$ can be relaxed substantially if model errors are directed perpendicular to the dominant parameter directions.","Fisher-matrix bias estimates agree with the bias-distance method when the signals are aligned in time and phase and the Fisher matrix is evaluated at both the true and best-fit parameters, making the two approaches interchangeable for practical bias-SNR estimation."],"supporting_citations":[{"why":"Introduces the mismatch indistinguishability criterion and the point that errors orthogonal to the model surface do not bias measurements.","marker":"[17]"},{"why":"Derives the chi-square generalisation of the indistinguishability criterion that the paper reinterprets with the bias distance.","marker":"[20]"},{"why":"Exemplifies the common $N/(2\\rho^2)$ criterion that the paper shows to be overly conservative for individual parameters.","marker":"[21]"},{"why":"Presents the equivalent Pythagorean and per-parameter bias-distance formulation that the paper derives and validates independently.","marker":"[22]"},{"why":"Earlier estimate of balance SNRs and detector accuracy needs that this work sharpens by separating faithfulness, effectualness, and bias distances.","marker":"[24]"},{"why":"Provides the Cutler-Vallisneri Fisher systematic-bias formalism used as a point of comparison.","marker":"[26]"},{"why":"Supplies the high-SNR Gaussian-likelihood argument and the split of waveform error into bias-inducing and non-bias-inducing parts.","marker":"[32]"},{"why":"Accurate surrogate model used to generate proxy true signals across parameter space for the bias-SNR survey.","marker":"[41]"}],"fun_headline_variants":["Corrected mismatch gives per-parameter GW bias SNRs","Beyond conservative mismatch: per-parameter bias SNR","Accurate GW bias SNRs via proper mismatch distances","Mismatch done right predicts when GW parameters bias"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the three waveform distances form a right triangle, with the error that shifts parameters perpendicular to the error that does not, so that subtracting one mismatch from the other yields the bias distance; if the model manifold is strongly curved or the parameter dependence nonlinear, that subtraction is only approximate.","fun_headline_variants_meta":{"raw":{"variants":["Corrected mismatch gives per-parameter GW bias SNRs","Beyond conservative mismatch: per-parameter bias SNR","Accurate GW bias SNRs via proper mismatch distances","Mismatch done right predicts when GW parameters bias"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000725,"raw_usage":{"total_tokens":3293,"prompt_tokens":1030,"completion_tokens":2263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":646,"completion_tokens_details":{"reasoning_tokens":2200}},"tokens_in":646,"tokens_out":2263,"duration_ms":21633,"temperature":1.0,"reasoning_tokens":2200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:24:03.942170+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a deliberately deficient model whose error is known to lie along a curved direction of parameter space, compute the per-parameter bias SNRs from the Pythagorean relation, and compare with a full Bayesian posterior run at the predicted SNR; a significant offset of the true parameter from the 90% boundary would disprove the relation's scope. A cheaper version is to measure the left- and right-hand sides of $M(h_{\\mathrm{bf}},h_s) \\approx M(s,h_s)-M(s,h_{\\mathrm{bf}})$ across a grid of high-mismatch signals and identify where the approximation fails.","supporting_citations":[],"review_version":1}