{"id":"e3f36747-f781-4f7d-8150-3f6bf497a3f8","arxiv_id":"2608.07237","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A flexible gravitational-wave test constrains spin-induced quadrupole and octupole moments, finds consistency with Kerr black holes in current data, and projects order 10^-2 to 10^-1 bounds for Einstein Telescope and Cosmic Explorer.","lead":"This paper builds a flexible, waveform-agnostic gravitational-wave test that looks for deviations in the spin-induced quadrupole and octupole moments of merging compact objects. Current LIGO-Virgo-KAGRA events are consistent with black holes, and the authors project that Einstein Telescope and Cosmic Explorer will squeeze these deviations to about one percent, sharp enough to tell black holes from neutron stars and exotic objects.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The XG forecast for |delta_kappa_a| rests on Fisher-matrix errors that the paper itself shows to be unreliable for the antisymmetric quadrupole parameter; the combined bound in Eq. (23) therefore needs a population-level validation before it can support the headline O(10^-2) claim.","rationale":"The reader's weakest-assumption analysis identifies exactly the load-bearing point: the FIM-based population forecast for delta_kappa_a is not reliable in cases the paper itself demonstrates are problematic. My reading of the manuscript confirms this. Section V A shows a direct Bayesian-versus-FIM disagreement for delta_kappa_a in the GW150914-like case, attributed to a pole in the delta_kappa_a coefficient, and a second disagreement when delta_kappa_s and delta_kappa_a vary simultaneously. The population forecast nevertheless uses single-parameter FIM errors for delta_kappa_a and combines them as though they were independent Gaussian standard deviations. The inverse-variance combination in Eq. (23) is especially sensitive to the smallest per-event variances, so even a small fraction of events with underestimated sigma_i can artificially tighten the combined bound. The paper's own argument that such cases cannot provide the best constraints addresses a single-event best bound, not a combined bound over the whole selected population. I do not think this warrants rejecting the paper. The observed-event analyses, the extensive injection studies, and the delta_kappa_s and delta_lambda_s forecasts are careful and transparent, and the authors explicitly acknowledge the FIM limitations. The appropriate verdict is CONDITIONAL, as the reader concluded: the delta_kappa_a component of the XG forecast needs additional validation before it can be quoted at face value. My stress-test does not move the verdict, so I mark the recommended verdict as UNCHANGED.","tokens_in":25859,"tokens_out":4780,"duration_ms":52074,"concrete_test":"Identify the 20 synthetic population events that contribute most to the inverse-variance sum for delta_kappa_a (rank by 1/sigma_i^2), run full Bayesian analyses with Bilby using the same SEOBNRv5HM ROM + FTI setup and XG PSDs, and recompute the combined 90% bound from the actual posterior distributions instead of Eq. (23). If the combined |delta_kappa_a| bound changes by more than a factor of about 2 relative to the quoted values 0.014 and 6.9e-3, the headline XG forecast for the antisymmetric quadrupole moment should be revised or explicitly restricted to single-parameter, well-validated regions of parameter space.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central forecast includes |delta_kappa_a| < 0.014 (ET-only) and |delta_kappa_a| < 6.9e-3 (ET+2CE), obtained by combining per-event Fisher information matrix errors with the inverse-variance formula in Eq. (23). The paper's own validation in Sec. V A, Figs. 15 and 16, shows that the FIM underestimates the error on delta_kappa_a for the GW150914-like injection because the delta_kappa_a coefficient in Eqs. (15)-(17) has a pole near the injected masses and spins, and that the FIM badly underestimates errors when delta_kappa_s and delta_kappa_a vary jointly. The caveat that the systems providing the best constraints should not suffer from this issue is not sufficient: Eq. (23) sums over all selected events, and per-event FIM errors that are too small for even a minority of population events can dominate the combined inverse-variance sum. The selection criteria (inspiral SNR, number of cycles, nonzero chi_eff) do not exclude events near such poles, and only two nearby injection points were validated with full Bayesian analyses against a population of 1249 or 4090 events spanning a wide range of masses, spins, and mass ratios. The delta_kappa_s and delta_lambda_s projections are better supported by the Bayesian checks, but the delta_kappa_a projection, which is part of the advertised quadrupole sensitivity, is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a test of spin-induced quadrupole and octupole moments of compact objects within the flexible theory-independent (FTI) framework, using the SEOBNRv5HM waveform model. The authors study measurability through Bayesian injections calibrated to GW150914-like and GW190412-like systems, apply the test to seven selected GWTC-3 events plus O4a posteriors, and present forecasts for next-generation detectors (ET and ET+2CE) using Fisher-matrix errors on a synthetic BBH population. The central claims are that current LVK data are consistent with Kerr BHs in GR, and that XG detectors will constrain |δκ_s| to O(10^-2), |δκ_a| to O(10^-2), and |δλ_s| to O(10^-1), about two orders of magnitude better than current constraints.","tokens_in":26165,"tokens_out":5530,"duration_ms":55359,"significance":"If the projections hold, this would be an important step toward routine tests of the Kerr nature of compact binaries, with direct implications for distinguishing black holes from neutron stars and exotic compact objects. The paper has notable strengths: the SIQM and SIOM coefficients are taken from independent PN derivations and are not circular; the injection studies are careful, including checks of tapering frequency, higher-order modes, and non-GR injections; and the comparison with LVK SIM results provides an explicit cross-model systematic check. The forecasts for δκ_s and δλ_s are supported by Bayesian/FIM agreement in representative cases. However, the forecast for δκ_a rests on Fisher errors that the paper itself shows to be unreliable for an important class of systems; this undermines one of the headline quadrupole constraints and needs to be fixed before the central XG claim is fully supported.","major_comments":[{"comment":"The combined bounds |δκ_a|<0.014 (ET-only) and |δκ_a|<6.9e-3 (ET+2CE) are obtained by inverse-variance weighting of per-event Fisher errors via Eq. (23). However, Sec. V A and Figs. 15 and 16 demonstrate that the FIM underestimates the error on δκ_a for the GW150914-like injection because the δκ_a prefactor in Eq. (15) has a pole near the injected masses and spins, and that the FIM also underestimates errors when δκ_s and δκ_a vary jointly. The selection criteria used in the population forecast (inspiral SNR>10, at least 5 inspiral cycles, χ_eff nonzero at 90% credibility) do not exclude events near such poles, and the inverse-variance sum in Eq. (23) can be dominated by a small number of events with underestimated σ_i. The claim that 'the systems that provide the best constraints should not suffer from this issue' is not sufficient, because all selected events enter the sum. The δκ_a projection therefore needs population-level validation, for example by Bayesian analysis of a representative subsample of selected events, or by restricting the forecast to systems where the δκ_a prefactor is not small; otherwise the δκ_a bounds should be removed from the headline claims.","section":"Sec. V B, Eq. (23)"},{"comment":"The paper validates the FIM against full Bayesian analyses for only two nearby injection points (GW150914-like and GW190412-like), while the synthetic population spans a wide range of masses, spins, and mass ratios and contributes 1249 (ET-only) or 4090 (ET+2CE) events to Eq. (23). Since the δκ_a prefactor in Eq. (15) can vanish for certain combinations of mass ratio and spins, a non-negligible fraction of the selected population may have non-Gaussian δκ_a posteriors even when their inspiral SNR and cycle counts pass the selection criteria. The authors should quantify how many selected events have small δκ_a prefactors, or otherwise demonstrate that the FIM errors for δκ_a are accurate for the events that dominate the combined inverse-variance sum.","section":"Sec. V A, Fig. 15"},{"comment":"The FIM calculations are performed with the gwbench package, but the manuscript does not state which waveform approximant gwbench uses. The injection and Bayesian recovery use SEOBNRv5HM ROM, so if gwbench uses a different waveform model (e.g., a TaylorF2 or IMRPhenom variant), the comparison between FIM and Bayesian errors could be affected by waveform systematics rather than by the Gaussianity assumption. This is particularly relevant for δκ_a, where the FIM already fails to reproduce the Bayesian width. Please specify the waveform model used in the FIM calculation and, if it differs from SEOBNRv5HM, verify the FIM results against SEOBNRv5HM for the representative points in Figs. 14-16.","section":"Sec. V A, Sec. V B"}],"minor_comments":[{"comment":"The phrase 'O(10−2) andO(10−1)' is missing a space after 'and'; please fix the LaTeX rendering.","section":"Abstract"},{"comment":"The priors for δκ_s, δκ_a, and δλ_s are stated, but the prior on δλ_a is not specified; please clarify whether δλ_a is always fixed to zero or given a prior in the analyses.","section":"Sec. II D"},{"comment":"The selection criterion 'χ_eff is nonzero at 90% credible level' should be made precise: does this mean that zero is excluded from the 90% credible interval, and is there a sign requirement on χ_eff?","section":"Sec. IV"},{"comment":"The quoted combined result δκ_s = −29^{+38}_{−54} should clarify whether this is the hyperparameter μ of the assumed Gaussian population distribution or a posterior-predictive value; the individual event posteriors often reach the prior boundary, so the dependence of this combined number on the prior range should be discussed.","section":"Sec. IV C"},{"comment":"The authors note that in the ET-only configuration the sky location was not recovered correctly in the Bilby runs due to the multibanded likelihood, but they assert that this does not affect intrinsic parameters. Since this assertion is used to justify the FIM/Bayesian comparison, a supplementary check showing that the intrinsic posteriors are unchanged between the correct and incorrect sky modes would strengthen the validation.","section":"Sec. V A"}],"recommendation":"major_revision","confidential_remarks":"The paper is well executed and the δκ_s and δλ_s forecasts appear solid, but the δκ_a forecast is a load-bearing part of the advertised O(10^-2) quadrupole sensitivity and is currently supported only by FIM errors that the paper itself shows to be unreliable in relevant parts of parameter space. I would ask the authors to either validate the population-level δκ_a bound with Bayesian analyses on a representative subsample, restrict the forecast to events with non-negligible δκ_a prefactors, or explicitly remove δκ_a from the combined XG claims. I do not see a circularity problem: the PN coefficients are independently derived and the method is benchmarked against injections and LVK results."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent, honest implementation paper. It does real service by putting SIQM/SIOM tests into the FTI framework with SEOBNRv5HM, including the 3.5PN SIQM term, and by cross-checking against the LVK IMRPhenom-based results. The observed-event analysis is clean and consistent with GR, and the delta_kappa_s and delta_lambda_s XG projections are plausible. The soft spot is the delta_kappa_a part of the forecast. The paper's own Bayesian-versus-FIM validation shows FIM badly underestimates the error for a GW150914-like system because of a pole in the delta_kappa_a coefficient near the injected masses and spins, and it also fails for joint delta_kappa_s/delta_kappa_a variation. That is not a minor footnote. The population bound then builds on per-event FIM errors summed via Eq. (23), so a minority of events with underestimated errors can tighten the combined delta_kappa_a bound more than the data justify. The paper's caveat that the best-constraining systems should not suffer is plausible but unverified; the selection criteria do not exclude pole-adjacent systems, and only two injections were validated with full Bayes. So I would treat |delta_kappa_s|<0.013 and |delta_lambda_s|<0.24 as fairly supported, and |delta_kappa_a|<0.014 as needing population-level Bayesian validation before it is quoted as headline sensitivity.\n\nOther soft spots are minor. There is no code or configuration release, which limits reproducibility. The hierarchical combination uses LVK FTI posteriors from another paper; that is fine methodologically but makes the combined result depend on that external work. The injection studies are zero-noise, standard for this kind of paper, but not a substitute for noise-realization checks. On the credit side, the paper is transparent about what is new and what already existed, it does not oversell the novelty, and it explicitly describes the test as a null test, demonstrating that large deviations are not measured accurately. The reliance on the authors' own FTI framework is not a circularity problem; that framework is published and the multipole coefficients used are independently derived.\n\nWho is this for: GW data analysts and theory groups building tests of the Kerr nature of compact objects, and anyone quoting XG sensitivity to this class of deviations. It deserves a serious referee. I would send it to review with a request that the authors either validate the delta_kappa_a population forecast with Bayesian runs across the selected-population parameter space or downgrade the delta_kappa_a claim to a FIM-only preliminary estimate.","headline":"A careful, honest FTI-based SIQM/SIOM implementation with solid injection checks, but the headline XG forecast leans on FIM errors for delta_kappa_a that the paper itself shows to be unreliable—worth refereeing, with that specific claim needing population-level validation.","tokens_in":26761,"tokens_out":1936,"would_cite":true,"duration_ms":20636,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["83C35","83C57"],"pacs":[],"model":"deepseek-v4-flash","headline":"Next-generation ground-based gravitational-wave detectors can constrain spin-induced quadrupole and octupole moment deviations to roughly $10^{-2}$ and $10^{-1}$, two orders of magnitude tighter than current bounds, allowing routine…","keywords":["gravitational waves","spin-induced quadrupole moment","spin-induced octupole moment","no-hair theorem","post-Newtonian approximation","parameterized tests of general relativity","next-generation ground-based detectors","Fisher information matrix"],"falsifier":"Run full Bayesian inference on the synthetic next-generation population for the events that dominate the combined bounds, and check whether the 90 percent credible intervals widen when non-Gaussian posteriors are used; the paper already demonstrates such a widening for the GW150914-like single-event $\\delta\\kappa_a$ case.","tokens_in":25608,"feed_emoji":"🔭","tokens_out":10412,"duration_ms":89635,"temperature":0.7,"pith_summary":"This paper argues that the spin-induced quadrupole and octupole moments of compact objects can be tested with gravitational waves through a flexible theory-independent framework that adds post-Newtonian phase corrections to an aligned-spin waveform. Applied to observed events, the test finds no deviation from the Kerr prediction. The paper's central forecast is that next-generation detectors such as Einstein Telescope and Cosmic Explorer will bound the quadrupole deviation to $|\\delta\\kappa_s| \\lesssim 10^{-2}$ and the octupole deviation to $|\\delta\\lambda_s| \\lesssim 10^{-1}$, roughly two orders of magnitude tighter than current constraints. If those forecasts hold, gravitational-wave observations will be able to tell whether a compact object is a Kerr black hole or something else, such as a neutron star or an exotic compact object.","feed_headline":"Gravitational waves could pin down black-hole 'hair' at 1 percent","feed_subtitle":"Next-gen detectors could beat current bounds by a hundredfold, testing whether black holes are truly Kerr.","key_machinery":"The central object is the flexible theory-independent phase correction $\\delta\\psi_{\\ell m}(f)$ added to the frequency-domain gravitational-wave phase. The deviation parameters enter through post-Newtonian coefficients: the quadrupole contributes at 2PN and 3PN, with a 3.5PN term shown to be negligible at current signal-to-noise ratios, and the octupole contributes at 3.5PN. A tapering function $W(f)$ smoothly switches the correction off before merger, and the framework's flexibility lets the user choose the baseline waveform and the taper location; the paper shows that higher tapering frequencies generally tighten the bounds. For the next-generation forecasts, the Fisher information matrix is the main forecasting tool, and the paper validates it against Bayesian inference for the single-parameter tests it uses.","core_discovery":"The central claim is that the spin-induced multipole moments of the binary components can be measured as fractional deviations from the Kerr values, with $\\delta\\kappa_s$ and $\\delta\\kappa_a$ for the quadrupole and $\\delta\\lambda_s$ for the octupole moments. Using a one-year synthetic population of merging black-hole binaries, the paper forecasts that a triangular Einstein Telescope alone will bound $|\\delta\\kappa_s|<0.013$, $|\\delta\\kappa_a|<0.014$, and $|\\delta\\lambda_s|<0.24$, improving to $|\\delta\\kappa_s|<6.5\\times10^{-3}$, $|\\delta\\kappa_a|<6.9\\times10^{-3}$, and $|\\delta\\lambda_s|<0.12$ when two Cosmic Explorer detectors are added. The paper establishes the test by building the corrections into the flexible theory-independent framework, validating Fisher-matrix error forecasts against full Bayesian analyses for selected cases, and applying the test to seven observed events from the first three observing runs plus later data. All observed results are consistent with general relativity, and the paper concludes that next-generation detectors will be able to routinely test the black-hole nature of compact binary coalescences.","pith_inferences":["If the Fisher-matrix failure the paper documents extends to the wider population, the combined bounds quoted for $\\delta\\kappa_a$ and for simultaneous variation of both quadrupole parameters are optimistic; a fully Bayesian population forecast would likely report wider error bars for those parameters.","The freedom to change the taper location is itself a diagnostic: comparing bounds from different taper choices across the same events would expose systematic modeling error, in the same way the paper compares two waveform approximants.","The XG-era population constraint will probably be set by the loudest high-spin systems rather than by the total event count, since one highly spinning event already outperforms stacked low-spin events; this favors detector designs that maximize detection of high-mass, high-spin inspirals.","The paper's quoted ranges for $\\kappa$ and $\\lambda$ across neutron stars, boson stars, and gravastars mean that future $\\delta\\kappa$ and $\\delta\\lambda$ bounds can be translated directly into statements about which classes of compact objects are excluded."],"forward_implications":["A triangular Einstein Telescope alone should bound $|\\delta\\kappa_s|<0.013$, $|\\delta\\kappa_a|<0.014$, and $|\\delta\\lambda_s|<0.24$ for one year of observations.","Adding two Cosmic Explorer detectors tightens those bounds by roughly a factor of two, to $|\\delta\\kappa_s|<6.5\\times10^{-3}$ and $|\\delta\\lambda_s|<0.12$, while detecting about four times as many usable events.","If those bounds hold, next-generation detectors can distinguish Kerr black holes from neutron stars and exotic compact objects, whose quadrupole and octupole parameters can differ from unity by orders of magnitude.","Combining many low-spin events adds little: a single highly spinning event can beat the combined population constraint, so the strongest tests will come from high-spin binaries.","The test detects injected quadrupole deviations as small as $\\delta\\kappa_s=\\pm2$, but for large deviations it recovers biased values, so it should be used as a null test rather than as a precise measurement of large non-Kerr moments."],"supporting_citations":[{"why":"Supplies the post-Newtonian spin-induced quadrupole phase corrections and the $\\delta\\kappa_s,\\delta\\kappa_a$ parameterization.","marker":"[3]"},{"why":"Introduces the flexible theory-independent method that adds PN coefficient deviations to the inspiral phase.","marker":"[26]"},{"why":"Provides the Fisher-matrix formalism used for the next-generation population forecasts.","marker":"[37]"},{"why":"The SEOBNRv5HM waveform model is the baseline general-relativity model used for injections and recovery.","marker":"[46]"},{"why":"The GWTC-3 test-of-general-relativity results whose quadrupole posteriors are directly compared with the new results.","marker":"[13]"},{"why":"The O4a parameterized-test posteriors that are combined with the GWTC-3 events in hierarchical inference.","marker":"[33]"},{"why":"The highly spinning event GW241011 that provides the tightest single-event spin-induced quadrupole constraint so far.","marker":"[14]"},{"why":"Earlier third-generation detector forecast for spin-induced moments whose bounds are compared with the present population results.","marker":"[17]"},{"why":"Provides the Einstein Telescope sensitivity curve used in the next-generation forecasts.","marker":"[78]"},{"why":"Provides the Cosmic Explorer sensitivity curves used for the two-CE network forecast.","marker":"[79]"}],"fun_headline_variants":["Gravitational waves could test black hole 'hair' at 1% level","Next-gen detectors to probe black hole multipole moments 100x tighter","Einstein Telescope may reveal if black holes obey no-hair theorem","GWs could catch black holes deviating from Kerr prediction","Testing black hole nature via spin multipole moments in gravitational waves"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quoted next-generation bounds rest on the assumption that Fisher-matrix errors accurately describe every selected population event, an assumption the paper itself shows fails for $\\delta\\kappa_a$ in a GW150914-like case and when $\\delta\\kappa_s$ and $\\delta\\kappa_a$ are varied together.","fun_headline_variants_meta":{"raw":{"variants":["Gravitational waves could test black hole 'hair' at 1% level","Next-gen detectors to probe black hole multipole moments 100x tighter","Einstein Telescope may reveal if black holes obey no-hair theorem","GWs could catch black holes deviating from Kerr prediction","Testing black hole nature via spin multipole moments in gravitational waves"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1662,"prompt_tokens":1104,"completion_tokens":558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":720,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":720,"tokens_out":558,"duration_ms":5594,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T12:04:31.739598+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run full Bayesian inference on the synthetic next-generation population for the events that dominate the combined bounds, and check whether the 90 percent credible intervals widen when non-Gaussian posteriors are used; the paper already demonstrates such a widening for the GW150914-like single-event $\\delta\\kappa_a$ case.","supporting_citations":[],"review_version":1}