{"id":"033a7c0b-7a26-4fef-a56b-a8b8e2ffe00d","arxiv_id":"2511.11886","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"An extended gravitational-wave amplitude test adds (4,4) and (3,2) modes, reports a new δA44 constraint from GW230814, and shows precession and eccentricity can mimic GR violations.","lead":"This paper extends a gravitational-wave test of general relativity to two additional higher-order waveform modes and benchmarks it against noise, numerical-relativity simulations, and real detections. It finds the test works for aligned and mildly precessing binaries, reports a new constraint on the (4,4)-mode amplitude from GW230814, and warns that strong precession or eccentricity can mimic violations of GR.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Event-level δA44 claim for GW230814 is uncalibrated: by the paper's own selection threshold, GW230814's ρ⊥44 is marginal, and no injection-recovery for that event's parameters has been performed.","rationale":"The reader identified the small 20-realization calibration and the reliance on IMRPhenomXPHM waveform accuracy as the weakest assumptions. My concern agrees with the calibration issue but sharpens it: the most load-bearing consequence is that the headline real-event δA44 constraint is not calibrated for the event actually used. The paper contains no GW230814-specific injection-recovery test, and its own Appendix B demonstrates how sensitive δA44 posteriors are to noise and degeneracies in the weak-HOM regime. The fact that GW230814's ρ⊥44 lower bound is below the paper's own Appendix A selection threshold is a concrete, internal inconsistency: the event is presented as ideal and used to make a 'strongest constraint' claim even though the paper's stated selection criterion would not admit it. This does not require changing the overall conditional verdict — the analysis is transparent and the authors do caveat many of these issues — but it does mean the headline result should not be accepted as a robust null-test constraint until the event-specific calibration is supplied. The proposed test is straightforward and would settle whether the interval reflects mode information or prior/degeneracy effects.","tokens_in":20620,"tokens_out":4249,"duration_ms":45082,"concrete_test":"Perform a GW230814-like injection-recovery campaign: inject an IMRPhenomXPHM waveform with the maximum-likelihood parameters of GW230814 and δA44 = 0 into at least 100 independent Gaussian noise realizations with the same O4 PSDs, priors, and analysis settings used for the real event; recover δA44 in each realization; and compare the empirical coverage of the nominal 90% and 99% credible intervals with the expected rates, and the distribution of recovered interval widths with the real-event posterior. If δA44 = 0 falls outside the nominal CI in significantly more than the expected fraction, or if the recovered posteriors are as broad and bimodal as the GW230814 posterior, the quoted δA44 interval is not a calibrated constraint. Also verify whether GW230814 actually passes the Appendix A ρ⊥44 selection criterion; if its 68% lower bound is below 2.145, the event should be classified as too","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline result is the real-event constraint δA44 = −0.30^{+1.16}_{−3.45} for GW230814, and the conclusion that this establishes the SMA test as a validated null test. The load-bearing support for this is statistical calibration. But that calibration (Sec. III B, Fig. 3) uses only 20 Gaussian noise realizations for a single (32+8) M⊙ configuration with ρ⊥ℓm ≈ 3, and it already shows δA33 touching the 3σ boundary because of degeneracies with inclination and reference phase (Appendix B). No equivalent injection-recovery calibration is performed for the GW230814-like configuration that produces the headline δA44 interval. This matters because the paper itself shows in Appendix B that for GW250114-like signals with modest ρ⊥44, noise and degeneracies produce broad, bimodal δA44 posteriors that carry little information. GW230814's reported ρ⊥44 = 3.39^{+0.41}_{−1.26} has a 68% lower bound of 2.13, which is below the paper's own event-selection threshold of 2.145 stated in Appendix A. The quoted interval may therefore be dominated by prior volume, phase degeneracies, and noise fluctuations rather than by genuine (4,4)-mode amplitude information. Without an event-specific injection-recovery study, the claim that this is the 'strongest constraint' on δA44 is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an extension of the subdominant-mode amplitude (SMA) test of general relativity to the (3,2) and (4,4) modes, in addition to the previously considered (2,1) and (3,3) modes. The test is benchmarked through Gaussian-noise injections, numerical-relativity simulations (SXS), injection-recovery of amplitude deviations, and responses to phase-modified waveforms. The authors then apply the test to several O4 events, reporting a constraint on the (4,4) amplitude deviation from GW230814, δA44 = -0.30^{+1.16}_{-3.45}, which they describe as the strongest to date, and a constraint on δA33 from GW241011. The paper also demonstrates that waveform systematics can mimic GR violations for strongly precessing or eccentric binaries, and that the test responds to phase perturbations as well as amplitude perturbations.","tokens_in":20972,"tokens_out":2357,"duration_ms":23135,"significance":"If the statistical calibration and event-level constraints hold, the SMA test would be a useful null test of GR that is complementary to standard phasing tests, and the extension to (3,2) and (4,4) modes broadens its applicability to more symmetric and face-on binaries. The paper's systematic benchmarking against SXS waveforms and its explicit demonstration of systematics-induced biases in high-mass precessing systems are valuable contributions. However, the headline robustness claim rests on a calibration with only 20 noise realizations and on event-level results for systems whose mode SNR is marginal by the paper's own selection criterion, so the empirical validation is weaker than the abstract suggests.","major_comments":[{"comment":"The statistical calibration uses only N=20 Gaussian noise realizations for a single binary configuration. The p-p plot shows δA33 touching the 3σ contour and the 60% CI contains the injected value in >80% of runs, which is a notable deviation from expectation. Since the central claim that the SMA test is a validated null test rests on this calibration, 20 realizations is too small to establish the false-alarm rate, especially for δA33, whose bimodal degeneracy with reference phase and inclination (Appendix B) is shown to produce broad, over-covering intervals. The paper should either increase N substantially or present a quantitative uncertainty on the calibration curve and discuss how the δA33 behavior affects the interpretability of event-level δA33 constraints.","section":"Sec. III B, Fig. 3"},{"comment":"The headline event-level result, δA44 for GW230814, is selected using the criterion that the 68% lower bound of ρ⊥44 exceeds 2.145 (the 90th percentile of a χ2 distribution). The paper reports ρ⊥44 = 3.39^{+0.41}_{-1.26}, whose 68% lower bound is 2.13, below the stated threshold. Thus GW230814 does not satisfy the paper's own selection criterion. Moreover, no injection-recovery study is performed for a GW230814-like configuration; the only weak-HOM injection study (Appendix B, Fig. 14) is for GW250114 and shows that such configurations yield broad, bimodal δA44 posteriors dominated by degeneracies and noise. Without event-specific calibration, the quoted interval δA44 = -0.30^{+1.16}_{-3.45} cannot be interpreted as a meaningful constraint on the (4,4) amplitude, and the claim that it is the 'strongest constraint' is not established.","section":"Sec. V A, Appendix A"},{"comment":"The injection-recovery tests for amplitude deviations use IMRPhenomXPHM both to inject and to recover the signals. This is a valid check of internal consistency, but it does not probe the ability of the test to recover deviations when the template family is imperfect, which is the relevant systematic for real events. The SXS injections in Sec. III C partially offset this, but they are performed only for GR-consistent signals and do not include nonzero δAℓm. The paper should either acknowledge this limitation explicitly in the interpretation of Fig. 8 or add at least one cross-family injection-recovery (e.g., an SXS waveform with a modified subdominant mode) to demonstrate that the recovery of amplitude deviations is not an artifact of template self-consistency.","section":"Sec. IV A, Fig. 8"}],"minor_comments":[{"comment":"Typo: 'ampltiude' should be 'amplitude'.","section":"Sec. VI"},{"comment":"Typo: 'feect' should be 'effect'.","section":"Fig. 4 caption"},{"comment":"Typo: 'thode' should be 'those'.","section":"Fig. 14 caption"},{"comment":"The p-p plot would benefit from a legend or explicit statement that the shaded bands are 1σ, 2σ, and 3σ for N=20; the current shading is not self-explanatory.","section":"Fig. 3"},{"comment":"The definition of ρ⊥ℓm as an orthogonal SNR is clear, but the dependence on the noise-weighted inner product should be stated explicitly, including the normalization convention, since the χ2 null distribution in Appendix A relies on this.","section":"Sec. II, Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a strong claim about the first informative constraint on δA44, but the event in question is marginal by the paper's own cut and is not calibrated with event-specific injections. The N=20 noise calibration is also thin for the central robustness claim. I would encourage the editor to request either a larger calibration set or a clear downgrading of the event-level claims, and to verify that the 'strongest constraint' wording is warranted once the selection and calibration issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, incremental extension of the SMA null test to the (4,4) and (3,2) modes, and the first real-data (4,4) amplitude bound. The bound is real but weak, and the paper conditions it better in the body than in the abstract. I'd send it to a referee, but I'd ask for a revised abstract and a bit more calibration work.\n\nWhat's new: the extension to (4,4) and (3,2) opens the test to more mass-symmetric and face-on binaries. The benchmarking is genuinely useful—SXS injections covering aligned, precessing, and eccentric cases, plus a phase-deformation study showing the test responds coherently to phase perturbations. The authors are honest about the regimes where IMRPhenomXPHM fails: GW231123 and SXS:BBH:0623 show obvious systematic biases, and they flag them rather than brush them under the rug. The treatment of the δA33–reference-phase degeneracy in Appendix B is clear, and the GW250114 \"bad\" case is a good cautionary example.\n\nSoft spots, in order of severity. First, the event-level headline, δA44 = −0.30+1.16−3.45 for GW230814, sits in a regime that their own selection criterion (Appendix A) regards as marginal: the 68% lower bound on ρ⊥44 is 2.13, below the 2.145 threshold. For GW250114 they ran a 20-realization noise-injection study to show that broad, bimodal δA44 posteriors arise from degeneracies and noise; they didn't do the same for GW230814-like parameters. Without that, the claim that this is the 'strongest constraint' on δA44 is not really established—it's a first, plausibly prior-dominated bound. Second, the statistical calibration itself rests on N=20 Gaussian noise realizations, and the δA33 curve touches the 3σ contour. That's not fatal, but it's thin, and the paper acknowledges it. Third, the injection-recovery in Sec. IVA uses the same model for injection and recovery, so it doesn't probe waveform systematics for deformed amplitudes; the SXS tests partially cover that, but not for injected deviations.\n\nThe same-model issue is minor, the N=20 is a caveat, and the missing GW230814 calibration is the one thing I'd want fixed. No code or posterior samples are released, which will slow independent checks.\n\nWho's it for: the LVK tests-of-GR community and anyone building null tests with waveform models. Read the body, not just the abstract. It deserves a serious referee, and with a softened abstract and a modest injection study added, I'd be comfortable accepting it in substantially this form.","headline":"Solid incremental extension of the SMA test to (4,4)/(3,2) with an honest benchmark, but the headline δA44 constraint is thinner and less calibrated than the abstract claims.","tokens_in":21466,"tokens_out":3899,"would_cite":true,"duration_ms":36392,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An improved subdominant-mode amplitude test of general relativity, extended to the (4,4) and (3,2) modes, gives the strongest constraint yet on the hexadecapolar (4,4) mode amplitude, δA44 = −0.30^{+1.16}_{−3.45}, consistent with GR.","keywords":["general relativity tests","gravitational-wave higher-order modes","subdominant-mode amplitudes","binary black hole mergers","waveform systematics","spin precession","orbital eccentricity","Bayesian parameter estimation"],"falsifier":"Run the SMA test on hundreds of Gaussian noise realizations of a GW230814-like signal and require the fraction of runs with δA44=0 outside the 99% CI to be ≈1%; if the bimodality produces an inflated false-alarm rate, the quoted CI is too narrow. Separately, reanalyze GW230814 with an independent waveform model and check whether the δA44 posterior and the GR-consistency conclusion are unchanged.","tokens_in":20519,"feed_emoji":"🌊","tokens_out":5920,"duration_ms":50937,"temperature":0.7,"pith_summary":"This paper extends a test of general relativity that rescales the amplitudes of subdominant gravitational-wave modes — the (2,1), (3,3), (4,4), and (3,2) harmonics — while holding the dominant quadrupole mode fixed. The authors show that the test recovers injected deviations, produces null results for GR-consistent signals in Gaussian noise for aligned and mildly precessing binaries, and yields the strongest constraint to date on the (4,4) mode amplitude from a real event, δA44 = −0.30^{+1.16}_{−3.45}, consistent with GR. They also show the test is sensitive to phase deviations and that strong spin precession or orbital eccentricity can mimic apparent GR violations when unmodeled. A sympathetic reader would care because it validates a complementary, amplitude-based null test of strong-field gravity and maps the regime where it can be trusted.","feed_headline":"Gravitational-wave test reports strongest (4,4) mode bound","feed_subtitle":"Amplitude-based test finds GR consistent in black hole mergers, and shows where waveform errors can masquerade as violations.","key_machinery":"The SMA modification of the waveform: h → dominant quadrupole terms + Σ_HOM (1+δAℓm) hℓm, with each subdominant mode amplitude modified independently while the (2,2) quadrupole mode is kept fixed. The test's statistical power is carried by the orthogonal mode SNR ρ⊥ℓm, which measures how much signal cannot be explained by the dominant quadrupole mode, and by the one-at-a-time Bayesian estimation of each δAℓm. A load-bearing degeneracy is between δA33, the inclination angle, and the reference orbital phase: when the (3,3) mode is weak, the reference phase becomes bimodal and δA33 develops a secondary peak near −2 that can mimic a deviation. The (3,2) mode is singled out because it contributes","core_discovery":"The paper claims that an improved subdominant-mode amplitude (SMA) test, which lets only the amplitudes of the (2,1), (3,3), (4,4), and (3,2) modes float while fixing the quadrupole (2,2) mode, is a reliable null test of general relativity in the aligned-spin and mildly precessing binary black hole regime. Benchmarked on Gaussian noise injections and numerical-relativity waveforms, the test returns unbiased posteriors and GR-consistent Bayes factors for those systems. Applied to the events GW241011 and GW230814, it gives δA33 = 0.00^{+0.46}_{−1.82} and δA44 = −0.30^{+1.16}_{−3.45}, the latter the tightest published bound on the (4,4) mode amplitude deviation, both consistent with GR. The aut","pith_inferences":["A larger Gaussian-noise injection campaign (hundreds of realizations) would calibrate the bimodal δA33 and δA44 tails; the current 20-realization p-p plot leaves the quoted 99% false-alarm rate uncertain.","The test's phase sensitivity suggests a cheap diagnostic for catalog events: check the secondary spin posterior from the GR fit — near-extremal spins, as seen in the eccentric and phase-deformed injections, signal unmodeled physics rather than a real Kerr black hole.","A hierarchical combination of δAℓm posteriors across many events could turn the per-event null test into a population-level bound, with the (3,2) mode as the most phase-sensitive channel."],"forward_implications":["The (3,2) mode extends the test to near-equal-mass and face-on binaries, where the (2,1) and (3,3) modes are weak, so amplitude tests can now cover a larger share of detected black hole mergers.","The constraint δA44 = −0.30^{+1.16}_{−3.45} for GW230814 is the strongest published bound on a (4,4) amplitude deviation and is consistent with GR.","The δA33 posterior for GW241011, 0.00^{+0.46}_{−1.82}, likewise does not reject GR and is among the tightest constraints on that mode.","For strongly precessing or eccentric high-mass systems, apparent deviations reported by the test (e.g., for GW231123) should be interpreted as waveform-modeling systematics rather than evidence against GR.","Because the test also responds to phase perturbations, a nonzero δAℓm is not proof of an amplitude anomaly; it can flag phase-level deviations that the model does not include."],"fun_headline_variants":["Tightest (4,4) mode amplitude bound, consistent with GR","Improved SMA test: reliable for aligned spins, biased otherwise","Phase errors can trigger false GR violations in SMA test","Subdominant-mode amplitude test: robust null test with caveats","SMA test: GR consistent in 4,4 mode, but biased for precessing"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The event-level GR-consistency claims assume that the waveform model used for parameter estimation is an accurate GR template for those signals and that the deviation-parameter null distribution, calibrated from just 20 noise realizations, fully captures the degeneracies that can mimic deviations.","fun_headline_variants_meta":{"raw":{"variants":["Tightest (4,4) mode amplitude bound, consistent with GR","Improved SMA test: reliable for aligned spins, biased otherwise","Phase errors can trigger false GR violations in SMA test","Subdominant-mode amplitude test: robust null test with caveats","SMA test: GR consistent in 4,4 mode, but biased for precessing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000767,"raw_usage":{"total_tokens":3259,"prompt_tokens":785,"completion_tokens":2474,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":2381}},"tokens_in":529,"tokens_out":2474,"duration_ms":16650,"temperature":1.0,"reasoning_tokens":2381,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:07:18.268137+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the SMA test on hundreds of Gaussian noise realizations of a GW230814-like signal and require the fraction of runs with δA44=0 outside the 99% CI to be ≈1%; if the bimodality produces an inflated false-alarm rate, the quoted CI is too narrow. Separately, reanalyze GW230814 with an independent waveform model and check whether the δA44 posterior and the GR-consistency conclusion are unchanged.","supporting_citations":[],"review_version":1}