{"id":"1d12ad2d-f995-45f5-bdd2-d9a16be4d8a0","arxiv_id":"2412.00320","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A framework is proposed to marginalize over black hole parameter uncertainties when setting exclusion limits on ultralight vector bosons from gravitational wave non-detections, with extensions to dark photon and dark Higgs models.","lead":"This paper develops a statistical way to turn the absence of a gravitational wave signal from a black hole 'boson cloud' into limits on invisible ultralight particles, even when the black hole's properties are uncertain. It also shows how those limits could probe unexplored dark photon and dark Higgs models with next-generation detectors.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Search-configuration grid built from independent marginal percentiles may not cover the joint signal parameter space, leaving Pdet calibration unvalidated beyond one synthetic event.","rationale":"The reader's weakest_assumption correctly identifies the search-configuration grid as the most load-bearing practical element: Pdet is the exclusion confidence, and the grid determines Pdet. However, the concern is not fatal to the central claim. The statistical framework in Eqs. (2)-(5) is a standard posterior-marginalized power calculation; the grid affects the numerical value of Pdet but not the validity of the mapping itself. Moreover, a suboptimal grid tends to lower Pdet, making the resulting exclusions conservative rather than overconfident. The paper's own check that more than 11 configurations do not change Pdet supports the grid's adequacy for the synthetic example, and the paper is explicit that the example is illustrative and that real constraints require event-specific studies. The dark-sector projections in Appendix B are flagged by the authors as order-of-magnitude estimates, which limits their precision but does not undermine the core methodology. Therefore no change to the ACCEPT verdict is warranted, though future applications should validate the grid construction on the actual remnant's posterior.","tokens_in":24078,"tokens_out":24062,"duration_ms":241420,"concrete_test":"For the GW170814-like posterior, draw 500 configurations directly from the joint distribution of (tstart, Tcoh, Tobs) using a covariance-aware sampling of the sample-derived optimal configurations (e.g., multivariate kernel density estimate), and recompute Pdet(mV) for the 11 mV values in Sec. IIIB. If any Pdet changes by more than the quoted beta-binomial 1-sigma error, the independent-percentile grid is inadequate. Additionally, apply the Appendix A construction to a second, qualitatively different CBC remnant (e.g., a high-mass, high-spin event such as GW190521) and check that the resulting 11 configurations recover injections from the full posterior at the nominal rate.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central mapping in Sec. IIC states that Pdet(mV), computed via Eq. (5), is the confidence level with which a null search excludes vector bosons of mass mV. This Pdet depends critically on the 11 search configurations constructed in Appendix A. The configurations are formed by taking values at the same percentiles from the marginal distributions of tstart, Tcoh, and Tobs independently, rather than from their joint distribution (Table II). Because the recovery is defined as success if at least one configuration finds the injection, any mismatch between this grid and the true joint distribution of optimal configurations for the posterior BH samples will bias Pdet. For the GW170814-like example the paper demonstrates that adding more than 11 configurations does not change Pdet, but this is a single event and a single convergence check; it does not establish that independent marginal percentiles adequately cover the joint space for other CBC remnants with different parameter correlations. If the grid under-covers the joint space, Pdet is underestimated and the quoted exclusion confidence is conservative rather than well-calibrated; if it over-covers high-recovery regions relative to the true search setup, Pdet could be overestimated. Either way, the claimed calibration of exclusion confidence is not robustly established for the general method.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a frequentist framework for interpreting null results in gravitational-wave follow-up searches for ultralight vector boson clouds around merger remnant black holes. Given the posterior distribution of remnant parameters from CBC parameter estimation, the authors define a marginalized detection probability Pdet(mV) (Eqs. 2-5), estimate it by injecting SuperRad waveforms into Gaussian noise realizations using public GW170814 posteriors, and interpret Pdet as the confidence level for excluding a vector of mass mV in the absence of a detection. The framework is demonstrated on a synthetic GW170814-like event, then extended to kinetically mixed dark photon and dark Higgs-Abelian sectors via the critical couplings εc, ρr, and ρs derived in Appendix B, with forecasts for next-generation detectors.","tokens_in":24329,"tokens_out":12316,"duration_ms":119819,"significance":"If the method is correct, it provides a practical and reproducible route from CBC follow-up null searches to quantitative boson-mass exclusions without requiring precise knowledge of the remnant's spin and mass. The paper's strengths include a transparent statistical construction, an injection study based on public posteriors and open software (LALSuite, SuperRad), explicit conservative assumptions in the dark-sector mapping, and falsifiable projections for next-generation detectors. The central derivation (Eqs. 2-5) is sound; the main open questions concern the generality of the search-configuration grid and the treatment of parameter uncertainty in the forecast section, rather than the statistical construction itself.","major_comments":[{"comment":"The central mapping from Pdet(mV) to exclusion confidence relies on the recovery rate being computed with a search configuration grid that adequately covers the joint distribution of optimal (tstart, Tcoh, Tobs) for the posterior BH samples. The grid is built by taking the same percentile from each marginal distribution independently rather than from the joint distribution, and the convergence check (that adding more than 11 configurations does not change Pdet) is performed only for the GW170814-like synthetic event. The statement in Appendix A that the procedure 'should be generally applicable to the majority of merger remnants' is therefore not supported by evidence beyond a single example. Please either add a quantitative coverage metric (e.g., the fraction of posterior samples whose optimal configuration lies within a chosen tolerance of the grid), demonstrate the procedure on additional synthetic remnants with different parameter correlations, or explicitly limit the claim to targets for which such a convergence test is performed.","section":"Sec. IIC and Appendix A (Table II, Fig. 6)"},{"comment":"The next-generation forecasts appear to use the horizon distance d_H from Ref. [62], which assumes the remnant parameters are known, together with catalog point estimates of (M, χ, d). This does not fold in the posterior parameter uncertainties that the rest of the paper emphasizes (Eqs. 2-5). Since parameter uncertainty acts as a mismatch that generally reduces the recovery probability relative to the perfectly-known case, the projected accessible parameter space in Fig. 4 is likely optimistic. Please either propagate the posterior-weighted Pdet into the horizon condition or state explicitly that the forecast neglects parameter uncertainty, and quantify the expected degradation by comparing, for a representative event, the Pdet from Eq. (5) with the perfect-knowledge recovery rate.","section":"Sec. IVC and Fig. 4"}],"minor_comments":[{"comment":"The claim that the percent error of the variance is ≲10% for NBH=200 is not backed by numerical values; please report the measured percent errors or show a quantitative convergence plot rather than relying on the visual comparison in Fig. 1.","section":"Sec. IIIA"},{"comment":"Because Nnoise=10, each Ndet(θi;mV) is an integer between 0 and 10, so the per-sample detection probability is coarsely quantized; please comment on whether this granularity affects the quoted 50% and 80% Pdet thresholds, and propagate the beta-binomial uncertainty to the boundaries of the excluded mass ranges.","section":"Sec. IIIB and Eq. (5)"},{"comment":"The statement that Pdet(mV) 'corresponds to the confidence level' should be stated as 1 minus the false-dismissal probability under the assumed noise and waveform models; the equivalence is exact only when the detection threshold and false-alarm probability are fixed, as done here, and would be clearer if phrased in those terms.","section":"Sec. IIC"},{"comment":"The safety factor δ=10^2 in Eq. (B14) is conservative but no sensitivity study is provided; since the accessible parameter region in Fig. 4 depends on ρs, please indicate how the curves shift for other choices such as δ=10 or δ=10^3.","section":"Sec. IVB and Appendix B"},{"comment":"The caption states that error bars are 1σ beta-binomial uncertainties but does not give the formula or the assumed prior; please add a reference or an explicit expression for the beta-binomial interval used.","section":"Fig. 2 caption"},{"comment":"The text says the mass range mV ∈ [0.6,1.1]mopt_V was 'chosen empirically,' and the subsequent claim that BHs within two standard deviations of the mean have optimally matched boson masses within this range is not demonstrated; please show the distribution of mopt_V over the NBH=200 posterior samples relative to the 11-value mV grid.","section":"Sec. IIIB"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a methodology paper that builds substantially on the authors' own prior work (Refs. [62] and [66]), which is natural in this subfield. The main risk is not the statistical construction but the generality of the configuration-grid prescription and the apparent neglect of parameter uncertainty in the forecast; both are fixable within the scope of a revision. The editor may wish to weigh whether a single synthetic demonstration is sufficient for the broad claims made in the abstract and Appendix A."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read.\n\nThe genuinely new thing is the posterior-marginalized injection framework in Sec. II. Instead of assuming perfectly known BH parameters as in the earlier search paper, they draw from the CBC posterior, inject SuperRad waveforms into Gaussian noise, and define Pdet as the recovery rate averaged over posterior samples and noise realizations. That closes an obvious gap, and Eqs. (2)–(5) give a clean, frequentist construction that doesn't smuggle in the target result. The demonstration on a GW170814-like event is honest: public posteriors, LALSuite, SuperRad, and beta-binomial uncertainties. The mappings to kinetically mixed dark photons and dark Higgs sectors are new, and the thresholds are deliberately conservative.\n\nSoft spots, in proportion. The search-configuration grid in Appendix A takes the 2nd, 10th, ..., 98th percentiles of tstart, Tcoh, and Tobs independently, not from their joint distribution, and defines recovery as success if any of the 11 configurations finds the injection. The paper shows for this one event that adding configurations doesn't change Pdet; that's a reasonable convergence check, but it's one event. If the independent grid over-covers high-recovery regions relative to the true joint distribution, Pdet would be overestimated; if it under-covers, the exclusion is conservative but miscalibrated the other way. I don't think this is fatal—the method is per-target, so a real follow-up would construct the grid from that event's posterior and check convergence again. But the general claim of well-calibrated exclusions would be stronger with a second event or a population-level check. The paper should either soften that claim or add a joint-distribution grid as a cross-check.\n\nThe other caveats—idealized noise, no data gaps, hand-picked beta2 and delta, order-of-magnitude dark-sector projections—are acknowledged in the text and are acceptable for a methodology paper as long as the projections aren't over-sold. They mostly aren't.\n\nWho this is for: people designing follow-up searches for boson clouds around merger remnants, and anyone building exclusion pipelines from null CW searches. It's a practical, citable framework. It deserves a serious referee; with a modest revision it would be a solid addition to the literature.\n\nRecommendation: send to peer review.","headline":"Posterior-marginalized exclusion framework is a real advance over perfectly-known-parameter searches; the main caveat is a configuration grid validated on a single synthetic event.","tokens_in":24868,"tokens_out":3407,"would_cite":true,"duration_ms":30405,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A null gravitational-wave search for a superradiance cloud around a merger remnant black hole can be converted, through a detection probability marginalized over the remnant's uncertain parameters, into a calibrated exclusion of…","keywords":["ultralight vector bosons","black hole superradiance","gravitational waves","merger remnant black holes","hidden Markov model search","dark photon","dark Higgs sector","frequentist upper limits"],"falsifier":"Run the same injection study on a second, different remnant (for example a heavier or more distant one), drawing search configurations from the joint distribution of start time, coherent time, and observing time instead of from independent percentiles; if the recovery rates disagree with the Appendix A grid beyond the quoted error bars, the $P_{\\rm det}$ calibration is not robust.","tokens_in":23856,"feed_emoji":"🕳️","tokens_out":11699,"duration_ms":97556,"temperature":0.7,"pith_summary":"The paper proposes a way to turn a null gravitational-wave search for a superradiant vector-boson cloud around a binary-merger remnant black hole into a calibrated exclusion interval on the boson mass. The key move is to treat the injection recovery rate, averaged over posterior samples of the remnant's uncertain mass, spin, distance, and orientation, as the confidence with which the absence of a signal rules out each mass. The method is demonstrated on a synthetic event similar to the detected merger GW170814, where a non-detection would exclude vector masses in a band around $4.6\\times10^{-13}$ eV at 80% confidence, and it is then extended to kinetically mixed dark photons and a dark Higgs sector. The paper stresses that no real constraints are derived yet; the contribution is the procedure, plus a forecast that next-generation detectors can probe previously unconstrained parameter space with a handful of follow-up targets.","feed_headline":"Null black-hole search can rule out ultralight bosons","feed_subtitle":"New method turns merger-remnant uncertainty into calibrated boson-mass bounds.","key_machinery":"The load-bearing object is the marginalized detection probability of Eq. (5), $P_{\\rm det}(m_V)=\\frac{1}{N_{\\rm BH}N_{\\rm noise}}\\sum_{i=1}^{N_{\\rm BH}}N_{\\rm det}(\\theta_i;m_V)$, where $\\theta_i$ are posterior samples of the remnant black hole parameters and $N_{\\rm det}$ counts recoveries in noise realizations. This identity is what folds parameter uncertainty into a frequentist exclusion. The underlying search is the hidden Markov model (HMM) tracking scheme of [62], which follows the quasi-monochromatic, upward-drifting signal over coherent segments. The second piece is the search-configuration grid of Appendix A: eleven configurations built by taking percentiles of the start time, coherent time, and observing time distributions independently, with an injection counted as recovered if at least one configuration finds it. The third piece is the SuperRad waveform model, a numerical waveform model for gravitational waves from superradiant vector clouds; it generates the injected signals and supplies the frequency-evolution information used to pick coherent times.","core_discovery":"On the paper's own terms, the central claim is that the detection probability $P_{\\rm det}(m_V)$, defined as the fraction of SuperRad injections recovered in $N_{\\rm noise}$ noise realizations averaged over $N_{\\rm BH}$ samples drawn from the compact-binary-coalescence posterior, is the confidence level with which a null search excludes vector bosons of mass $m_V$. This makes Eq. (5) into a frequentist statement: no detection means the existence of vectors with mass $m_V$ is excluded at confidence $P_{\\rm det}(m_V)$. Using $N_{\\rm BH}=200$ posterior samples and $N_{\\rm noise}=10$ noise realizations per mass value, the paper shows the marginalized recovery rate is stable and can be computed for the 11 masses in the promising window $m_V\\approx[0.6,1.1]m_V^{\\rm opt}$. It further establishes that the same null result remains valid for a kinetically mixed dark photon only for kinetic mixing $\\epsilon<\\epsilon_c$, and for a dark Higgs sector only for couplings $\\lambda v^4>\\max(\\rho_r,\\rho_s)$, with explicit conditions derived in Appendix B.","pith_inferences":["Editorial inference: the same posterior-marginalized recovery-rate prescription could be applied to other superradiance probes, such as black hole spin measurements or stochastic background searches, to convert their parameter uncertainties into boson-mass exclusions.","Editorial inference: the Appendix A grid takes percentiles of each search parameter independently, so its calibration could be checked by drawing a grid from the joint distribution of start time, coherent time, and observing time; the paper validates the grid only on the synthetic GW170814-like example.","Editorial inference: counting an injection as recovered when any one of eleven configurations finds it makes $P_{\\rm det}$ an upper envelope over configurations, so a real search that commits to a single configuration may see lower recovery and weaker exclusion.","Editorial inference: with next-generation detectors the number of high-signal-to-noise remnants grows, so the same method could map the vector mass range continuously across a population rather than around individual remnant masses."],"forward_implications":["A null follow-up of a detected compact-binary merger remnant can be reported as a confidence level on $m_V$, with the remnant's parameter uncertainty already included rather than treated as a systematic caveat.","The computation is feasible: 200 posterior samples and 10 noise realizations per mass value suffice for the synthetic target, because $N_{\\rm BH}=200$ reproduces the posterior variance to about 10%.","The same null result doubles as a bound on a kinetically mixed dark photon for $\\epsilon<\\epsilon_c$ and on a dark Higgs sector for $\\lambda v^4>\\max(\\rho_r,\\rho_s)$.","Projections using the currently detected population of remnants show that next-generation detectors can reach dark photon parameter space unconstrained by cosmic microwave background observations, up to kinetic mixing $\\epsilon\\sim10^{-6}$, and Higgs couplings down to $\\lambda^{1/4}v\\sim\\mathcal{O}(10)$ MeV.","Improved frequency-evolution models would refine the search configuration but, as the paper notes, do not change the qualitative conclusions."],"supporting_citations":[{"why":"Supplies the hidden Markov model long-transient search method that the marginalization procedure builds on.","marker":"[62]"},{"why":"Provides the SuperRad waveform model used to generate every injected synthetic signal.","marker":"[66]"},{"why":"Defines the kinetic-mixing regimes and the pair-plasma dissipation model used to set the critical mixing $\\epsilon_c$.","marker":"[72]"},{"why":"Provides the dark-Higgs self-interaction channel and scaling used to estimate the modified frequency evolution and $\\rho_r$.","marker":"[69]"},{"why":"Supplies the string-formation threshold applied to define the critical coupling $\\rho_s$.","marker":"[74]"},{"why":"Supplies the GW170814 posterior samples used to build the synthetic target and to test the required sample size.","marker":"[87]"},{"why":"Provides the cosmic microwave background constraint used as the baseline for the next-generation detector projections.","marker":"[100]"}],"fun_headline_variants":["Null GW search method yields ultralight boson limits","Merger remnant silence can set vector boson bounds","Marginalizing black hole uncertainty tightens boson constraints","No signal from mergers can constrain ultralight bosons","Next-gen GW detectors could break boson model degeneracies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the calibration of the detection probability: the search-configuration grid is built by taking percentiles of each search parameter separately and is validated on a single synthetic event, so if a real signal would not be caught by that grid, the confidence quoted from a null search would be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Null GW search method yields ultralight boson limits","Merger remnant silence can set vector boson bounds","Marginalizing black hole uncertainty tightens boson constraints","No signal from mergers can constrain ultralight bosons","Next-gen GW detectors could break boson model degeneracies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2190,"prompt_tokens":954,"completion_tokens":1236,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":1156}},"tokens_in":570,"tokens_out":1236,"duration_ms":11810,"temperature":1.0,"reasoning_tokens":1156,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T05:30:19.488891+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same injection study on a second, different remnant (for example a heavier or more distant one), drawing search configurations from the joint distribution of start time, coherent time, and observing time instead of from independent percentiles; if the recovery rates disagree with the Appendix A grid beyond the quoted error bars, the $P_{\\rm det}$ calibration is not robust.","supporting_citations":[],"review_version":1}