{"id":"5f4fa935-e3b9-4f0a-9ff9-3f8057886e7c","arxiv_id":"2608.07312","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"No continuous gravitational-wave signal was found in any of the 562 radiometer-triggered candidates after vetoes and an O4b follow-up; the search reaches 95% detection efficiency for strains around 0.63 to 6.3 times 10^-25.","lead":"Astronomers followed up 562 frequency-sky patches where LIGO's O4a radiometer search reported sub-threshold hints of persistent gravitational waves. After vetoes and a re-check in later O4b data, no hint survived, so they report no continuous-wave detection and measure the search's sensitivity with injected signals.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"O4b injection-recovery validation is reported only as 'vast majority'; without per-candidate counts, the non-recovery of all 21 outliers in O4b cannot be interpreted as ruling out real signals.","rationale":"The paper is generally careful and the null result is plausible. The outlier count is consistent with the 5% false-alarm expectation (24 vs 27±5), and the vetoes are physically motivated. The off-target exponential tail model is the reader's identified weakness, but it is empirically calibrated by this count consistency, and contamination of off-target samples by real astrophysical signals is negligible because the F-statistic's sky response is narrow (the off-target samples are ≥10 deg away in sky and share only the frequency band). The more load-bearing step is the O4b confirmation: the conclusion that none of the 21 outliers persist is used in the abstract and Sec. V C to conclude no convincing signal. That conclusion depends on O4b having comparable sensitivity for each of those specific candidates. The injection test at 110% h95 is a reasonable validation, but 'vast majority' is too vague; without exact per-candidate counts, the reader cannot verify that a real signal in one of the outlier candidates would have been recovered in O4b. This is a concrete, checkable gap, not a fatal flaw. I therefore recommend accepting conditionally on reporting the exact injection-recovery statistics.","tokens_in":24517,"tokens_out":23977,"duration_ms":201543,"concrete_test":"Locate the O4b injection-recovery results for the 21 outliers and report the exact number of injections recovered in O4b out of 210, plus per-candidate counts. For any candidate with failed injections, compare the failed injections' frequency, sky position, and fdot to the outlier's parameters; also check whether the outlier's O4b statistic is consistent with the distribution of non-recovered injections. If the failure rate is non-negligible or the failed injections resemble the outlier, the O4b non-recovery cannot be used to dismiss that outlier.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central conclusion in Sec. V C rests on the claim that the 21 O4a outliers are not recovered in O4b, interpreted as evidence against a real CW signal. This interpretation requires that a signal detectable in O4a would also be detectable in O4b for each of the 21 candidates. The only validation offered is the statement that 'the vast majority' of the 210 simulated signals (10 per candidate, injected at 110% of h95) are detected in both O4a and O4b. No exact recovery counts or per-candidate results are given. If even a few candidates had failed injections, an outlier in one of those candidates could be a real signal that O4b is simply too insensitive to see; the non-recovery would then be meaningless. Moreover, the outliers' amplitudes are unknown and may lie below h95, where O4b detection efficiency is lower than the 110% h95 test. The O4b follow-up is thus a load-bearing part of the null argument, and its validation is currently unquantified.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a follow-up search for continuous gravitational waves targeting 562 sub-threshold candidates from the O4a LVK all-sky all-frequency radiometer analysis. For each candidate, the authors run an F-statistic search over a 1/32 Hz band and a ~13 sq-deg sky pixel, with a hidden Markov model to track frequency wandering. Using O4a LIGO data, they find 24 outliers with trials-corrected false-alarm probability below 5%; three are removed by detector-artifact vetoes, leaving 21. These 21 outliers are re-analyzed in independent O4b data, and none are recovered at comparable significance. The authors therefore conclude that no convincing CW signals are present. They also report extensive injection-recovery sensitivity estimates: 95% detection efficiency at h0 ~ (0.63-6.3)e-25 for isolated neutron stars and h0 ~ (0.82-7.1)e-25 for long-period binaries, plus validation against four of five pulsar hardware injections.","tokens_in":24774,"tokens_out":8361,"duration_ms":78454,"significance":"If the result holds, the paper demonstrates that a model-agnostic radiometer pipeline can be systematically followed up with an HMM-based coherent search over hundreds of candidates, and it places meaningful strain upper limits in the 20-1726 Hz band. The strengths of the manuscript are its large empirical injection campaign (500 injections per candidate), the use of external O4b data for an independent cross-check, validation with hardware injections, and the transparent reporting of outlier parameters and false-alarm estimates. The main null conclusion is also supported by the binomial expectation: 24 outliers from 533 clean candidates is consistent with the 27 +/- 5 expected at a 5% false-alarm threshold. The principal weakness is that the O4b validation is reported only as \"the vast majority\" of simulated signals being recovered, without per-candidate counts; this leaves the O4b non-recovery argument, which the paper uses to rule out real signals, under-quantified.","major_comments":[{"comment":"The O4b follow-up is load-bearing for the statement that the 21 O4a outliers are not consistent with astrophysical signals, but its validation is not quantified. The text says only that \"the vast majority\" of the 210 simulated signals (10 per candidate, injected at 110% of h95) are confidently detected in both O4a and O4b, without giving exact counts or a per-candidate breakdown. If some candidates had low O4b recovery efficiency, then an O4a outlier in one of those candidates could be a real signal that O4b is simply too insensitive to detect; the non-recovery would not be informative. In addition, the injection amplitude is fixed at 110% of the estimated 95% efficiency strain, while the amplitudes of the actual outliers are unknown and may lie below h95, where O4b detection efficiency is lower. I request that the authors report, for each of the 21 candidates, the number of injections detected in O4a and O4b, define \"confidently detected\" (e.g., above the O4b threshold with a specified margin), and discuss how the results would change if a few candidates had substantially lower O4b recovery. This is a local, fixable issue that does not invalidate the overall null result, especially since the outlier count is already consistent with noise, but it is necessary to fully support the O4b-based conclusion.","section":null}],"minor_comments":[{"comment":"The per-candidate h0 ranges are said to be \"adjusted according to the average noise level,\" but the adjustment rule is not specified. Please state how the sampled strain range is chosen and confirm that h95 falls within the sampled range for all candidates, so that the logistic regression is not extrapolating.","section":null},{"comment":"The caption contains a typo: \"Cumululative density\" should be \"Cumulative density.\"","section":null},{"comment":"The off-target false-alarm procedure assumes that no signal or unmodeled artifact exists near the off-target templates. The manuscript states this assumption, but it would help to quantify the risk: for example, by reporting how many off-target samples lie within a known spectral artifact and whether excluding those samples changes the fitted lambda_hat.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid null-result follow-up with an extensive empirical validation. The only substantive issue is the unquantified O4b injection-recovery validation; with a per-candidate table of recovery counts and a clear definition of 'confidently detected,' the O4b argument would be complete. I do not see a need for re-review beyond the revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, transparent null result, and the reader's ACCEPT is about right. The new content is the O4a application of the O3 HMM follow-up pipeline to 562 ASAF candidates, the 21 surviving outliers, their non-recovery in O4b, and per-candidate sensitivity estimates including long-period binaries. The outlier count (24 from 533 clean candidates) matches the 27±5 expected at 5% false-alarm, three are vetoed as H1 artifacts, and the most significant outlier sits near the 60 Hz mains. That combination is a clean null argument.\n\nWhat it does well: methods are unusually detailed; the disturbed-candidate cut (Ltail > 7.5) is validated against the known-lines list; sensitivity comes from 500 injections per candidate plus four of five hardware injections recovered. No code is shipped, but the description is enough to reimplement.\n\nSoft spots: the stress-test note about O4b is fair but not fatal. The paper says only that 'the vast majority' of 210 injections at 110% h95 were recovered in both O4a and O4b, with no per-candidate counts. Since O4b non-recovery of all 21 outliers is presented as a check, a reader would want exact numbers. But the null does not stand on O4b alone; the outlier-count consistency and vetoes do the load-bearing work. The off-target exponential tail model is an assumption, but it is disclosed and standard for this pipeline family. The Ltail cut is data-driven, though the cross-check against vetted lines makes it credible. Minor: the 2.3 polarization-averaging factor is empirical, and the binary sensitivity covers only Pb > 1 yr.\n\nWho it is for: people doing CW follow-up and ASAF candidate work; it closes the O4a candidate list and provides O4a sensitivity benchmarks. It is an extension rather than a new method, but at-scale deployment is worth recording.\n\nRecommendation: send it to peer review. A serious referee should ask for the O4b injection counts and perhaps a per-candidate table, but the analysis is sound.","headline":"Careful O4a follow-up of 562 radiometer candidates; the null result is supported by outlier-count consistency, though the O4b validation is under-quantified.","tokens_in":25317,"tokens_out":2196,"would_cite":true,"duration_ms":18970,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["04.80.Nn","95.55.Ym"],"model":"deepseek-v4-flash","headline":"Following up 562 radiometer candidates, this search finds no continuous gravitational waves in LIGO O4a data.","keywords":["continuous gravitational waves","hidden Markov model","F-statistic","radiometer search","all-sky all-frequency","LIGO O4","neutron stars","gravitational wave data analysis"],"falsifier":"Re-run the identical pipeline on the 21 surviving outliers using O4b data with the C01 calibration and a lower false-alarm threshold; recovery of any of them at $p_{\\rm corr}^{\\rm fa}<5\\%$ would contradict the claim that all outliers are noise. A second check is to inject a simulated continuous wave at the quoted 95% sensitivity strain amplitude into one of the 562 O4a frequency-sky pixels and verify recovery; failure to recover it would falsify the sensitivity claim.","tokens_in":24265,"feed_emoji":"🔭","tokens_out":7285,"duration_ms":64212,"temperature":0.7,"pith_summary":"This paper follows up 562 sub-threshold candidates from the O4a all-sky all-frequency radiometer search, using an F-statistic matched filter with a hidden Markov model that lets the signal frequency wander between segments. Its central result is a null: 24 candidates produce outliers with false-alarm probability below 5%, three are vetoed as detector artifacts, and the remaining 21 are not recovered at comparable significance in independent O4b data. The paper concludes that no convincing continuous gravitational wave is present in these 562 frequency-sky pixels, and that all outliers are consistent with noise fluctuations. It also establishes sensitivity: 95% detection efficiency for strain amplitudes $h_0 \\sim (0.63\\text{--}6.3)\\times10^{-25}$ for isolated neutron stars and $h_0 \\sim (0.82\\text{--}7.1)\\times10^{-25}$ for long-period binaries across 20–1726 Hz. The search matters because it targets signals with irregular phase evolution that standard phase-coherent searches can miss.","feed_headline":"Null result: 562 candidate spots hide no continuous gravitational waves","feed_subtitle":"None of 21 outliers recur in independent O4b data, and injected signals show the search sees down to h0 ~ 6e-26.","key_machinery":"The central machinery is the F-statistic matched filter run in 24-hour coherent segments, combined with a hidden Markov model and the Viterbi algorithm to select the most probable segment-wise frequency path for each sky-position and spin-down template. The detection statistic is the normalized Viterbi log-likelihood $\\bar{L}=L/N_T$. Significance is assigned by fitting an exponential tail to the distribution of detection statistics at off-target sky positions, then applying a trials-factor correction to obtain false-alarm probabilities. Three vetoes—known spectral lines, single-interferometer power imbalance, and Doppler-modulation-off consistency—filter detector artifacts, and a large injection campaign with logistic-regression fitting yields the reported 95% efficiency strain amplitudes.","core_discovery":"The paper claims that no convincing continuous gravitational wave is detected in a directed follow-up of 562 sub-threshold candidates produced by the O4a LIGO–Virgo–KAGRA all-sky all-frequency radiometer analysis. Each candidate is a narrow $1/32$ Hz frequency band paired with a roughly $13$ deg$^2$ sky pixel. Searching O4a data from the two Advanced LIGO detectors with an F-statistic matched filter equipped with hidden Markov model frequency tracking, the paper finds 24 outliers below a trials-corrected false-alarm probability of 5%; three fail instrumental vetoes, and the 21 surviving outliers are not recovered with comparable significance when re-analyzed in O4b data. The observed outlier count (24) is consistent with the binomially expected number ($27\\pm5$) at the 5% threshold, supporting the interpretation that all outliers are noise fluctuations. Injection-recovery tests place the search's 95% efficiency sensitivity at $h_0\\sim(0.63\\text{--}6.3)\\times10^{-25}$ for isolated neutron stars and $h_0\\sim(0.82\\text{--}7.1)\\times10^{-25}$ for neutron stars in long-period ($P_b>1$ yr) binaries.","pith_inferences":["Editorial inference: the null result suggests, though the paper does not claim, that sub-threshold radiometer candidates in O4a are dominated by scattered noise artifacts rather than by a population of missed quasi-monochromatic sources.","Editorial inference: because the off-target noise model is the load-bearing premise, a natural test is to inject weak signals into off-target sky positions in O4b and measure how much the fitted exponential tail shifts; the published off-target distributions could be reused for this check.","Editorial inference: the search's long-period binary sensitivity is capped by the HMM's one-frequency-bin-per-segment drift limit, so a higher-order frequency-jump model could probe whether rapidly wandering sources are being missed.","Editorial inference: combining O4a and O4b data with a longer coherence time for the same 562 candidates would likely push the quoted strain sensitivity lower, since the paper demonstrates the pipeline can be deployed at scale but does not run this combined search."],"forward_implications":["The 562 followed-up frequency-sky pixels are effectively excluded as sites of a detectable continuous wave at the quoted strain amplitudes in O4a data.","Because none of the 21 surviving outliers persists in O4b data, any claimed detection at these parameters would have to explain why the signal vanished in later data; the paper's injection recovery shows a real detectable signal would persist.","The outlier count matches the expected noise count at the 5% false-alarm threshold, so the search behaves consistently with its stated statistical calibration.","The sensitivity improvement over the earlier O3 follow-up and over the ASAF upper limits sets a new reference point for radiometer-triggered continuous wave searches at O4a sensitivity."],"supporting_citations":[{"why":"Identifies the 562 O4a sub-threshold ASAF candidates that this paper follows up.","marker":"[51]"},{"why":"Introduces the HMM-based F-statistic follow-up pipeline and its veto framework, which this work adapts.","marker":"[65]"},{"why":"Establishes the hidden Markov model formalism for tracking a wandering-frequency continuous wave signal.","marker":"[81]"},{"why":"Applies the HMM and Viterbi algorithm to continuous wave searches and defines the Viterbi log-likelihood statistic.","marker":"[82]"},{"why":"Defines the single-interferometer veto criterion used to identify outliers consistent with one detector's noise.","marker":"[86]"},{"why":"Supplies the vetted narrowband spectral artifact list used by the known-lines veto.","marker":"[85]"},{"why":"Provides the public O4a calibrated strain data used for the primary search.","marker":"[67]"},{"why":"Provides the public O4b strain data used for the independent secondary follow-up.","marker":"[68]"}],"fun_headline_variants":["Null result: no continuous gravitational waves in 562 O4a candidates","Follow-up search of 562 candidate spots finds no CW signals","Sensitive to 6e-26 strain, 562 O4a candidates show no CWs","No convincing continuous gravitational waves from O4a candidates","Search of 562 radiometer candidates yields no CW detections"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The false-alarm probabilities come from fitting the tail of the detection statistic to off-target sky positions, under the assumption that no real signal or unmodeled artifact contaminates those off-target positions; if that assumption fails, the outlier list and the conclusion drawn from it could change.","fun_headline_variants_meta":{"raw":{"variants":["Null result: no continuous gravitational waves in 562 O4a candidates","Follow-up search of 562 candidate spots finds no CW signals","Sensitive to 6e-26 strain, 562 O4a candidates show no CWs","No convincing continuous gravitational waves from O4a candidates","Search of 562 radiometer candidates yields no CW detections"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0009,"raw_usage":{"total_tokens":3976,"prompt_tokens":1150,"completion_tokens":2826,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":766,"completion_tokens_details":{"reasoning_tokens":2732}},"tokens_in":766,"tokens_out":2826,"duration_ms":19362,"temperature":1.0,"reasoning_tokens":2732,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T10:29:29.384494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the identical pipeline on the 21 surviving outliers using O4b data with the C01 calibration and a lower false-alarm threshold; recovery of any of them at $p_{\\rm corr}^{\\rm fa}<5\\%$ would contradict the claim that all outliers are noise. A second check is to inject a simulated continuous wave at the quoted 95% sensitivity strain amplitude into one of the 562 O4a frequency-sky pixels and verify recovery; failure to recover it would falsify the sensitivity claim.","supporting_citations":[{"cited_title":"Small glitches and other rotational irregularities of the Vela pulsar","cited_arxiv_id":"2007.02921","evidence_quote":"Identifies the 562 O4a sub-threshold ASAF candidates that this paper follows up."}],"review_version":1}