{"id":"128f7c7b-380f-46cf-88ab-2ea259f852bb","arxiv_id":"2411.18370","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A generic GPU search engine plus a fast semianalytic sensitivity estimator could make blind continuous gravitational-wave searches much cheaper to run and characterize.","lead":"The paper introduces fasttracks, a GPU-accelerated engine that evaluates detection statistics for generic continuous gravitational-wave searches, and reports large speedups without physical approximations. It also presents a fast semianalytic sensitivity-estimation method and tests both in a minimal all-sky search using LIGO-Virgo-KAGRA O3 data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Semianalytic sensitivity estimator is validated on only one of eight frequency bands; its transfer to other bands is assumed, not demonstrated.","rationale":"The paper is an honest methodological contribution: the GPU brute-force engine is simple and the semianalytic sensitivity estimator extends prior work, with a real-data demonstration. In stress-testing the strongest claim, I focused on the statistical-equivalence assertion because it is the 'one-stop' part of the paper: it promises a negligible-cost sensitivity estimate without large injection campaigns. The reader identified the transfer of the Gaussian-noise model to real O3 data as the weakest assumption; I agree. I considered two alternative concerns. (1) The GPU-efficiency comparison to model-specific pipelines is not same-baseline, but this is explicitly framed as an order-of-magnitude statement and does not underpin the sensitivity-estimate novelty. (2) The mismatch distribution p(m|o) is defined via actual noisy s values (Eq. 24), so it might depend on signal amplitude; however, the Fig. 4 validation spans a broad range of D and matches, which suggests the amplitude dependence is not dominant for that band. The remaining and most consequential risk is band dependence of the non-Gaussianity veto. If the number-count threshold does not restore Gaussian statistics in other bands, the semianalytic pdet is biased, and the paper's central claim of a generic estimator fails. A single additional validation band would substantially de-risk the claim; the seven other bands in Fig. 6 are already injection-tested, so the comparison is cheap. I therefore maintain the reader's CONDITIONAL verdict: the claim is plausible and well-supported on one band, but acceptance as a fully generic method should await the additional check.","tokens_in":25972,"tokens_out":8262,"duration_ms":72908,"concrete_test":"Run the cows3 semianalytic estimator for at least one additional band in Fig. 6 (e.g., 260.625 Hz) using the same search setup, the same oversampling o=3.44, and the actual loudest-candidate list T from the real-data search after applying the weighted number-count threshold and the optimal Nreject/Ncand selection. Compare the predicted pdet(D) curve (or D95%) with the injection-measured curve for that band. If the predicted D95% deviates from the measured value by more than the combined statistical uncertainty (sigmoid-fit covariance plus Monte Carlo sampling error, approximately 0.2-0.3 in D), the transfer of the Gaussian model outside the 110.5 Hz band is unsupported. Running the same check for all seven remaining bands would fully settle the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the semianalytic sensitivity estimator (Sec. V) is statistically equivalent to injection campaigns for arbitrary short-coherence searches rests on validating the Gaussian-noise model p(z|D,o,lambda) against real-data injections after the number-count veto and box rejection. This validation is performed only for the 110.5 Hz band (Fig. 4). The estimator uses Eqs. (21)-(23) and (46), which assume Gaussian noise statistics with a fixed 1.012 PSD-bias factor, and a step detection rule (Eq. 47) that does not itself model the weighted number-count veto introduced in Sec. VIB. The agreement in Fig. 4 may reflect that on this particular band the veto sufficiently suppresses non-Gaussianities so that surviving tracks behave Gaussianly. The paper provides no independent comparison of the semianalytic prediction with the injection-measured detection probabilities for the other seven bands in Fig. 6; it only compares those measured values with a previous search [130]. If the veto is less effective on other bands, the semianalytic pdet would be biased, undermining the generic 'one-stop' sensitivity-estimate claim. The paper itself notes in Sec. VIB that real-data non-Gaussianities make the Sec. V results inapplicable unless the veto restores Gaussian behavior, which is precisely the unvalidated link.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a 'one-stop' strategy for blind searches of long-duration gravitational-wave signals. It introduces fasttracks, a GPU-accelerated brute-force engine for short-coherence detection statistics; a random template bank built from uniform-density coordinates; an ahead-of-time box partition that replaces clustering-based post-processing; and a semianalytic sensitivity-estimation method that generalizes earlier work by Wette and Dreissigacker et al. to weighted short-coherence statistics and box-based selection rules. The authors demonstrate the approach on an all-sky search for binary neutron-star continuous-wave signals in eight frequency bands from the first half of O3, validate the sensitivity estimator against injections at 110.5 Hz, and compare the resulting depth estimates with a previous search.","tokens_in":26255,"tokens_out":8177,"duration_ms":72868,"significance":"If the central claims hold, the paper offers a substantial practical advance: a generic, signal-model-agnostic search pipeline whose sensitivity can be estimated without large injection campaigns, complementing the existing literature on model-specific GPU implementations and analytic sensitivity estimation. The open-source releases of fasttracks and cows3, the explicit Monte Carlo formulation of pdet in Eq. (48), and the real-data injection tests are concrete strengths. The estimator is not circular: p(q|λ) and p(m|o) are constructed from independent simulated priors and then compared with an injection campaign. However, the validation is currently too narrow to support the full 'generic one-stop' claim, and the GPU efficiency comparison lacks a same-hardware baseline against existing pipelines.","major_comments":[{"comment":"The semianalytic sensitivity estimator is validated against injection-measured detection probabilities only for the 110.5 Hz band (Fig. 4). For the seven other bands, Fig. 6 compares only injection-measured D95 values with a previous search [130]; no semianalytic pdet curves or D95 predictions are shown for those bands. Since Sec. VIB states that real-data non-Gaussianities make the Sec. V results inapplicable unless the number-count veto restores Gaussian behavior, the transfer of p(m|o) and the Gaussian-noise p(z|D,o,λ) model to the other bands is exactly the unverified link. Moreover, Eq. (47) models the detection rule as a step threshold on z and does not include the weighted number-count condition used in the actual search, so the agreement at 110.5 Hz may reflect properties of that particular band rather than a general equivalence between the sampler and an injection campaign. To support the paper's central claim, the semianalytic D95 should be compared with the injection-measured values for all eight bands, or at least for a random subset not used to tune the veto.","section":"Sec. V, Sec. VIB, Figs. 4 and 6"},{"comment":"The claim that GPU-accelerated brute-force template evaluation provides 'comparable computing efficiencies to using model-specific optimizations' is not directly established by the measurements shown. Fig. 2 benchmarks fasttracks on an H100 against a CPU only; the comparison with [53,94] is qualitative and uses different hardware, workloads, and SFT configurations. The fitted cost model in Eq. (26) is stated without a derivation or a test against the data points, and the figure has no error bars despite being averaged over 10 realizations. I recommend either adding a same-hardware benchmark of at least one existing pipeline with the same SFT data and template bank, or explicitly restricting the claim to the CPU-versus-GPU speedup of fasttracks.","section":"Sec. III, Fig. 2, Eq. (26)"},{"comment":"The treatment of the PSD-estimation bias is internally inconsistent. Sec. IIB states that estimating Sn with a running median introduces a 'small upward bias in µG and σG', but Eq. (46) applies the 1.012 factor only to µG, not to σG. If the bias is common to both moments, the standardized statistic z in Eq. (36) and hence the threshold rule in Eq. (47) are miscalibrated; if the bias affects only the mean, the text should say so and the numerical value should be justified. The current statement 'about 1.2% for µG' is not derived or evidenced in the paper, despite being a multiplicative calibration of the central sensitivity estimate.","section":"Sec. IIB, Eq. (46)"}],"minor_comments":[{"comment":"There are several typos: 'NVDIA' should be 'NVIDIA' in Sec. III, 'Tukey widow' should be 'Tukey window' in Sec. VI, 'unaplicable' should be 'inapplicable' in Sec. VIB, and 'resoltuion' should be 'resolution' in Appendix A.","section":"Secs. III, VI, Appendix A"},{"comment":"The injection-based pdet points are shown without binomial error bars, although each point is based on 500 injections; adding confidence intervals would make the agreement with the semianalytic curves more interpretable.","section":"Figs. 4 and 5"},{"comment":"The statement that p(m|o) is 'compatible with Weibull distributions' is not supported by any quantitative fit; the authors should either provide fit parameters and a goodness-of-fit measure or soften the statement to a qualitative observation.","section":"Sec. IVA, Fig. 3"},{"comment":"The cost-model formula in Eq. (26) is ambiguous because the denominator is not clearly parenthesized; the intended expression for log10 cλ should be written out unambiguously.","section":"Eq. (26)"},{"comment":"The statement that the obtained sensitivity depths are 'broadly consistent' with [130] is not quantified, and the blue crosses are described as being shown only for completeness; either remove the consistency claim or support it with a numerical comparison.","section":"Sec. VIC, Fig. 6"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the core ideas are promising. The main risk is that the headline claims outrun the evidence: the semianalytic sensitivity estimator is validated on a single frequency band, and the GPU efficiency comparison lacks a same-hardware baseline against model-specific pipelines. These are fixable with additional validation or by explicitly narrowing the claims. No concerns about citation practices or novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line up front: this is a solid methods paper for blind continuous-wave searches, and the core claims are honestly scoped. What's actually new is the combination of a generic brute-force GPU engine (fasttracks) that vectorizes any short-coherence detection statistic, an ahead-of-time box partitioning in density coordinates that replaces clustering, and a generalized semianalytic sensitivity estimator (cows3) that samples p(z|D,o,lambda) by Monte Carlo integration of the analytic Gaussian model. The individual ingredients are known, but the packaging is genuinely useful and the open-source releases make it reproducible.\n\nThe derivations in Sec. II are standard and internally consistent. The sensitivity estimator is not circular: mismatch and orientation priors are simulated independently, and the validation against injections is genuine. The real-data demonstration on O3 data is a plus.\n\nThe soft spots are real but not fatal. First, the semianalytic pdet is validated against injections only at 110.5 Hz (Fig. 4). For the other seven bands, Fig. 6 compares measured injection-based D95% with the previous search [130], not with the semianalytic prediction. The paper itself notes that non-Gaussianities make the Sec. V results inapplicable unless the weighted number-count veto restores Gaussian behavior. That link is exactly the unvalidated transfer. It could hold everywhere, but the evidence only supports it for one band. That's a moderate gap, not a fatal one.\n\nSecond, the GPU benchmark (Fig. 2) compares fasttracks on an H100 with CPU runs of the same code, and then claims parity with other GPU pipelines [53,94] that ran on different hardware. Without a same-baseline comparison, the 'comparable efficiency' claim is weaker than the text suggests. This is a minor-to-moderate issue.\n\nMinor points: no commit hash or direct reproduction scripts are cited, just GitHub URLs, and the cost model in Eq. (26) has fit coefficients without quoted uncertainty. Neither changes the thrust.\n\nWho this is for: anyone building blind or semi-directed long-duration searches, especially for next-generation detectors, and anyone who needs quick sensitivity estimates without large injection campaigns. It deserves a serious referee. I would send it to review with a request for either more band-level validation of the estimator or a clear statement that only the one band is validated, plus a same-hardware baseline if feasible.","headline":"A credible, useful methods paper: generic GPU engine plus a semianalytic sensitivity estimator demonstrated on one band; the main gap is the unvalidated transfer of that estimator to other bands.","tokens_in":26722,"tokens_out":2236,"would_cite":true,"duration_ms":21189,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single GPU engine can now search for continuous gravitational waves generically, matching specialized pipelines while estimating sensitivity without large injection campaigns.","keywords":["continuous gravitational waves","GPU computing","semicoherent search","sensitivity estimation","random template banks","binary neutron stars","short-coherence detection statistics","number count veto"],"falsifier":"Run the identical search configuration with 500 injections in one of the other seven frequency bands, such as 187.0 Hz, and compare the measured detection probability against the semianalytic curve; a deviation in the 95% sensitivity depth larger than the fit's quoted uncertainty would falsify the claim that the Gaussian model transfers once the number-count veto is applied.","tokens_in":25785,"feed_emoji":"📡","tokens_out":5915,"duration_ms":52168,"temperature":0.7,"pith_summary":"This paper claims that a brute-force, GPU-parallelized evaluation of short-coherence detection statistics can match the computational efficiency of purpose-built continuous-wave search pipelines without any model-specific optimization, and that the sensitivity of such a generic search can be estimated semianalytically by sampling from a four-parameter Gaussian distribution rather than running injection campaigns. The engine evaluates any statistic that is a weighted sum of normalized-power values along a time–frequency track, so it applies to isolated neutron stars, binary systems, and other long-duration signals alike. The sensitivity estimator replaces the costly step of simulating and re-analyzing thousands of signals with Monte Carlo draws that are statistically equivalent to those injections, including the effect of random template banks and ahead-of-time post-processing boxes. The scheme is demonstrated in an all-sky search for unknown binary neutron stars in real data from the third observing run, where a minimal number-count veto brings the measured detection probability into agreement with the semianalytic estimate.","feed_headline":"GPU brute force matches specialized CW searches","feed_subtitle":"A generic engine evaluates any short-coherence statistic and estimates sensitivity without mass injection campaigns.","key_machinery":"The load-bearing object is the short-coherence detection statistic s(λ) = Σ w_{Xα} s[t_{Xα}; f(t_{Xα}; λ)], evaluated with bulk vectorized array operations; its Gaussian approximation in the many-SFT limit reduces the signal distribution to four moments q = {μG, σG, ρ̂1², ρ̂2²}, with μG carrying a 1.012 PSD-estimation bias factor. Around this, the paper builds a uniform-density coordinate system ξ(λ) from the local template density ϱ(λ), which turns a non-uniformly populated parameter space into a hyperbox where random template banks with oversampling o are trivial to draw and where mismatch distributions p(m|o) are precomputed once in Gaussian noise. Post-processing is replaced by an ahead-of-time partition of ξ-space into fixed-size hyperboxes; the detection rule is then a per-box threshold τ(λ) on the standardized statistic z, which makes the detection probability a cheap Monte Carlo integral over q, m, and λ. A minimal weighted number count, computed from a binarized spectrogram with threshold 3.2, acts as the persistence veto that makes the Gaussian model applicable to real data.","core_discovery":"The central claim is twofold. First, for a short-coherence detection statistic that combines SFT power along a template track, batch evaluation on a GPU makes a brute-force template loop two to three orders of magnitude faster than a saturated multi-core CPU implementation and comparable to existing GPU-accelerated pipelines that exploit parameter-space structure; the per-template cost is captured by an empirical model in Eq. (26). Second, the detection-statistic distribution under the signal hypothesis is determined, in the large-NSFT Gaussian limit, by only four quantities q = {μG, σG, ρ̂1², ρ̂2²} built from detector weights and per-SFT signal power, so drawing a detection statistic for a signal of depth D reduces to sampling a mismatch m from the template-bank oversampling distribution p(m|o), a sky and orientation population, and a Gaussian with those moments. The paper states this equivalence directly: sampling p(z|D, o, λ) is statistically equivalent to injecting a CW signal in Gaussian noise at depth D and retrieving the maximum detection statistic z from a random template bank with oversampling o. In real data from the third observing run, imposing a minimal weighted number count suppresses non-Gaussian artifacts well enough that the measured detection probabilities agree with the semianalytic estimate, and the resulting 95% sensitivity depths are consistent with those of a prior all-sky binary search.","pith_inferences":["The equivalence between sampling p(z|D, o, λ) and injection campaigns suggests that, in regimes where the Gaussian-noise model is trusted, injection campaigns could be reduced to validation sets rather than used as the primary sensitivity estimator, a step the authors validate only in one band.","The ξ-coordinate random template bank sidesteps metric-based lattice placement; if template counting rather than metric coverage is the binding constraint, the oversampling o needed for a given mismatch may itself become the search's main tuning knob, something the paper does not directly quantify in terms of missed volume.","Because the whole pipeline is expressed as array operations, the same engine could be paired with automatic differentiation of the frequency model to build template banks or to optimize follow-up, an extension the paper does not pursue.","The other seven frequency bands' agreement with the semianalytic estimate is implied by the Gaussian transfer but not independently demonstrated, so the method's portability across bands is an extrapolation to check."],"forward_implications":["Any short-coherence search, regardless of the source model, can be deployed by supplying only a frequency-evolution model f(t; λ); no pipeline-specific optimization is needed to reach competitive GPU efficiency.","Sensitivity estimation for a blind search becomes a laptop-scale calculation, so optimal candidate-selection and box-rejection strategies can be solved in minutes instead of by dedicated injection campaigns.","Requiring a minimal weighted number count removes most non-Gaussian artifact contamination, restoring agreement between measured and modeled detection probability.","The per-template cost model in Eq. (26) lets future searches choose template-batch sizes and SFT counts to saturate a given GPU, and predicts that further gains come from reducing the number of semicoherent segments.","The same machinery transfers to long-duration signals from next-generation ground-based and space-borne detectors, where compact-binary coalescence signals linger in band and resemble continuous waves."],"supporting_citations":[{"why":"Supply the original semi-analytic sensitivity-estimation methods that this work generalizes to short-coherence weighted statistics.","marker":"[55, 56]"},{"why":"Provide the follow-up framework and the sensitivity-simulator implementation that the new estimator builds on.","marker":"[57, 58]"},{"why":"Gives the number-count statistic, the running-median PSD estimate, and the short-coherence setup used in the search.","marker":"[34]"},{"why":"Is the GPU-accelerated binary-system search pipeline whose computational efficiency is compared against.","marker":"[53]"},{"why":"Provides the improved all-sky binary search method and the short-segment cost scaling discussed in the paper.","marker":"[89]"},{"why":"Supplies the real-data search setup and sensitivity-depth values the results are compared with.","marker":"[130]"},{"why":"Establish the random-template-bank framework and expected mismatch behavior used for the oversampling prescription.","marker":"[118, 121]"},{"why":"Supplies the parameter-space metric and resolution for binary-system CW searches that Eq. (28) agrees with.","marker":"[67]"}],"fun_headline_variants":["GPU brute force makes CW searches generic and fast","Fasttracks: semi-analytic sensitivity for blind CW searches","GPU computing speeds up all-sky continuous-wave searches","Generic CW detection without mass injection campaigns","Brute-force on GPU matches specialized CW pipelines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The semianalytic sensitivity estimator assumes that the Gaussian-noise distributional model, including the 1.012 PSD-bias factor and the precomputed mismatch distribution p(m|o), remains valid in real detector data once a minimal number-count veto is imposed; this transfer is checked against injections only in the 110.5 Hz band, not in the other seven bands.","fun_headline_variants_meta":{"raw":{"variants":["GPU brute force makes CW searches generic and fast","Fasttracks: semi-analytic sensitivity for blind CW searches","GPU computing speeds up all-sky continuous-wave searches","Generic CW detection without mass injection campaigns","Brute-force on GPU matches specialized CW pipelines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3295,"prompt_tokens":1005,"completion_tokens":2290,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":2217}},"tokens_in":621,"tokens_out":2290,"duration_ms":15376,"temperature":1.0,"reasoning_tokens":2217,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:16:21.220784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical search configuration with 500 injections in one of the other seven frequency bands, such as 187.0 Hz, and compare the measured detection probability against the semianalytic curve; a deviation in the 95% sensitivity depth larger than the fit's quoted uncertainty would falsify the claim that the Gaussian model transfers once the number-count veto is applied.","supporting_citations":[{"cited_title":"Measuring neutron-star distances and properties with gravitational-wave parallax","cited_arxiv_id":"2212.07506","evidence_quote":"Is the GPU-accelerated binary-system search pipeline whose computational efficiency is compared against."},{"cited_title":"Estimating the sensitivity of wide-parameter-space searches for gravitational-wave pulsars","cited_arxiv_id":"1111.5650","evidence_quote":"Supplies the parameter-space metric and resolution for binary-system CW searches that Eq. (28) agrees with."}],"review_version":1}