{"id":"3e5c2d02-0ab8-49c9-af6f-2b9528a849dc","arxiv_id":"2412.10501","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An existing semi-coherent hierarchical search recovers the five loudest injected stellar-origin binary black holes in LISA Data Challenge 1b Yorsh, with SNR as low as 12.94.","lead":"This paper applies a previously developed GPU semi-coherent search to LISA Data Challenge 1b Yorsh and recovers five injected stellar-origin black hole binaries with SNR between about 12 and 25. It reports an effective detection threshold around SNR 12, lower than earlier estimates, on idealized challenge data with simplified LISA response models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The confidence threshold for 'confident detection' is imported from a different search (Fig. 5 of Ref. [17]) and applied to all Yorsh tiles without validation on Yorsh-like noise, so the claim of detecting an SNR 12.94 source is not independently calibrated.","rationale":"The paper is a useful and honest benchmark application of a previously published search to a public LDC. The recovery results in Tables II and III are internally consistent and partly corroborated by parameter estimation posteriors near injected values. However, the load-bearing step is significance assignment: the Υ ≥ 100 threshold is imported from Fig. 5 of Ref. [17] and applied to all tiles with no evidence that the background distribution transfers to the different Yorsh noise realization, PSD, TDI response, and search settings. The word 'confident' in the strongest claim is therefore uncalibrated. The paper explicitly states this limitation in Sec. III, which is to its credit, but the limitation is still decisive for the threshold-SNR claim because source #5 at SNR 12.94 is the lone support for extending the detection threshold below about 15. If the Yorsh background were recomputed and Υ ≥ 100 remained a valid threshold, the central claim would be strongly supported; if not, the claim would need to be downgraded to a recovery benchmark rather than a confident detection. I agree with the reader's weakest-assumption analysis and with the CONDITIONAL verdict: the recovery results are plausible and valuable, but the broader threshold claim needs Yorsh-specific background validation before it can be considered established.","tokens_in":10040,"tokens_out":1643,"duration_ms":13687,"concrete_test":"Run the SC search-1.5 pipeline on N ≥ 20 no-signal Yorsh-like datasets (same duration, cadence, PSD, TDI channels, and tile priors as the five detected sources) and record the maximum Υ_{N=1} per tile. If the 99% quantile of this background exceeds 100, the confident-detection threshold is too low and the claims about source #5 (SNR 12.94) and the threshold SNR are not established; if the quantile is well below 100, the imported threshold is conservative and the central claim survives. Run the test for tile #5 specifically, since the SNR 12.94 claim rests on that tile's background.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that SC search-1.5 confidently detects all five Yorsh sources with injected SNR ≥ 12.94. Significance is assigned by comparing the semi-coherent statistic Υ_{N=1} to a threshold Υ ≥ 100, where that threshold is justified by a background noise distribution 'from Fig. 5 in Ref. [17]' applied unchanged to all tiles (Sec. III). Ref. [17] used a different search, waveform, and synthetic data; Yorsh has a different PSD, TDI-2 response, and signal content. The false-alarm rate corresponding to Υ ≥ 100 on Yorsh-like noise is therefore unknown. If the Yorsh background is heavier-tailed, the five 'confident' detections could be noise fluctuations, especially the marginal SNR 12.94 source #5; if lighter-tailed, the claimed threshold SNR is unvalidated rather than measured. The paper is transparent about importing the background, but the headline interpretation rests on it. A secondary concern is hand-placed tiles around known injections, which weakens the search as a blind benchmark; the imported-background issue is more load-bearing because it directly underpins the 'confident detection' designation and the threshold-SNR conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the authors' semi-coherent hierarchical search (SC search) to the LISA Data Challenge 1b Yorsh dataset, using two approximate LISA response models. The more accurate model, SC search-1.5, returns candidate detections for the five injected SoBBH sources with SNR ≥ 12.94, recovering chirp mass to within ~0.002 M_sun and merger time to within hours; rapid MCMC parameter estimation gives posteriors consistent with the injections. The paper interprets this as evidence that the threshold SNR for detecting SoBBHs in LISA data is lower than some previous estimates. The main caveats are that detection significance is assigned using a background distribution imported from Ref. [17] rather than a Yorsh-specific noise background, and that search tiles were hand-placed around known injections.","tokens_in":10220,"tokens_out":4116,"duration_ms":40109,"significance":"If the results are taken at face value, they constitute a useful validation of a hierarchical semi-coherent search on a community LDC dataset, with public code and data products (Refs. [30, 46]) that aid reproducibility. The recovery accuracy reported in Tables II and III is internally consistent and the rapid parameter estimation is a practical contribution. However, the headline threshold-SNR conclusion is not fully supported: the false-alarm calibration is imported from a different search on different synthetic data, and the hand-placed tiles mean the exercise is not a blind search. These caveats temper the significance but do not destroy the value of the recovery results.","major_comments":[{"comment":"The threshold used to classify detections as confident is imported from Fig. 5 of Ref. [17] and applied unchanged to all Yorsh tiles, with no noise-only background computed for the Yorsh PSD, TDI-2 injections, or the SC search-1.5 response. Since the false-alarm rate of the Υ statistic is a property of the search and the data, the statement that source #5 at injected SNR 12.94 is 'confidently detected' is not calibrated, and the threshold-SNR conclusion in Sec. VI rests on this unvalidated input. The authors should either compute Yorsh-specific background distributions for at least representative tiles or explicitly weaken the significance language and the threshold-SNR claim.","section":"III"},{"comment":"The search tiles were chosen by hand around each individual injection, with the flow width set so that the prior on the derived parameter tc brackets the actual merger time (Fig. 2). This makes the search non-blind: the reported detection efficiencies and the threshold-SNR interpretation apply only to a search already directed to the correct region of parameter space. The text notes that Yorsh is not a blind challenge, but the abstract and Sec. VI do not carry this qualifier; this limitation should be stated explicitly wherever the headline results are summarized.","section":"III"}],"minor_comments":[{"comment":"The phrase 'all five sources in the data challenge with injected signal-to-noise ratios ≳ 12' is ambiguous because Yorsh contains eight SoBBH injections, three of which are not found by either search; the paper should say 'the five loudest injections' or 'all sources with injected SNR ≥ 12.94'.","section":"Abstract and Sec. IV"},{"comment":"The table layout contains stray punctuation and spacing, for example 'δMc, [M⊙]' in the SC search-1.5 header and entries such as '40689 .' and '11.60'; these should be cleaned for readability.","section":"Table II"},{"comment":"The semi-coherent statistic Υ_{N=1} is not defined in this paper; a one-sentence definition or an explicit equation reference to Ref. [17] would make the methods section more self-contained.","section":"III"},{"comment":"Reporting the maximum Υ value for every candidate, including the non-detections, would let readers see how far each candidate is from the adopted threshold; currently only the found/not-found status and matched-filter SNR are given.","section":"IV, Table II"},{"comment":"Appendix A correctly states that no detailed convergence checks were performed for the rapid parameter estimation, but the main text says the parameter estimation 'confirms' the detections; 'is consistent with' would be more proportionate given the absence of convergence checks.","section":"Appendix A and Sec. IV"}],"recommendation":"major_revision","confidential_remarks":"The recovery results appear internally sound, and I do not see grounds for rejection. The decisive issue is the transfer of the false-alarm threshold from Ref. [17] to Yorsh: if the authors cannot produce a Yorsh-specific background, they should revise the significance claims rather than present the threshold as established. The heavy reliance on the authors' own prior work is not inappropriate in itself, but the threshold transfer needs explicit validation or a clear downgrade of the claim. Fit to the journal is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, honest benchmark paper, and the central recovery result is very likely right. The significance threshold is imported from the authors' earlier search rather than calibrated on Yorsh-like noise, so the 'SNR ~12.94 detection' is not as statistically clean as the prose implies. But the PE follow-up, which lands on the injected parameters for all five sources, carries the main claim.\n\nThe genuinely new piece is the first SoBBH search run on an official LDC dataset, with waveform and response mismatch between injection and search. That is worth having. It shows the semi-coherent hierarchy with the approximate rigid-rotation TDI-1.5 response and Yorsh PSD finds five sources with injected SNR >= 12.94, recovers chirp masses to ~0.001 Msun and merger times to hours. The waveform extension with aligned spins is a concrete advance, checked in the circular limit. The paper is transparent: it states Yorsh is not blind, says tiles were hand-placed, and gives computational costs. Code and data products are public. That is proper work.\n\nThe main soft spot is the one the stress test flags. Sources with U >= 100 are called confidently detected because of a background distribution from Fig. 5 of Ref. [17], which used a different waveform, response, PSD, and noise realization. No Yorsh-specific noise background is computed, so the false alarm rate attached to 'confident detection' is unknown. I do not think this sinks the paper: the MCMC posteriors concentrate on injected values for all five detections, which is not what noise triggers should do. But it does mean the 'threshold SNR around 12' claim is not a measured, calibrated threshold; it is a suggestion based on a transfer assumption.\n\nA secondary point: hand-placed tiles around known injections mean the search would not work as a blind benchmark. The paper owns this, but anyone citing the SNR threshold should be careful. The self-citation pattern is not itself a problem—this is their pipeline, and the code and prior papers are public. The imported threshold is the issue, not the self-citation.\n\nThis paper is for LISA data analysis practitioners and anyone estimating SoBBH detection rates. It deserves a serious referee. I would send it to review and ask the authors to either compute a Yorsh-specific noise background or explicitly reframe the significance as tentative. I would cite it as a benchmark.","headline":"A useful, honest first benchmark of SoBBH searching on an official LISA Data Challenge, but the 'confident detection' label leans on an unvalidated background from an earlier paper.","tokens_in":10819,"tokens_out":2204,"would_cite":true,"duration_ms":22098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A semi-coherent search recovers all five stellar-origin black hole binaries in LISA Data Challenge 1b with injected SNR at least 12.94, indicating a lower detection threshold.","keywords":["LISA Data Challenge","stellar-origin binary black holes","semi-coherent search","gravitational wave data analysis","time-delay interferometry","particle swarm optimization","parameter estimation","Yorsh"],"falsifier":"Generate many Yorsh-like datasets containing only the same instrumental noise and no injected signals, run the identical search (SC search-1.5 with the same tiles and threshold), and count how often a pure-noise candidate reaches $\\Upsilon \\geq 100$; if the false-alarm rate is materially higher than the imported background predicts, the confident-detection labels are not calibrated for Yorsh.","tokens_in":9762,"feed_emoji":"🔭","tokens_out":13000,"duration_ms":105301,"temperature":0.7,"pith_summary":"This paper reports the first search for stellar-origin binary black holes inside a multipurpose LISA Data Challenge, specifically the 1b Yorsh dataset. The authors apply their GPU-accelerated hierarchical semi-coherent search, using templates from an eccentric post-Newtonian waveform and a simplified model of the LISA time-delay interferometry response, to hunt for the injected signals. The better of the two searches, SC search-1.5, confidently identifies all five injected sources with SNR at least 12.94 and recovers their chirp masses to within about $0.002\\,M_\\odot$ and merger times to within hours, with the two sources that merge during the mission recovered to within minutes. The authors read this as evidence that the SNR threshold for detecting stellar-origin binary black holes in LISA data is lower than some earlier estimates, which matters because the expected number of detected sources depends strongly on that threshold.","feed_headline":"Semi-coherent search finds all five LISA black-hole binaries","feed_subtitle":"The five recovered signals have injected signal-to-noise ratios of 12.94 or higher, pointing to a lower LISA detection threshold.","key_machinery":"The load-bearing object is the semi-coherent detection statistic $\\Upsilon$: a matched-filter score built by splitting the data into segments, comparing eccentric post-Newtonian template waveforms (TaylorF2Ecc, with aligned-spin terms up to 2.5 PN order) against the A/E/T time-delay-interferometry channels, and maximizing over parameters. The search uses a particle-swarm optimizer to explore each narrow tile in chirp-mass and starting-frequency space, and a fast frequency-domain model of the LISA response (TDI-1.5, a rigid rotating-constellation version of time-delay interferometry) rather than the full TDI-2 response used in the injections. Candidates with $\\Upsilon \\geq 100$ are declared confidently detected based on a noise background imported from an earlier search, then followed up with a fast ensemble MCMC for parameter estimation.","core_discovery":"The central claim is that a semi-coherent matched-filter search can find the five loudest stellar-origin binary black holes hidden in the Yorsh LISA Data Challenge 1b dataset, even though the search's waveform and instrument-response models differ from those used to generate the data. In the best-performing version (SC search-1.5, using the TDI-1.5 response and the Yorsh noise power spectral density), all five sources with injected SNR of 12.94 or higher are confidently detected; chirp masses are recovered within about $0.002\\,M_\\odot$ and merger times within hours, with the two sources that merge during the mission recovered to within minutes. The paper interprets these detections as evidence that the minimum SNR needed to detect stellar-origin binary black holes in LISA is lower than previous estimates. Rapid MCMC parameter estimation confirms that the detections correspond to the injected sources, with the expected small biases from waveform and response modeling.","pith_inferences":["Because the search used one hand-picked tile per injected source rather than covering the whole parameter space automatically, the demonstrated sensitivity applies to sources whose approximate location in chirp-mass and frequency space is already known; a fully blind search still needs automated tiling and per-tile background calibration.","The success with mismatched models suggests the search is robust to waveform and response inaccuracies for the loudest sources, but it also means the reported parameter biases, such as underestimated distances and slightly nonzero eccentricities, are partly systematic offsets from model mismatch rather than purely statistical errors.","If the lower threshold holds in realistic data with gaps, noise uncertainties, and confusion from other source classes, previous forecasts of LISA's stellar-origin binary black hole detection yield may need to be revised upward; this is directly testable in future LISA Data Challenges that include these sources and more realistic noise.","A natural next step would be to re-run the same pipeline on the same Yorsh data with the full TDI-2 response to separate the response-model contribution to parameter biases from waveform-model effects."],"forward_implications":["If the threshold SNR for detecting stellar-origin binary black holes in LISA is around 12 rather than higher, the expected number of such sources in real LISA data increases, since source counts depend steeply on this threshold.","Simplified LISA response models (TDI-1 and TDI-1.5) are sufficient for detecting some sources, especially at lower frequencies, so parts of a global fit may be able to use cheaper response models without losing the sources.","The automated rapid parameter estimation after each detection delivers posteriors narrow enough to initialize more detailed global-fit parameter estimation.","For the two sources that merge within the LISA mission lifetime, merger time is recovered to within minutes, enabling multi-band follow-up planning.","The per-tile computational cost of about one to three days on a GPU makes a full survey of roughly 100 to 1000 search tiles over the stellar-origin binary black hole parameter space feasible."],"supporting_citations":[{"why":"Describes the GPU-accelerated hierarchical semi-coherent search algorithm that this paper applies to Yorsh.","marker":"[16]"},{"why":"Defines the semi-coherent statistic, the search procedure, and supplies the noise background distribution used to set the confidence threshold.","marker":"[17]"},{"why":"The LISA Data Challenge 1b Yorsh dataset analyzed here.","marker":"[18]"},{"why":"Supplies the IMRPhenomD waveform used to generate the injected signals, against which the search's recovery is compared.","marker":"[27]"},{"why":"Supplies the TaylorF2Ecc post-Newtonian waveform model used as the search template.","marker":"[31]"},{"why":"Provides the frequency-domain LISA response model used to compute the TDI channels in the search.","marker":"[34]"},{"why":"Provides the implementation of the fast frequency-domain LISA response model used by both searches.","marker":"[36]"},{"why":"Gives a previous higher estimate of the threshold SNR for detecting these sources, which the present results are compared against.","marker":"[39]"}],"fun_headline_variants":["Semi-coherent LISA search finds all five stellar-origin binaries","LISA challenge search detects all five black-hole binaries","LISA search finds five, lowering estimated detection threshold","All five binaries recovered in LISA challenge search"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The significance threshold that decides which candidates count as confidently detected was calibrated on a background of noise triggers measured in a different, earlier search on different synthetic data, and that same threshold is applied to every search tile here without re-measuring the background for Yorsh.","fun_headline_variants_meta":{"raw":{"variants":["Semi-coherent LISA search finds all five stellar-origin binaries","LISA challenge search detects all five black-hole binaries","LISA search finds five, lowering estimated detection threshold","All five binaries recovered in LISA challenge search"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00108,"raw_usage":{"total_tokens":4491,"prompt_tokens":891,"completion_tokens":3600,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":3534}},"tokens_in":507,"tokens_out":3600,"duration_ms":23291,"temperature":1.0,"reasoning_tokens":3534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:53:12.255438+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate many Yorsh-like datasets containing only the same instrumental noise and no injected signals, run the identical search (SC search-1.5 with the same tiles and threshold), and count how often a pure-noise candidate reaches $\\Upsilon \\geq 100$; if the false-alarm rate is materially higher than the imported background predicts, the confident-detection labels are not calibrated for Yorsh.","supporting_citations":[{"cited_title":"Bandopadhyay and C","cited_arxiv_id":null,"evidence_quote":"Defines the semi-coherent statistic, the search procedure, and supplies the noise background distribution used to set the confidence threshold."},{"cited_title":"LISA Data Challenges,","cited_arxiv_id":null,"evidence_quote":"The LISA Data Challenge 1b Yorsh dataset analyzed here."}],"review_version":1}