{"id":"e0aaa4eb-4d66-4fb7-a253-ee4d72771665","arxiv_id":"2607.18136","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"After native-background recalibration and empirical physical vetoes, DANTE V3 finds zero coincident novel glitch morphologies in LIGO O4a, with per-detector 90% upper limits of 5.63–5.83 yr-1.","lead":"An unsupervised deep-learning pipeline for LIGO glitch hunting, DANTE V3, finds that after re-scoring candidates against the detector's own background and applying physical vetoes, no discrete novel glitch family and no coincident candidate survive in early O4a data. It quotes 90% upper limits of about 5.6–5.8 per year on new uncatalogued instrumental morphologies and argues that previous unsupervised hits were mostly drifting-noise artifacts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rate upper limits use zero-event Poisson despite one surviving L1 outlier; L1 R90 should be ~9.5 yr^-1 not 5.63.","rationale":"The reader's weakest assumption was the pooled-cross-correlation null, but the rate-limit inconsistency is more decisive and more concrete: it is an internal contradiction between the paper's own classification of the surviving L1 singleton and the zero-event Poisson calculation. The central claim includes two quantitative limits; the L1 limit is computed as if no novel instrumental morphology survived, while the text explicitly states one did. Recomputing with N=1 changes the limit by ~70%, so the quoted number is wrong as stated. This does not necessarily invalidate the broader qualitative conclusion (low rate, no discrete family), which is why the verdict remains CONDITIONAL rather than REJECT, but the paper cannot be accepted without correcting this. The pooled-null concern is also valid but is partially defended by the per-event comparison (11.8% vs 12.3%) and the injection-recovery tests; the Poisson-counting error has no such defense. Hence I flag the rate limit as the single most load-bearing concern. The reader mentioned the inconsistency in the rationale but not as the weakest assumption, hence 'partial' agreement.","tokens_in":23630,"tokens_out":5546,"duration_ms":57740,"concrete_test":"Recompute the L1 Poisson upper limit using N=1 (the surviving uncatalogued outlier, GPS 1382955228.0) instead of N=0, with T=149.4 d: λ90=0.5*chi2(0.9, 4)=3.889, so R90=3.889/(149.4/365)≈9.5 yr−1. Audit the disposition ledger to confirm whether any other DSD-ROBUST survivor (e.g., the three Family A isolates) is classified as uncatalogued and should increment the count; if N≥1 on L1, the abstract's limits are understated and must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper quotes 90% Poisson upper limits on the emergence of novel uncatalogued instrumental morphologies as R90≤5.83 yr−1 (H1) and R90≤5.63 yr−1 (L1), computed from N_unexplained=0 via λ90=−ln(0.1)≈2.303 (Sec. VII B, Eq. 3). But Sec. VII A explicitly classifies the lone surviving L1 event (GPS 1382955228.0) as an 'uncatalogued instrumental morphological outlier' that passes the PEM veto and has no cross-detector counterpart. By the paper's own definition, this is exactly the population the limit is meant to bound: a novel, uncatalogued, per-detector instrumental morphology surviving the DSD and PEM vetoes. For L1, N=1, not 0. The correct 90% Poisson upper limit for one observed event (0.5*chi2_0.9(4)=3.889) gives R90≈3.889/(149.4/365)≈9.5 yr−1, roughly 70% higher than quoted. The abstract's 'no coincident event' is not equivalent to 'no novel morphology'; the zero-event calculation is internally inconsistent with the paper's own surviving outlier, overstating the constraint and undermining the headline rate claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DANTE V3, a fully retrospective unsupervised detector-characterization search for novel instrumental glitches in early LIGO O4a data. A multi-scale Q-transform/MIL pipeline extracts 10,372 candidates; a block-bootstrap Domain Shift Defense against a vector-quantized native background index leaves 2,937 robust survivors. The survivors form one macro-cluster; the authors show that this topology is shared by pristine background and explicitly retire an earlier morphological-diffusivity test as circular. An empirical family-wise PEM veto removes one singleton via a control-line coupling, while one L1 singleton survives as an uncatalogued instrumental outlier. A physical cross-correlation coincidence test, validated with 2,400 injections and a 200-trial null, finds no cross-detector coincidence among 8,749 candidates. The paper quotes 90% Poisson upper limits on the rate of novel instrumental morphologies of R90≤5.83 yr−1 (H1) and R90≤5.63 yr−1 (L1).","tokens_in":23952,"tokens_out":12914,"duration_ms":137819,"significance":"If the results hold, the main contributions are methodological: a self-critical validation of an unsupervised glitch search under domain shift, including a directly measured coincidence-statistic efficiency (2,400 injections, ten morphologies), a 200-trial null, an empirically calibrated PEM null, and a like-for-like falsification control showing that the single-macro-cluster topology is a property of the embedding geometry rather than of anomaly status. The paper is unusually transparent: it ships pinned open-source code and archived artifacts, discloses bugs and retires its own earlier statistic. However, the headline L1 rate limit is internally inconsistent with the paper's own surviving L1 outlier, and the pooled coincidence null needs stationarity support. The central negative result is defensible, but one principal quantitative claim is currently overstated.","major_comments":[{"comment":"The zero-event Poisson limit for L1 is inconsistent with Sec. VII A, where GPS 1382955228.0 survives as an uncatalogued instrumental morphological outlier — exactly the population Eq. (3) is said to bound ('novel, uncatalogued, per-detector instrumental morphologies that survive the DSD and PEM vetoes'). Table I's 'Final unexplained events: 0' contradicts this. With N=1, the 90% Poisson upper limit is λ90=0.5·χ²_0.9(4)=3.889, giving R90≈3.889/(149.4/365)≈9.5 yr⁻¹, not 5.63 yr⁻¹. Either count the L1 singleton or explicitly relabel the limits as coincident-only; as written, the abstract overstates the constraint by about 70%.","section":"Sec. VII B / Eq. (3) / Table I"},{"comment":"The 'no coincident event' conclusion depends on a pooled coincidence null. Because each event contributes only 4–8 usable time shifts, the threshold τ_cc=0.405 is the 99th percentile of the pooled null max-statistic distribution, not a per-event false-alarm threshold. This assumes exchangeability of the null across 8,749 events in a dataset the paper itself emphasizes is non-stationary. The authors should demonstrate that the null max-statistic distribution is stable across sessions or months (e.g., a per-session comparison, or a random-effects model). Without this, the 0.149% on-source exceedance is not a calibrated per-candidate false-alarm rate, and the global threshold could be miscalibrated in quiet versus noisy epochs. Limitation 11 notes the per-event significance issue but does not resolve it; this is load-bearing for the headline null result.","section":"Sec. IV B / Limitation 11"}],"minor_comments":[{"comment":"The text says the coincidence stage recovers coherent waveform injections 'with efficiency consistent with zero (ε_coh ≈0)', which directly contradicts Table IV, where all ten morphologies reach ε_coh=100% at sufficient SNR. Please rephrase to indicate that efficiency is strongly morphology- and SNR-dependent.","section":"Sec. VII B"},{"comment":"With 9,999 Mantel permutations, the smallest attainable p-value is 10^-4; reporting p<10^-4 is not possible. Please report the exact p-value (or use a larger permutation count).","section":"Sec. V"},{"comment":"The abstract says 'below both the 1% nominal false-alarm rate and the 1.01% realized by the null itself.' Since τ_cc is a pooled threshold rather than an event-specific false-alarm probability, please clarify in the abstract or text that this is a population-level statement, not a per-candidate significance.","section":"Sec. IV B / Abstract"},{"comment":"The tail counts 13 (on-source) vs 88 (null) are small; adding Poisson confidence intervals or a zoom panel would help the reader assess the statistical weight of the 'deficit' claim.","section":"Fig. 2 / Sec. IV B"},{"comment":"Please state explicitly whether the released code tag 3.5.0 uses the corrected frequency-band constant in the coincidence test and in Table VII, so that the reported frequency parameters and the coincidence null are reproducible from the archived artifacts.","section":"Sec. VIII, Limitation 16"}],"recommendation":"major_revision","confidential_remarks":"The paper is unusually transparent and the methodology is largely sound; the rate-limit error is simple to fix but appears in the abstract and cannot be treated as a minor issue. The pooled-null stationarity concern should be addressed with a concrete test. I do not recommend rejection: the central negative result is supported by the injection campaign and seems likely to survive revision. Please also ensure Table I's '0' is reconciled with the surviving L1 singleton, and that the corrected L1 limit is reported consistently in the abstract, main text, and conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Cirfeta's DANTE V3 is the rare null result that does its homework. The main negative finding — no discrete novel glitch family, no coincident cross-detector event in early O4a — is supported by measured recovery efficiency, a like-for-like background falsification control, and explicit retirement of the author's own earlier statistics when they turned out to be confounded. That alone makes it worth a referee's time. But the headline rate limits are internally inconsistent: Sec. VII A classifies the surviving L1 event (GPS 1382955228.0) as an uncatalogued instrumental outlier that passed the PEM veto, and Sec. VII B then quotes a zero-event Poisson limit for L1. If that event belongs to the population of 'novel uncatalogued instrumental morphologies surviving DSD and PEM vetoes' — and by the paper's own definition it does — then L1 has N=1, not N=0. The correct 90% Poisson upper bound for one observed event, 3.889/(149.4/365) ≈ 9.5 yr^-1, is about 70% higher than the quoted 5.63. The abstract's 'no coincident event' is not the same as 'no novel morphology', and the current wording conflates them.\n\nWhat is genuinely new and solid: multi-scale extraction with DINOv2/MIL (0.5–4 s), the block-bootstrap DSD, the empirically calibrated family-wise PEM null, and a physical cross-correlation coincidence test that is validated with 2,400 injections rather than asserted. The paper also demonstrates that the embedding-similarity coincidence test cannot separate signal from null, and it measures the recovery efficiency of the coincidence stage across ten morphologies with morphology-dependent saturation SNRs. That is the right way to state a null result. The falsification test showing pristine background is even more monolithic than the survivors is a clean, persuasive control.\n\nThe soft spots are mostly acknowledged. The pooled coincidence null uses only 4–8 time shifts per event, so per-event significance is not resolvable; the paper says so in Limitation 11 and thresholds on the pooled distribution instead. The PEM veto assumes the public auxiliary channels are safe, which cannot be verified from outside the collaboration (Limitation 7); that weakens the veto verdicts but not the main negative result. The rate-limit error is the one serious unacknowledged problem. It is fixable and should be fixed before publication.\n\nFor a reading group on detector characterization or unsupervised anomaly detection, this paper is valuable: it shows how to calibrate against domain shift with a native background and how to retire a statistic honestly. I would cite it for the methodological lessons. It deserves a serious peer review, with the rate-limit correction required.","headline":"A self-correcting null result that earns serious review, but the quoted rate limits silently drop the paper's own surviving L1 outlier.","tokens_in":24416,"tokens_out":4910,"would_cite":true,"duration_ms":45302,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Early O4a data, once recalibrated against native background and physically vetoed, show no discrete novel glitch family and no cross-detector coincident event; 90% upper limits are 5.83 yr−1 (H1) and 5.63 yr−1 (L1).","keywords":["gravitational-wave detector characterization","glitch morphology","unsupervised anomaly detection","domain shift","multi-scale time-frequency analysis","coincidence veto","rate upper limits","O4 observing run"],"falsifier":"For a random sample of candidates, build a dense per-event null using 100+ time shifts instead of the 4–8 available (for example, shifts of ±0.1, ±0.2, ..., ±8 s). If more than about 1% of on-source events exceed their per-event 99th percentile, the pooled-threshold design is masking event-level coincidences and the zero-coincidence claim fails. A second check: compare per-event null means between early and late O4a sessions; if they differ beyond sampling noise, the pooled null is not exchangeable.","tokens_in":23498,"feed_emoji":"📡","tokens_out":8185,"duration_ms":79889,"temperature":0.7,"pith_summary":"This paper tries to establish whether the early O4a gravitational-wave data contain any genuinely new, uncatalogued instrumental glitch morphologies once the detectors' drifting noise is properly accounted for. It builds a multi-scale unsupervised search that flags over 10,000 candidates, then re-scores them against a native background index using a block-bootstrap domain-shift defense; only 28.3% survive, and those survivors form one diffuse cluster with no compact substructure, the same topology as pristine background. After empirically calibrated environmental-channel vetoes and a physical cross-correlation coincidence test, no cross-detector coincident event remains: 0.149% of 8,749 candidates exceed the pooled null threshold, below the 1% false-alarm rate of the null itself. The paper therefore reports 90% upper limits of 5.83 yr−1 (H1) and 5.63 yr−1 (L1) on the rate of novel uncatalogued instrumental glitch emergence, as detector-characterization statements rather than astrophysical limits.","feed_headline":"Zero novel glitch families survive O4a vetting","feed_subtitle":"Recalibrating to native noise and applying physical vetoes leaves no coincident candidates; rate under 6/year.","key_machinery":"The argument is carried by the DANTE V3 pipeline. Its load-bearing parts: (1) multi-scale Q-transform spectrograms at 0.5–4 s plus a 32 s window, fed into a frozen self-supervised vision transformer with Top-k multiple-instance pooling to flag candidates; (2) a vector-quantized native background index (K=1216) with block-bootstrap percentile thresholds to re-score candidates and reject domain-shift artifacts; (3) an empirical family-wise PEM coherence veto calibrated per event from time-shifted surrogate pairs; and (4) a physical coincidence statistic—normalized cross-correlation of whitened strain in the candidate's band over the light-travel lag window—thresholded at the pooled 99th percen","core_discovery":"After expanding the search to four temporal scales and correcting for the non-stationary drift of the detector noise manifold with a native-background index, the author finds that the apparent novel glitch families reported by uncalibrated unsupervised pipelines do not survive: 71.7% of candidates are statistically indistinguishable from background, and the rest coalesce into a single macro-cluster with no discrete recurring substructure—a topology shared by pristine background. A physically calibrated PEM veto removes one isolated candidate via control-line coupling, leaving one uncatalogued instrumental outlier. The author then replaces the embedding-similarity coincidence test, which cann","pith_inferences":["If the central claim holds, earlier unsupervised reports of exotic glitch families in early O4a may have been measuring uncorrected domain shift; re-running those analyses with a native-background re-scoring stage should collapse their surviving families accordingly.","The surviving L1 singleton—the highest-anomaly candidate, outside the macro-cluster, with no PEM coupling—is a concrete target for future public O4a glitch catalogs: matching it, if such a catalog becomes available, would test whether it is truly uncatalogued.","A causal variant of the domain-shift defense, with a rolling trailing-window background index, could turn this retrospective survey into low-latency glitch monitoring; the paper flags this as a natural extension, and its viability would hinge on how much the 28.3% survival fraction drifts with shorter baselines.","The coincidence result's pooled null is the natural stress point: a public release of per-event dense shift ladders would either confirm or refute the zero-coincidence claim, since the paper itself notes that per-candidate significance requires a denser shift ladder than the 4–8 shifts used here."],"forward_implications":["Raw unsupervised candidate counts in O4a are dominated by noise-manifold drift: roughly 72% of multi-scale candidates are indistinguishable from native background after recalibration, so any unsupervised glitch claim on this epoch must include a native-background defense.","Within the morphologies the pipeline can see, there is no evidence of a discrete, recurring novel instrumental glitch family in early O4a; all survivors merge into one macro-cluster that pristine background exhibits even more strongly.","No cross-detector coincident unmodeled transient was found among 8,749 candidates, with the physical statistic validated to recover 100% of injected coincident structured morphologies at sufficient SNR; the null is a statement with demonstrated power for structured transients and is weaker for incoherent broadband bursts.","The 90% upper limits R90 ≤ 5.83 yr−1 (H1) and ≤ 5.63 yr−1 (L1) bound the rate of novel uncatalogued per-detector instrumental morphologies over CAT1-gated livetime, and carry no astrophysical interpretation.","The embedding-similarity coincidence test is retired; a physical normalized cross-correlation over the light-travel lag window is the validated replacement, because injected identical waveforms score 1.00 while independent noise scores 0.043."],"fun_headline_variants":["No new glitch families in LIGO O4a after recalibration","Unsupervised glitch search finds nothing left after physical vetting","Calibration kills ghost glitches: zero coincident in O4a","DANTE V3 yields zero novel glitch families in LIGO O4a","O4a glitch hunt: no novel morphologies survive physical checks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 'no coincident event' result rests on the assumption that the pooled time-shifted noise measurements across 8,749 events form one valid reference distribution, even though each event supplies only 4–8 shifted samples; if the detectors' noise statistics change between events, the shared threshold could conceal a real coincidence or create false assurance.","fun_headline_variants_meta":{"raw":{"variants":["No new glitch families in LIGO O4a after recalibration","Unsupervised glitch search finds nothing left after physical vetting","Calibration kills ghost glitches: zero coincident in O4a","DANTE V3 yields zero novel glitch families in LIGO O4a","O4a glitch hunt: no novel morphologies survive physical checks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000186,"raw_usage":{"total_tokens":1240,"prompt_tokens":899,"completion_tokens":341,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":241}},"tokens_in":643,"tokens_out":341,"duration_ms":381001,"temperature":1.0,"reasoning_tokens":241,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:53:43.306648+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a random sample of candidates, build a dense per-event null using 100+ time shifts instead of the 4–8 available (for example, shifts of ±0.1, ±0.2, ..., ±8 s). If more than about 1% of on-source events exceed their per-event 99th percentile, the pooled-threshold design is masking event-level coincidences and the zero-coincidence claim fails. A second check: compare per-event null means between early and late O4a sessions; if they differ beyond sampling noise, the pooled null is not exchangeable.","supporting_citations":[],"review_version":1}