{"id":"1c32a57b-9adc-4b4c-be71-1c987d520a92","arxiv_id":"2506.22588","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Sequential anytime-valid tests for sparse Gaussian anomalies become possible exactly at the fixed-sample detection time t*=T*rho(beta*), and an adaptive mixture test achieves the same threshold without knowing the anomaly parameters.","lead":"This paper builds sequential tests that watch many data streams and can detect a rare weak signal while keeping false alarms controlled no matter when you decide to stop. It also identifies the earliest moment such detection becomes mathematically possible, and provides a practical adaptive test that needs no prior knowledge of the signal's rarity or strength.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The displayed δ-grid in Eq. (9) has the wrong sign: with the formula as printed, T_i=2lnK/δ_i^2 lies in [1/C,1], so the proof of Theorem 2.6 cannot construct a grid point satisfying (7) when T*>1; the adaptive threshold theorem is therefore not established as stated.","rationale":"The reader's weakest_assumption correctly locates the fragile step. The oracle theorems (2.1, 2.3) are internally coherent, and the proof strategy (truncated first moments, Cramér–Chernoff bounds, Rényi-divergence lower bounds) is substantially worked out; the simulations are described in enough detail to re-implement, and the fixed-sample by-product discussion is honest. The load-bearing problem is localized to the adaptive construction: every positive guarantee for the unknown-parameter test passes through Proposition 2.4 via a nearest-grid-point argument, and the sign mismatch in Eq. (9) breaks that argument for the main case T*>1. I do not see a reason to reject the paper: the mismatch is identifiable, and the proof strongly suggests the intended grid uses exp(−i/⌈εK⌉). But the theorem as stated is not established until the definition and proof are aligned; hence the conditional verdict should stand. I agree with the reader's weakest_assumption; the only addition is that a corrected negative exponent would still leave T*<1 outside the stated proof, so the theorem's wording 'T*<C' should also be scoped to T*≥1 or handled separately.","tokens_in":55179,"tokens_out":17831,"duration_ms":193236,"concrete_test":"Analytical check: with Eq. (9) as printed, for K=1000, β*=0.85, T*=40, C=5T*, compute the nearest grid point to (ε*,δ*) and evaluate |d_T|K^{1−β°}; the plus sign gives T°≤1, so d_T≥1−1/sqrt(40)≈0.842 and the product diverges, violating (7). Repeat using the proof's T_i=exp(i/⌈Kε°⌉), i.e. δ_i=sqrt(2lnK exp(−i/⌈εK⌉)): the nearest grid point then satisfies T°<T* and |d_T|K^{1−β°}=O(1), restoring the proof. A more expensive but decisive simulation: implement both grids for K=1000, T*=40, β*=0.85 and compare cumulative rejection rates of the mixture SPRT; the printed grid should miss the threshold t* badly, while the corrected grid reproduces Theorem 2.7.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central adaptive claim (Theorems 2.6 and 2.7) depends on the grid G_C containing a point (ε°,δ°) that meets the misspecification window of Proposition 2.4: |d_T|K^{1−β°}=O(1) and |d_β|lnK→0. The displayed grid in Eq. (9) sets δ_i = sqrt(2 lnK exp(i/⌈εK⌉)); hence T_i = 2lnK/δ_i^2 = exp(−i/⌈εK⌉), so all grid points have T°∈[1/C,1]. The proof of Theorem 2.6 instead uses T_i = exp(i/⌈Kε°⌉), with grid points up to T°=C, and needs to bracket δ* from below by a grid point with T°≤T*. For any T*>1, taking the displayed definition literally forces T°≤1, so d_T=1−sqrt(T°/T*) is bounded away from zero (at best T°=1), and |d_T|K^{1−β°} diverges. Condition (7) then fails, so Proposition 2.4 cannot be invoked and the adaptive log-growth/stopping guarantees are not established for the stated regime T*<C. The bracketing inequality in the proof is also written with δ_i and δ_{i+1} in the wrong order for the intended T_i<T*≤T_{i+1}. This looks correctable (the intended grid presumably has exp(−i/⌈εK⌉)), but as printed it is a genuine definition–proof mismatch in the load-bearing construction. A secondary gap remains even after the sign fix: such exponential grids cover T≈[1,C], while the theorem allows any T*<C including T*<1.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops anytime-valid (AV) tests for detecting sparse anomalies among K Gaussian streams, under the contamination model where each stream is anomalous with probability ε and has mean shift δ. The authors analyze the oracle likelihood-ratio martingale E*_t and show that, for ε*=K^{-β*} and δ*=sqrt(2(1/T*) ln K), the expected log-growth E*[ln E*_t] transitions from 0 to infinity at t*=T*ρ(β*), and that the corresponding SPRT stops with probability tending to 0 before t* and to 1 after t* (Theorems 2.1 and 2.3). They then construct an adaptive mixture martingale E_t(Π) based on a K-dependent grid prior and claim that it attains the same threshold behavior without knowing (ε*,δ*) (Theorems 2.6 and 2.7). Numerical simulations compare the proposed tests with fixed-sample likelihood-ratio and higher-criticism benchmarks.","tokens_in":55636,"tokens_out":15034,"duration_ms":150616,"significance":"If fully established, the paper would be a substantial contribution: it transfers the fixed-sample sparse-detection phase transition into a sequential anytime-valid setting, provides an explicit and computationally tractable adaptive construction (O(K^{3/2} ln K) per step), and backs the theory with finite-K simulations. The oracle analysis is largely self-contained, and the lower-bound part legitimately imports the fixed-sample HC risk result rather than assuming the desired conclusion. The adaptive construction is a genuine methodological advance, and the Monte Carlo calibration of thresholds is a reasonable and clearly described practical choice. The main caveat is that the adaptive threshold theorems are not established as printed because the displayed grid in Eq. (9) is inconsistent with the proof of Theorem 2.6 and does not cover the full stated parameter range.","major_comments":[{"comment":"As printed, the grid in Eq. (9) reads δ_i = sqrt(2 ln K exp(i/⌈εK⌉)) (or equivalently sqrt(2 ln K) exp(i/(2⌈εK⌉))), so that T_i = 2 ln K / δ_i^2 = exp(-i/⌈εK⌉) lies in [1/C,1]. The proof of Theorem 2.6, however, uses T_i = exp(i/⌈Kε°⌉) and needs a grid point with T° ≤ T* and T* < C. For T* > 1 the displayed grid cannot provide such a point, condition (7) of Proposition 2.4 cannot be verified, and Theorem 2.6 is not established for the stated regime. This is a load-bearing definition-proof mismatch, though it appears correctable by writing the intended grid as δ_i = sqrt(2 ln K / exp(i/⌈εK⌉)).","section":"Eq. (9) and proof of Theorem 2.6 (Section B.4)"},{"comment":"The bracketing inequality in the proof is also reversed. The proof states δ_i < δ* ≤ δ_{i+1} with T_i = exp(i/⌈Kε°⌉) and T_{i+1} = exp((i+1)/⌈Kε°⌉). Since T_i is increasing in i, δ_i = sqrt(2 ln K / T_i) is decreasing in i, so the displayed ordering is impossible for the claimed choice T°=T_i ≤ T*. The intended ordering should be δ_{i+1} < δ* ≤ δ_i, which is what would make T_i ≤ T* and d_T = 1 - sqrt(T_i/T*) nonnegative and of order K^{-(1-β°)}.","section":"Proof of Theorem 2.6 (Section B.4)"},{"comment":"Even after correcting the sign in Eq. (9), the proposed grid with T_i = exp(i/⌈εK⌉) and i ≥ 1 covers only T° ≥ exp(1/⌈εK⌉) ≈ 1. The theorems assume only T* < C with C > 1 and allow 0 < T* < 1. For such T* no grid point can satisfy T° ≤ T* together with |d_T| K^{1-β°} = O(1), so the proof of Theorem 2.6 (and hence Theorem 2.7, which relies on the same grid point) does not cover the full stated range. The statements should either restrict T* to T* > 1 or extend the grid to cover T° ∈ [1/C,C].","section":"Theorem 2.6 and Theorem 2.7, parameter range T*<C"}],"minor_comments":[{"comment":"The text defines β* = ln(1/ε*), but since ε* = K^{-β*}, this should read β* = ln(1/ε*)/ln K.","section":"Section 1.3, paragraph after Theorem 2.1"},{"comment":"The definition says 'Let C > 0', but the grid construction and Theorem 2.6 require C > 1; also the K-dependence of G_C and Π is suppressed, which can confuse the reader. Writing G_C(K) and Π_K would clarify the construction.","section":"Definition 2.5"},{"comment":"Eq. (56) writes P*{τ* ≤ t} = P*{max_{s≤t} E_s(Π) ≥ 1/α}, but τ* was defined earlier for the oracle SPRT; the stopping time here should be τ_Π.","section":"Proof of Theorem 2.7, Eq. (56)"},{"comment":"The theorem states α ∈ (0,1], while all other level statements use α ∈ (0,1); this should be harmonized.","section":"Theorem 2.7"},{"comment":"There is a typo 'As we will wee' that should read 'As we will see'.","section":"Section 2.2, opening paragraph"}],"recommendation":"major_revision","confidential_remarks":"The central idea is important and the main proof strategy appears sound. The Eq. (9) sign error, the reversed bracketing inequality, and the T*<1 gap are localized and correctable, but they currently invalidate the adaptive theorems as stated. I would be willing to accept a revised version that fixes these points and either proves or explicitly restricts the parameter range."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core result is real and worth engaging with. The paper gives the first anytime-valid tests for sparse anomaly detection in the large-K regime, and the oracle analysis is the genuine contribution: they show the likelihood-ratio martingale has log-growth 0 before t*=T*rho(beta) and infinity after, and that the SPRT stops accordingly. The lower bound imports the fixed-sample HC risk result, which is legitimate rather than circular. The adaptive mixture martingale with a K-dependent discrete grid is a sensible construction and provably matches the oracle threshold up to a couple of proof bugs. The O(K^{3/2}) computational claim is a nice practical bonus, and the simulations include honest comparisons against HC-Bonferroni and plug-in methods. No code or data, but the simulation details are adequate for reimplementation.\n\nOn the soft spots. The stress-test's main claim is a misreading: the grid in Eq. (9) is delta_i = sqrt(2 ln K / exp(i/ceil(eps K))), so T_i = 2lnK/delta_i^2 = exp(i/ceil(eps K)), which ranges from about 1 to C, not [1/C,1]. The grid sign is fine. What is actually wrong is the bracketing inequality in the proof of Theorem 2.6: it writes delta_i < delta* <= delta_{i+1}, but since T_i increases with i, delta_i decreases, so the correct bracket is delta_{i+1} < delta* <= delta_i (equivalently T_i <= T* < T_{i+1}). The intended argument is obvious and fixable. Second, the construction genuinely does not cover T*<1: the grid has no points with T_i < 1, so for T*<1 the misspecification condition (7) can fail because d_T is bounded away from zero and K^{1-beta} diverges. The theorem only assumes T*<C with C>1, so this is a real gap, though narrow; restricting to T*>=1 or extending the grid with negative indices would fix it. Third, the claim that O(K^{3/2}) complexity is optimal over all discrete mixtures is asserted without proof; plausible but not established.\n\nWho is this for? Anyone working in anytime-valid inference or high-dimensional sparse detection. The oracle threshold result alone is worth a serious referee, and the adaptive construction is correct up to the fixable issues above. I would send it to peer review and expect acceptance after a minor revision that fixes the bracket order and states the T*>=1 regime (or handles T*<1).","headline":"First anytime-valid treatment of sparse anomaly detection with a real oracle threshold theorem; the adaptive construction has a proof typo and a T*<1 gap, but the stress-test's main sign complaint misreads the grid.","tokens_in":56129,"tokens_out":5928,"would_cite":true,"duration_ms":58133,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F03","62L10","60G40"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that anytime-valid tests for sparse anomalies have a sharp detection moment: before t*, no test martingale accumulates evidence; after it, the oracle and an adaptive mixture martingale both stop with probability tending…","keywords":["anytime-valid tests","sparse normal means","minimax hypothesis testing","sequential testing","likelihood ratio martingale","mixture SPRT","higher criticism","detection boundary"],"falsifier":"Take K=$10^{5}$, β*=0.75, T*=40, so ε*=$K^{{-β*}}$ and δ*=√(2 ln K/40), and simulate the adaptive mixture martingale E_t(Π) under P*. The theorems predict P*{τ_Π≤t}→0 for t<t* and →1 for t>t*, with t*=40ρ(0.75)=40(1−√0.25)^2=10; a simulation whose transition occurs at a different time would refute the claimed threshold. A separate check is whether the printed grid contains any δ with 2 ln K/$δ^{2}$>1, as the proof requires.","tokens_in":54988,"feed_emoji":"🔍","tokens_out":7935,"duration_ms":81893,"temperature":0.7,"pith_summary":"This paper establishes a sharp time threshold for sequentially detecting sparse anomalies among K Gaussian data streams. Under the scaling ε*=$K^{{-β*}}$ and δ*=√(2 ln K / T*), it shows that before t*=T*ρ(β*) no test martingale can accumulate evidence in expected-log terms, while after t* the oracle likelihood-ratio martingale grows without bound and its one-sided SPRT stops with probability tending to one. The main constructive contribution is a mixture likelihood-ratio martingale over a K-dependent grid that attains the same two-sided threshold without knowing ε* or δ*, and can be computed in O($K^{{3/2}}$) operations per time step. Because any anytime-valid test yields a fixed-sample test, this also produces an adaptive fixed-sample test matching the known detection phase transition.","feed_headline":"Adaptive sequential test hits oracle's anomaly-detection threshold","feed_subtitle":"A mixture martingale stops with probability one right after detection becomes possible, without knowing the sparsity or signal.","key_machinery":"The central object is the likelihood-ratio test martingale, a nonnegative martingale starting at one under the null; with known parameters it is E_t^*=∏_i[(1−ε*)+ε* p_{δ*}(X_{i,1..t})/p_0(X_{i,1..t})], and the adaptive construction replaces the unknown parameters by a uniform mixture E_t(Π) over a grid G_C. The analysis runs through Kullback-Leibler divergence, second-moment concentration bounds, and the piecewise detection boundary ρ(β), which locates the threshold t*=T*ρ(β*). The grid prior is designed so that, under the misspecification bounds |1−√(T_K/T*)|$K^{{1−β_K}}$=O(1) and |β_K−β*|ln K→0, at least one grid point is close enough to the truth for Proposition 2.4 to transfer the oracle threshold behavior to the mixture martingale.","core_discovery":"At fixed time t and growing K, the problem is governed by a detection moment t*=T*ρ(β*), where ρ(β*) is the piecewise boundary (β*−1/2 on (1/2,3/4) and (1−√(1−β*))^2 on [3/4,1]). Theorem 2.1 shows that the log-optimal oracle likelihood ratio E_t^* satisfies E*[ln E_t^*]→0 for t<t* and →∞ for t>t*, so no test martingale can have nontrivial expected-log growth before the threshold and the oracle must grow after it. Theorem 2.3 shows the same dichotomy for the stopping time of the oracle one-sided SPRT. Theorem 2.6 and Theorem 2.7 extend both statements to the adaptive mixture martingale E_t(Π), a uniform mixture of likelihood ratios over a grid of (ε,δ) values, which achieves the same threshold behavior without knowing the true parameters. The paper claims the sequential version of sparse-anomaly testing has its own phase transition, related but not implied by the fixed-sample transition.","pith_inferences":["The proof's grid as displayed appears to cover only T_i≤1, while Theorem 2.6 needs grid points with T_i up to C; if the displayed formula is literal, the adaptive guarantee for T*>1 is an unproven step rather than a closed result.","The threshold structure suggests a practical design rule: choose a desired sparsity-sensitivity pair (β*,T*), read off t*=T*ρ(β*) as the promised detection time, and use the mixture martingale without estimating parameters.","The detection-but-not-identification window between t* and T* implies that early rejection can be trusted as evidence that anomalies exist, but not as a list of which streams are anomalous; identification would require the stronger regime r>β.","A natural extension is to test whether the same threshold transfers to rank-based or distribution-free stream statistics; the Gaussian machinery suggests controlling a Kullback-Leibler-type divergence is the key ingredient, but that transfer is not shown in the paper."],"forward_implications":["Continuous monitoring of many streams for sparse anomalies acquires a sharp detection moment: before t*, no test martingale can have nontrivial expected-log growth, so early stopping is genuinely impossible, not merely hard.","After t*, the mixture SPRT stops with probability tending to one while retaining type-I error control at every stopping time, so practitioners can monitor indefinitely without a time-horizon-dependent Bonferroni penalty.","The mixture likelihood ratio yields a fixed-sample adaptive test whose power transitions at the same boundary as the oracle likelihood ratio and the higher-criticism test, as a direct by-product.","The O(K^{3/2} ln K) per-time-step computational cost makes the adaptive test feasible for large K, and the cost depends only logarithmically on the horizon bound C.","Finite-K simulations show the low-to-high power transition sharpening as K grows, consistent with the asymptotic dichotomy."],"supporting_citations":[{"why":"Supplies the fixed-sample detection boundary ρ(β), the higher-criticism statistic, and the asymptotic power phase transition that the sequential results extend and benchmark against.","marker":"[15]"},{"why":"Establishes the minimax detectability boundary for sparse signal detection that fixes the β-r parametrization used throughout.","marker":"[27]"},{"why":"Provides adaptive fixed-sample tests achieving the boundary, the starting point for the paper's adaptive construction.","marker":"[29]"},{"why":"Sharpens the power boundary and documents that no adaptive fixed-sample test dominates uniformly, justifying the choice of benchmarks.","marker":"[39]"},{"why":"Supplies the anytime-valid test-martingale framework and the log-optimality and e-value analysis that the sequential design builds on.","marker":"[40]"},{"why":"Gives the admissibility result showing anytime-valid tests must be based on nonnegative martingales, grounding the martingale construction.","marker":"[41]"},{"why":"Identifies the likelihood ratio as the log-optimal test martingale, the basis for the oracle E_t^*.","marker":"[33]"},{"why":"Introduces mixture sequential probability ratio tests, the template for the adaptive mixture martingale.","marker":"[43]"},{"why":"Supplies the SPRT framework and its stopping-time analysis, used for τ* and τ_Π.","marker":"[48]"},{"why":"Supplies the martingale maximal inequality giving the type-I error guarantee P0{sup_t E_t ≥ 1/α}≤α used for all anytime-valid thresholds.","marker":"[51]"}],"fun_headline_variants":["Sequential sparse-anomaly test matches oracle's threshold","Anytime-valid test reaches oracle's detection threshold","Phase transition in sequential sparse anomaly testing","Mixture martingale hits oracle threshold without parameter knowledge","Sequential test achieves oracle's sparse detection limit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The adaptive guarantee rests on the K-dependent grid prior Π=Uniform(G_C) containing a point whose misspecification bias satisfies |1−√(T_K/T*)|$K^{{1−β_K}}$=O(1) and |β_K−β*| ln K→0, with T*<C. The paper's displayed grid G_{δ,C}(ε) uses δ_i=√(2 ln K exp(i/⌈εK⌉)), which places T_i=2 ln K/$δ_i^{2}$ in [1/C,1], while the proof of Theorem 2.6 requires grid values with T_i up to C; read literally, the displayed grid cannot cover T*>1 and the stated regime of Theorem 2.6 is not established.","fun_headline_variants_meta":{"raw":{"variants":["Sequential sparse-anomaly test matches oracle's threshold","Anytime-valid test reaches oracle's detection threshold","Phase transition in sequential sparse anomaly testing","Mixture martingale hits oracle threshold without parameter knowledge","Sequential test achieves oracle's sparse detection limit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000437,"raw_usage":{"total_tokens":2244,"prompt_tokens":989,"completion_tokens":1255,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":605,"completion_tokens_details":{"reasoning_tokens":1181}},"tokens_in":605,"tokens_out":1255,"duration_ms":10883,"temperature":1.0,"reasoning_tokens":1181,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:04:15.237266+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take K=$10^{5}$, β*=0.75, T*=40, so ε*=$K^{{-β*}}$ and δ*=√(2 ln K/40), and simulate the adaptive mixture martingale E_t(Π) under P*. The theorems predict P*{τ_Π≤t}→0 for t<t* and →1 for t>t*, with t*=40ρ(0.75)=40(1−√0.25)^2=10; a simulation whose transition occurs at a different time would refute the claimed threshold. A separate check is whether the printed grid contains any δ with 2 ln K/$δ^{2}$>1, as the proof requires.","supporting_citations":[{"cited_title":"and JIN, J","cited_arxiv_id":null,"evidence_quote":"Supplies the fixed-sample detection boundary ρ(β), the higher-criticism statistic, and the asymptotic power phase transition that the sequential results extend and benchmark against."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the minimax detectability boundary for sparse signal detection that fixes the β-r parametrization used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides adaptive fixed-sample tests achieving the boundary, the starting point for the paper's adaptive construction."},{"cited_title":"and STEWART, M","cited_arxiv_id":null,"evidence_quote":"Sharpens the power boundary and documents that no adaptive fixed-sample test dominates uniformly, justifying the choice of benchmarks."},{"cited_title":"(1939).Étude Critique de la Notion de Collectif.Thèses de l’entre-deux-guerres218","cited_arxiv_id":null,"evidence_quote":"Supplies the martingale maximal inequality giving the type-I error guarantee P0{sup_t E_t ≥ 1/α}≤α used for all anytime-valid thresholds."}],"review_version":1}