{"id":"81f15790-0709-4c2d-beaa-5ff3533099dc","arxiv_id":"2502.10380","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors construct a flexible family of sharp asymptotic time-uniform confidence sequences for a mean, based on weighted Wiener-process and Brownian-bridge supremum limits.","lead":"This paper presents a new family of confidence sequences, which are statistical intervals that stay valid no matter when data collection stops. The method lets researchers choose the shape of the interval boundaries flexibly and proves the resulting intervals reach the nominal coverage level asymptotically.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 4.1's normalization is inconsistent with Donsker: √m∑ diverges at t≈m, and Eq (1)'s interval does not invert to the sup|ρW| limit. The printed central claim is unsupported; the finite-endpoint issue is secondary.","rationale":"The paper's contribution is the flexible sharp confidence sequence in Theorem 2.2, built on Theorem 4.1. The reader's finite-endpoint concern is legitimate: for ρ with finite endpoint, C_t collapses to a point for t > m e_ρ, and Definition 2.1/Theorem 2.2 do not include the convention that no data are collected after that point. But that issue is a defined convention fix and does not affect the recommended infinite-endpoint family. The normalization mismatch is more severe and affects every ρ, including the recommended family. Standard Donsker scaling is m^{-1/2} for partial sums, while Theorem 4.1 as displayed uses √m times the raw partial sum; the expression therefore diverges in distribution at t ≈ m. The proof repeats this normalization in equation (4), so it is not a one-off typo in the theorem statement. Inverting Eq (1) also gives a boundary c√t ρ(t/m) rather than the cρ(s) that would match sup|ρW|, so the claimed quantile is not the right critical value for the displayed intervals. I have no reason to doubt the authors' intent; the construction may well be salvageable by correcting b_t to something like √m/(tρ(t/m)) and the theorem's normalization to m^{-1/2}. But as printed, the central theorem is not internally consistent and cannot be accepted or conditionally accepted on the basis of the current text. Hence the verdict should be UNVERDICTED pending a source-level check of the normalization, rather than CONDITIONAL on the finite-endpoint convention.","tokens_in":9893,"tokens_out":21261,"duration_ms":230290,"concrete_test":"Simulate iid N(0,1) data with m=10^4 and ρ(s)=(1+s)^{-1}. First compare the empirical distribution of M_m = sup_{t≥1} |ρ(t/m)√m σ̂_t^{-1} ∑_{i=1}^t (X_i−μ)| over many replications with the claimed limit sup_{y>0}|ρ(y)W(y)|; under the printed normalization M_m should be of order m and not converge to that law. Second, compute the empirical coverage of the intervals (1) over all t for μ=0 and α=0.05. If coverage is close to 1−α the printed normalization is correct and this concern fails; if it is clearly different, the formula in Theorem 4.1 or Eq (1) must be corrected before the paper's central claim can be used.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central result Theorem 2.2 (and its test dual Theorem 3.1) is stated to be a direct consequence of Theorem 4.1. As printed, Theorem 4.1 cannot hold. For iid observations with finite variance, Donsker's theorem gives m^{-1/2}∑_{j=1}^{⌊ms⌋}(X_j−μ) ⇒ σ W(s). The theorem instead contains √m ∑_{j=1}^t (X_j−μ), which at t=m is m times a tight sequence and hence diverges in probability; the proof's equation (4) uses the same normalization. Thus the asserted convergence to sup_{y>0}|ρ(y)W(y)| is not a consequence of the stated assumptions. Independently, Eq (1) does not invert to that limit: the coverage event |μ̂_t−μ| ≤ σ̂_t c √(m/t) ρ(t/m) is algebraically equivalent to S_t/(σ̂_t√m) ≤ c√t ρ(t/m). For t ≍ m the right-hand side grows like √m, so the asymptotic false-rejection probability under the printed formulas is not α; it tends to a different value. Therefore at least one of Eq (1), Theorem 4.1, or the proof has the wrong normalization. As written, the paper does not establish its advertised sharp anytime-valid class. The finite-endpoint objection raised by the reader is genuine but is a convention-level fix compared with this normalization problem.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a new class of asymptotic time-uniform confidence sequences for the mean of iid observations. The sequence is defined by C_t(m;α) = μ̂_t ± σ̂_t c_α(ρ) √(m/t) ρ(t/m), where ρ is a user-chosen weight function satisfying growth conditions A1 and A2, and c_α(ρ) is the (1−α)-quantile of sup_{y>0} |ρ(y)W(y)| for a standard Wiener process. The main theorem (Theorem 2.2) claims that this sequence is asymptotically sharp, i.e., lim_{m→∞} P(μ ∈ C_t for all t∈N) = 1−α. The paper also presents a dual sequential testing formulation (Theorem 3.1), a stopping guarantee under alternatives (Theorem 3.2), and a limit theorem (Theorem 4.1) that is stated to be the underlying asymptotic tool. A special family of weight functions is shown to reduce the critical constant to a Brownian-bridge supremum (Proposition 4.3).","tokens_in":10132,"tokens_out":16798,"duration_ms":144499,"significance":"If the central claim were correct, the paper would offer a flexible and easily implementable class of anytime-valid confidence sequences, with the critical constant obtained from a standard Wiener-process or Brownian-bridge supremum. The limit theorem (Theorem 4.1) is plausible and its proof sketch is coherent for infinite-endpoint weight functions, and the connection to change-point monitoring is a valuable perspective. However, the advertised statistical construction does not follow from the limit theorem, and the stated confidence sequence does not achieve the claimed coverage probability. The error is load-bearing, so the paper does not establish its main contribution.","major_comments":[{"comment":"The claimed inversion of Theorem 4.1 is algebraically incorrect. The coverage event μ ∈ C_t(m;α) is equivalent to |S_t|/(σ̂_t √m) ≤ c_α(ρ) √t ρ(t/m). For t = ⌊my⌋, the left-hand side converges weakly to |W(y)|, while the right-hand side is c_α(ρ) √m √y ρ(y), which diverges for any y with ρ(y)>0. Consequently, the probability that any false rejection occurs tends to 0, not to α, so the sequence (1) is not sharp; its coverage tends to 1. The correct inversion of the limit sup_y |ρ(y)W(y)| would require a half-width proportional to √(m/t) / (√t ρ(t/m)), not √(m/t) ρ(t/m). This invalidates the central claim of Theorem 2.2 and its dual Theorem 3.1.","section":"Section 2, Eq. (1) and Theorem 2.2"},{"comment":"For weight functions with finite e_ρ, ρ(s)=0 for s>e_ρ, so b_t(m;ρ)=0 for t>m e_ρ. Then C_t collapses to the degenerate interval {μ̂_t}, which has probability zero of containing a fixed μ under the model. The sentence 'at most ⌊m e_ρ⌋ data points are collected' is not encoded in Definition 2.1, which requires coverage for all t∈N, nor in the statement of Theorem 2.2. Without an explicit convention such as C_t = R after the endpoint, the all-t coverage claim is false for finite-endpoint weight functions.","section":"Section 2, Definition 2.1 and Theorem 2.2 (finite endpoint e_ρ)"},{"comment":"The proof of Theorem 4.1 appears to handle only the case where the weight function decays sufficiently at infinity. For a finite endpoint e_ρ with ρ(e_ρ)>0, the tail bound in Eq. (8) requires sup_{s>V} s^{1−γ2} ρ(s) → 0 as V increases, which fails if ρ has a positive limit at e_ρ. The theorem is therefore not established for all ρ allowed by Assumption B; it is only clearly valid for infinite-endpoint ρ satisfying the stated decay.","section":"Section 4, proof of Theorem 4.1"}],"minor_comments":[{"comment":"The normalization in Theorem 4.1 is 1/√m times the partial sum, which is the correct Donsker scaling; the skeptical reading that the theorem contains √m times the sum is not supported by the manuscript text.","section":"Section 4, Eq. (4)"},{"comment":"The paper would benefit from an explicit statement that for finite-endpoint ρ the procedure stops at ⌊m e_ρ⌋ and the confidence sequence is defined as the whole real line afterwards; currently this is only mentioned informally.","section":"Section 2, paragraph after Eq. (1)"},{"comment":"There are minor inconsistencies in author name spellings, for example 'Stoehr' vs 'Stöhr', and the use of a PhD thesis (Stöhr 2019) for a key lemma; the authors should provide a more accessible reference if possible.","section":"References"}],"recommendation":"reject","confidential_remarks":"The paper's limit theorem (Theorem 4.1) seems plausible, but the mapping from that theorem to the proposed confidence sequence is wrong, making the central claim false as stated. This is not a minor typo: the confidence interval in Eq. (1) is not the inversion of the weighted sup-norm statistic, and the stated coverage probability is not attained. A correction would require redefining the interval width, which changes the core contribution. The finite-endpoint issue further weakens the paper. I recommend rejection, though the authors may be able to salvage a version of the idea with a corrected width."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The centerpiece is a genuinely new flexible boundary class for asymptotic time-uniform confidence sequences, but the printed interval formula does not invert the limit theorem the authors prove. As written, the confidence sequence is not sharp — the coverage probability tends to 1, not 1−α — so the main advertised result needs correction.\n\nWhat I like: the construction via generalized Hájek–Rényi inequalities plus a functional CLT is different from Waudby-Smith et al.'s strong-invariance route. The special weight family ρ(s)=(1+s)^{γ1+γ2−1/2}/s^{γ1} reduces the critical constant to a Brownian-bridge supremum, which is computable. The duality to sequential change-point monitoring is well explained, and the hierarchical-testing connection is a nice byproduct. No circularity: the constants come from external Wiener-process quantiles.\n\nThe soft spots are real. The algebra in Eq (1) is off. The coverage event |μ̂_t−μ| ≤ σ̂_t c √(m/t) ρ(t/m) is equivalent to |S_t|/(σ̂_t√m) ≤ c √t ρ(t/m). For t ≈ ms the right side grows like c √m √s ρ(s), so the probability that all these inequalities hold tends to 1. The limit theorem, by contrast, gives sup_y |ρ(y)W(y)| for the statistic ρ(t/m) S_t/(σ̂_t√m). Inverting that statistic correctly gives an interval width of √m/(t ρ(t/m)), not the printed √(m/t) ρ(t/m). The theorem and the confidence sequence are not connected as stated. Also, the statement of Theorem 4.1 as printed shows √m multiplying the sum, which would diverge; the proof uses /√m. That is likely a typo, but it must be fixed before the claim can be evaluated.\n\nThe reader's finite-endpoint concern is secondary but genuine: with a finite endpoint e_ρ, the interval collapses to a point for t > m e_ρ, and the all-t coverage definition needs an explicit convention (e.g., C_t = R after the endpoint). Remark 4.2 is a remark, not a theorem, so the time-series generalization shouldn't be cited as proven.\n\nBottom line: the proof machinery is sound for the corrected setting, and the parametric family is practical. This paper deserves a serious referee, but the central formula has to be fixed and the finite-endpoint convention added before it can be used. I'd recommend major revision, not acceptance.","headline":"Promising new boundary class for asymptotic confidence sequences, but the printed interval width doesn't invert the limit theorem, so the sharpness claim doesn't hold as stated.","tokens_in":10704,"tokens_out":12650,"would_cite":false,"duration_ms":114107,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G10","62G15","62G20","62L10"],"pacs":[],"model":"deepseek-v4-flash","headline":"A flexible family of boundary functions yields sharp asymptotic confidence sequences that hold uniformly over all stopping times.","keywords":["Anytime-valid inference","Time-uniform confidence sequence","Nonparametric","Sequential testing","Sharp asymptotic confidence sequence","Weighted Wiener process","Hájek–Rényi inequality","Brownian bridge"],"falsifier":"Take $\\rho(s)=\\mathbb{1}_{(0,1]}(s)$, so $e_\\rho=1$, choose $m=100$, $\\alpha=0.05$, and simulate iid standard normal data. Compute $C_t(100;0.05)=\\hat\\mu_t\\pm \\hat\\sigma_t c_{0.05}(\\rho)\\sqrt{100/t}$ for $t=1,2,\\ldots$ and check whether $\\mu=0$ lies in all intervals. Since for $t>100$ the interval has zero width and $\\hat\\mu_t\\neq 0$ almost surely, the simultaneous coverage will be 0, directly contradicting Theorem 2.2 unless $C_t$ is explicitly set to $\\mathbb{R}$ after $t=100$; repeating the simulation with that convention should yield coverage near 0.95.","tokens_in":9591,"feed_emoji":"⏱","tokens_out":16868,"duration_ms":133030,"temperature":0.7,"pith_summary":"This paper constructs a new family of anytime-valid confidence intervals for the mean of independent, identically distributed data: intervals $C_t(m;\\alpha) = \\hat\\mu_t \\pm \\hat\\sigma_t \\cdot c_\\alpha(\\rho) \\cdot \\sqrt{m/t}\\,\\rho(t/m)$ that are guaranteed to contain the true mean at every time $t$ simultaneously, with probability $1-\\alpha$ in the limit as the initial sample size $m$ grows. The shape of the boundary is chosen by a weight function $\\rho$ that only needs to satisfy mild growth conditions near zero and infinity, and the critical constant $c_\\alpha(\\rho)$ comes from the distribution of a supremum of a weighted Wiener process. The main theorem states that this sequence is a sharp asymptotic $(1-\\alpha)$-confidence sequence, meaning the simultaneous coverage probability converges to exactly $1-\\alpha$, not just at least. This matters because the existing nonparametric confidence sequences had a fixed boundary shape, whereas here the practitioner can spend the error budget unevenly over early and late times. The same construction dualizes to sequential tests with asymptotic level $\\alpha$, including one-sided hierarchical tests whose family-wise error rate is controlled.","feed_headline":"Weight functions unlock sharp anytime-valid confidence sequences","feed_subtitle":"Pick the boundary shape that suits your spending plan; coverage converges to exactly 1–α at every stopping time.","key_machinery":"The central mechanism is the weighted partial-sum process $\\rho(t/m)\\sqrt{m}\\,\\hat\\sigma_t^{-1}\\sum_{j=1}^t (X_j-\\mu_X)$ and its weak limit $Z_\\rho=\\sup_{y>0}|\\rho(y)W(y)|$. The interval width is $b_t(m;\\rho)=\\sqrt{m/t}\\,\\rho(t/m)$, so $\\rho$ directly shapes the boundary curve and can be chosen flexibly; the constant $c_\\alpha(\\rho)$ is the $(1-\\alpha)$-quantile of $Z_\\rho$. The proof works by splitting the time axis into early, middle, and late parts: (A1) and a generalized Hájek--Rényi inequality control $\\rho(s)$ near $s=0$, a functional central limit theorem controls the middle, and (A2) together with an extension of the Hájek--Rényi inequality to unbounded domains controls $s\\to\\infty$. For the recommended family, the change of variable $x=y/(1-y)$ turns the infinite-horizon Wiener supremum into the finite Brownian-bridge supremum in Proposition 4.3, which is the identity that makes the quantiles numerically tractable.","core_discovery":"The paper's central claim is Theorem 2.2: under the assumption that $(X_t)$ is iid with finite variance and a strongly consistent variance estimator, for any weight function $\\rho$ satisfying (A1) and (A2), the sequence defined by (1) is a sharp asymptotic $(1-\\alpha)$-confidence sequence in the sense of Definition 2.1, with $c_\\alpha(\\rho)$ equal to the $(1-\\alpha)$-quantile of $\\sup_{y>0} |\\rho(y)W(y)|$. The underlying limit theorem (Theorem 4.1) shows that $\\sup_{t\\ge \\ell_m} |\\rho(t/m)\\sqrt{m}\\,\\hat\\sigma_t^{-1}\\sum_{j=1}^t (X_j-\\mu_X)|$ converges in distribution to $\\sup_{y>0}|\\rho(y)W(y)|$ whenever $\\ell_m/m\\to 0$, using a functional central limit theorem for the middle time range and generalized Hájek--Rényi inequalities for the tails. For the concrete weight family $\\rho(s)=(1+s)^{\\gamma_1+\\gamma_2-1}/s^{\\gamma_1}$, $0\\le\\gamma_1,\\gamma_2<1/2$, Proposition 4.3 identifies this limit with $\\sup_{0\\le x\\le 1}|B(x)|/(x^{\\gamma_1}(1-x)^{\\gamma_2})$, where $B$ is a Brownian bridge, so the quantiles can be computed on a finite interval. The paper also shows the intervals are dual to sequential tests whose probability of ever rejecting under the null tends to $\\alpha$ and whose power under a fixed alternative tends to 1 when $\\rho$ decays slowly enough.","pith_inferences":["For a weight function with finite endpoint $e_\\rho$, the displayed interval has zero width for $t> m\\,e_\\rho$, so the theorem's 'for all $t\\in\\mathbb{N}$' statement really requires the convention that no more than $\\lfloor m e_\\rho\\rfloor$ observations are collected, or equivalently that $C_t=\\mathbb{R}$ afterwards; the paper notes this convention informally but does not put it into Definition 2.1","The Brownian-bridge representation means the critical constants $c_\\alpha(\\gamma_1,\\gamma_2)$ depend only on $\\alpha$ and the two tuning parameters, so they can be tabulated once and reused for every $m$; this would make the method easy to deploy in practice.","The same weighted-supremum limit should extend to other parameters that are functionals of partial sums (regression coefficients, U-statistics, and so on) whenever a functional central limit theorem and Hájek--Rényi-type bounds are available, so the construction is likely to generalize beyond the location model."],"forward_implications":["Choosing $\\gamma_2>0$ in the recommended family makes the interval half-width shrink to zero as $t\\to\\infty$, which is exactly what gives the sequential test power 1 under any fixed alternative (Theorem 3.2).","Because the coverage statement holds uniformly over all $t$, the intervals can be monitored continuously without any multiple-testing penalty; this is the anytime-valid property promised by a confidence sequence.","The dual tests control the family-wise error rate level $\\alpha$ when applied simultaneously to a hierarchy of one-sided hypotheses $\\mu_1<\\cdots<\\mu_k$ (Corollary 3.3), so the method supports multiple comparisons over candidate means.","The nonparametric nature means the method applies to any stationary time series after replacing the variance by the long-run variance, per Remark 4.2, not just to iid data."],"supporting_citations":[{"why":"Defines sharp asymptotic confidence sequences and provides the earlier boundary-constant construction whose shape this paper generalizes.","marker":"Waudby-Smith et al. (2024)"},{"why":"Supplies the classical maximal inequality used to control the early-time tail in the proof of Theorem 4.1.","marker":"Hájek and Rényi (1955)"},{"why":"Extends the Hájek--Rényi inequality to unbounded time domains, controlling the late-time tail above t > mV.","marker":"Frank (1966)"},{"why":"Provides the law of the iterated logarithm for Wiener processes and the Brownian-bridge transformation used in Proposition 4.3.","marker":"Csörgő and Révész (1981)"},{"why":"Its Lemma B.2 completes the weak-convergence argument that assembles the three time regimes in the proof of Theorem 4.1.","marker":"Stöhr (2019)"},{"why":"Gives an adaptive quantile computation for Brownian-bridge suprema, making the critical constant c_alpha(rho) numerically available for the proposed family.","marker":"Franke et al. (2022)"}],"fun_headline_variants":["Flexible weights sharpen anytime-valid confidence sequences","New weight class yields sharp time-uniform confidence bounds","Sharp confidence sequences from flexible weight functions","Weight-based sequences achieve sharp anytime-valid coverage","Anytime-valid confidence sequences refined by weight choices"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the asymptotic limit theorem in Theorem 4.1 is uniform over all $t\\in\\mathbb{N}$; for weight functions with a finite endpoint $e_\\rho$ this requires the convention that at most $\\lfloor m e_\\rho\\rfloor$ observations are collected, because the interval formula $\\sqrt{m/t}\\,\\rho(t/m)$ gives zero width for $t>m e_\\rho$ and the paper does not encode that convention into Definition 2.1 or Theorem 2.2.","fun_headline_variants_meta":{"raw":{"variants":["Flexible weights sharpen anytime-valid confidence sequences","New weight class yields sharp time-uniform confidence bounds","Sharp confidence sequences from flexible weight functions","Weight-based sequences achieve sharp anytime-valid coverage","Anytime-valid confidence sequences refined by weight choices"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000432,"raw_usage":{"total_tokens":2222,"prompt_tokens":981,"completion_tokens":1241,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":597,"completion_tokens_details":{"reasoning_tokens":1180}},"tokens_in":597,"tokens_out":1241,"duration_ms":9393,"temperature":1.0,"reasoning_tokens":1180,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T18:19:56.951762+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $\\rho(s)=\\mathbb{1}_{(0,1]}(s)$, so $e_\\rho=1$, choose $m=100$, $\\alpha=0.05$, and simulate iid standard normal data. Compute $C_t(100;0.05)=\\hat\\mu_t\\pm \\hat\\sigma_t c_{0.05}(\\rho)\\sqrt{100/t}$ for $t=1,2,\\ldots$ and check whether $\\mu=0$ lies in all intervals. Since for $t>100$ the interval has zero width and $\\hat\\mu_t\\neq 0$ almost surely, the simultaneous coverage will be 0, directly contradicting Theorem 2.2 unless $C_t$ is explicitly set to $\\mathbb{R}$ after $t=100$; repeating the simulation with that convention should yield coverage near 0.95.","supporting_citations":[],"review_version":1}