Pith. sign in

REVIEW 3 major objections 3 minor 21 references

A new and flexible class of sharp asymptotic time-uniform confidence sequences

T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A flexible family of boundary functions yields sharp asymptotic confidence sequences that hold uniformly over all stopping times.

desk verdict Promising new boundary class for asymptotic confidence sequences, but the printed interval width doesn't invert the limit theorem, so the sharpness claim doesn't hold as stated. read the letter →

arxiv 2502.10380 v2 pith:H4VS5ECZ submitted 2025-02-14 math.ST stat.MEstat.MLstat.TH

classification math.STstat.MEstat.MLstat.TH MSC 62G1062G1562G2062L10
keywords Anytime-validinferenceTime-uniformconfidencesequenceNonparametricSequentialtestingSharpasymptoticWeightedWienerprocessHájek–RényiinequalityBrownianbridge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper constructs a new family of anytime-valid confidence intervals for the mean of independent, identically distributed data: intervals $C_t(m;\alpha) = \hat\mu_t \pm \hat\sigma_t \cdot c_\alpha(\rho) \cdot \sqrt{m/t}\,\rho(t/m)$ that are guaranteed to contain the true mean at every time $t$ simultaneously, with probability $1-\alpha$ in the limit as the initial sample size $m$ grows. The shape of the boundary is chosen by a weight function $\rho$ that only needs to satisfy mild growth conditions near zero and infinity, and the critical constant $c_\alpha(\rho)$ comes from the distribution of a supremum of a weighted Wiener process. The main theorem states that this sequence is a sharp asymptotic $(1-\alpha)$-confidence sequence, meaning the simultaneous coverage probability converges to exactly $1-\alpha$, not just at least. This matters because the existing nonparametric confidence sequences had a fixed boundary shape, whereas here the practitioner can spend the error budget unevenly over early and late times. The same construction dualizes to sequential tests with asymptotic level $\alpha$, including one-sided hierarchical tests whose family-wise error rate is controlled.

What carries the argument

The central mechanism is the weighted partial-sum process $\rho(t/m)\sqrt{m}\,\hat\sigma_t^{-1}\sum_{j=1}^t (X_j-\mu_X)$ and its weak limit $Z_\rho=\sup_{y>0}|\rho(y)W(y)|$. The interval width is $b_t(m;\rho)=\sqrt{m/t}\,\rho(t/m)$, so $\rho$ directly shapes the boundary curve and can be chosen flexibly; the constant $c_\alpha(\rho)$ is the $(1-\alpha)$-quantile of $Z_\rho$. The proof works by splitting the time axis into early, middle, and late parts: (A1) and a generalized Hájek--Rényi inequality control $\rho(s)$ near $s=0$, a functional central limit theorem controls the middle, and (A2) together with an extension of the Hájek--Rényi inequality to unbounded domains controls $s\to\infty$. For the recommended family, the change of variable $x=y/(1-y)$ turns the infinite-horizon Wiener supremum into the finite Brownian-bridge supremum in Proposition 4.3, which is the identity that makes the quantiles numerically tractable.

What would settle it

Take $\rho(s)=\mathbb{1}_{(0,1]}(s)$, so $e_\rho=1$, choose $m=100$, $\alpha=0.05$, and simulate iid standard normal data. Compute $C_t(100;0.05)=\hat\mu_t\pm \hat\sigma_t c_{0.05}(\rho)\sqrt{100/t}$ for $t=1,2,\ldots$ and check whether $\mu=0$ lies in all intervals. Since for $t>100$ the interval has zero width and $\hat\mu_t\neq 0$ almost surely, the simultaneous coverage will be 0, directly contradicting Theorem 2.2 unless $C_t$ is explicitly set to $\mathbb{R}$ after $t=100$; repeating the simulation with that convention should yield coverage near 0.95.

Watch

Extended reading notes

Core claim

The paper's central claim is Theorem 2.2: under the assumption that $(X_t)$ is iid with finite variance and a strongly consistent variance estimator, for any weight function $\rho$ satisfying (A1) and (A2), the sequence defined by (1) is a sharp asymptotic $(1-\alpha)$-confidence sequence in the sense of Definition 2.1, with $c_\alpha(\rho)$ equal to the $(1-\alpha)$-quantile of $\sup_{y>0} |\rho(y)W(y)|$. The underlying limit theorem (Theorem 4.1) shows that $\sup_{t\ge \ell_m} |\rho(t/m)\sqrt{m}\,\hat\sigma_t^{-1}\sum_{j=1}^t (X_j-\mu_X)|$ converges in distribution to $\sup_{y>0}|\rho(y)W(y)|$ whenever $\ell_m/m\to 0$, using a functional central limit theorem for the middle time range and generalized Hájek--Rényi inequalities for the tails. For the concrete weight family $\rho(s)=(1+s)^{\gamma_1+\gamma_2-1}/s^{\gamma_1}$, $0\le\gamma_1,\gamma_2<1/2$, Proposition 4.3 identifies this limit with $\sup_{0\le x\le 1}|B(x)|/(x^{\gamma_1}(1-x)^{\gamma_2})$, where $B$ is a Brownian bridge, so the quantiles can be computed on a finite interval. The paper also shows the intervals are dual to sequential tests whose probability of ever rejecting under the null tends to $\alpha$ and whose power under a fixed alternative tends to 1 when $\rho$ decays slowly enough.

Load-bearing premise

The load-bearing premise is that the asymptotic limit theorem in Theorem 4.1 is uniform over all $t\in\mathbb{N}$; for weight functions with a finite endpoint $e_\rho$ this requires the convention that at most $\lfloor m e_\rho\rfloor$ observations are collected, because the interval formula $\sqrt{m/t}\,\rho(t/m)$ gives zero width for $t>m e_\rho$ and the paper does not encode that convention into Definition 2.1 or Theorem 2.2.

Editorial extensions

If this is right

  • Choosing $\gamma_2>0$ in the recommended family makes the interval half-width shrink to zero as $t\to\infty$, which is exactly what gives the sequential test power 1 under any fixed alternative (Theorem 3.2).
  • Because the coverage statement holds uniformly over all $t$, the intervals can be monitored continuously without any multiple-testing penalty; this is the anytime-valid property promised by a confidence sequence.
  • The dual tests control the family-wise error rate level $\alpha$ when applied simultaneously to a hierarchy of one-sided hypotheses $\mu_1<\cdots<\mu_k$ (Corollary 3.3), so the method supports multiple comparisons over candidate means.
  • The nonparametric nature means the method applies to any stationary time series after replacing the variance by the long-run variance, per Remark 4.2, not just to iid data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • For a weight function with finite endpoint $e_\rho$, the displayed interval has zero width for $t> m\,e_\rho$, so the theorem's 'for all $t\in\mathbb{N}$' statement really requires the convention that no more than $\lfloor m e_\rho\rfloor$ observations are collected, or equivalently that $C_t=\mathbb{R}$ afterwards; the paper notes this convention informally but does not put it into Definition 2.1
  • The Brownian-bridge representation means the critical constants $c_\alpha(\gamma_1,\gamma_2)$ depend only on $\alpha$ and the two tuning parameters, so they can be tabulated once and reused for every $m$; this would make the method easy to deploy in practice.
  • The same weighted-supremum limit should extend to other parameters that are functionals of partial sums (regression coefficients, U-statistics, and so on) whenever a functional central limit theorem and Hájek--Rényi-type bounds are available, so the construction is likely to generalize beyond the location model.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper proposes a new class of asymptotic time-uniform confidence sequences for the mean of iid observations. The sequence is defined by C_t(m;α) = μ̂_t ± σ̂_t c_α(ρ) √(m/t) ρ(t/m), where ρ is a user-chosen weight function satisfying growth conditions A1 and A2, and c_α(ρ) is the (1−α)-quantile of sup_{y>0} |ρ(y)W(y)| for a standard Wiener process. The main theorem (Theorem 2.2) claims that this sequence is asymptotically sharp, i.e., lim_{m→∞} P(μ ∈ C_t for all t∈N) = 1−α. The paper also presents a dual sequential testing formulation (Theorem 3.1), a stopping guarantee under alternatives (Theorem 3.2), and a limit theorem (Theorem 4.1) that is stated to be the underlying asymptotic tool. A special family of weight functions is shown to reduce the critical constant to a Brownian-bridge supremum (Proposition 4.3).

Significance. If the central claim were correct, the paper would offer a flexible and easily implementable class of anytime-valid confidence sequences, with the critical constant obtained from a standard Wiener-process or Brownian-bridge supremum. The limit theorem (Theorem 4.1) is plausible and its proof sketch is coherent for infinite-endpoint weight functions, and the connection to change-point monitoring is a valuable perspective. However, the advertised statistical construction does not follow from the limit theorem, and the stated confidence sequence does not achieve the claimed coverage probability. The error is load-bearing, so the paper does not establish its main contribution.

major comments (3)
  1. [Section 2, Eq. (1) and Theorem 2.2] The claimed inversion of Theorem 4.1 is algebraically incorrect. The coverage event μ ∈ C_t(m;α) is equivalent to |S_t|/(σ̂_t √m) ≤ c_α(ρ) √t ρ(t/m). For t = ⌊my⌋, the left-hand side converges weakly to |W(y)|, while the right-hand side is c_α(ρ) √m √y ρ(y), which diverges for any y with ρ(y)>0. Consequently, the probability that any false rejection occurs tends to 0, not to α, so the sequence (1) is not sharp; its coverage tends to 1. The correct inversion of the limit sup_y |ρ(y)W(y)| would require a half-width proportional to √(m/t) / (√t ρ(t/m)), not √(m/t) ρ(t/m). This invalidates the central claim of Theorem 2.2 and its dual Theorem 3.1.
  2. [Section 2, Definition 2.1 and Theorem 2.2 (finite endpoint e_ρ)] For weight functions with finite e_ρ, ρ(s)=0 for s>e_ρ, so b_t(m;ρ)=0 for t>m e_ρ. Then C_t collapses to the degenerate interval {μ̂_t}, which has probability zero of containing a fixed μ under the model. The sentence 'at most ⌊m e_ρ⌋ data points are collected' is not encoded in Definition 2.1, which requires coverage for all t∈N, nor in the statement of Theorem 2.2. Without an explicit convention such as C_t = R after the endpoint, the all-t coverage claim is false for finite-endpoint weight functions.
  3. [Section 4, proof of Theorem 4.1] The proof of Theorem 4.1 appears to handle only the case where the weight function decays sufficiently at infinity. For a finite endpoint e_ρ with ρ(e_ρ)>0, the tail bound in Eq. (8) requires sup_{s>V} s^{1−γ2} ρ(s) → 0 as V increases, which fails if ρ has a positive limit at e_ρ. The theorem is therefore not established for all ρ allowed by Assumption B; it is only clearly valid for infinite-endpoint ρ satisfying the stated decay.
minor comments (3)
  1. [Section 4, Eq. (4)] The normalization in Theorem 4.1 is 1/√m times the partial sum, which is the correct Donsker scaling; the skeptical reading that the theorem contains √m times the sum is not supported by the manuscript text.
  2. [Section 2, paragraph after Eq. (1)] The paper would benefit from an explicit statement that for finite-endpoint ρ the procedure stops at ⌊m e_ρ⌋ and the confidence sequence is defined as the whole real line afterwards; currently this is only mentioned informally.
  3. [References] There are minor inconsistencies in author name spellings, for example 'Stoehr' vs 'Stöhr', and the use of a PhD thesis (Stöhr 2019) for a key lemma; the authors should provide a more accessible reference if possible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the critical constants are external Wiener-process quantiles, the proof uses external invariance principles and Hájek–Rényi inequalities, and the self-citations are discussion-level pointers rather than load-bearing inputs.

full rationale

The paper's central claim is that the intervals C_t(m; alpha) in Eq. (1) form a sharp asymptotic confidence sequence when c_alpha(rho) is the (1-alpha)-quantile of sup|rho W|. This is not circular: the quantile is not fitted to data, and the coverage statement is meant to follow from the functional limit theorem in Theorem 4.1, whose proof invokes the classical functional central limit theorem, Hájek–Rényi inequalities, and the law of the iterated logarithm from standard external references (Hájek and Rényi 1955, Frank 1966, Csörgő and Révész 1981). The cited Stöhr (2019, Lemma B.2) is a technical continuous-mapping-type lemma used only to assemble the limiting argument, not an assumption of the target result. Self-citations to Aue and Kirch (2024), Kirch and Tadjuidje Kamgaing (2015), Kirch and Stoehr (2022), and Hlávka et al. (2012) appear in the discussion and outlook or in Remark 4.2 as pointers to known invariance principles and monitoring methodology; they do not carry the proof of Theorem 2.2. The finite-endpoint issue noted in the reader's take is a convention-level correctness gap rather than a circularity: if t exceeds m*e_rho, the printed interval degenerates, so the claim 'for all t in N' needs the supplementary convention that C_t = R after the endpoint. That is a fix to the statement, not a reduction of the output to an input. Likewise, the possible normalization mismatch between Eq. (1) and Theorem 4.1 flagged by the skeptic is a mathematical-correctness concern; if real, it would make the derivation false rather than tautological. Under the circularity rubric, no step equates the claimed result to an input by construction, so the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper does not introduce new particles, forces, dimensions, or physical entities. It introduces a family of weight functions and a new proof strategy, with user-selected shape parameters gamma1 and gamma2. The central claim rests on standard iid assumptions, stated growth conditions on rho, and classical stochastic-process limit theorems. The finite-endpoint coverage convention is an unstated assumption that should be made explicit.

free parameters (3)
  • gamma1 = user-specified, 0 <= gamma1 < 1/2
    Controls the shape of the boundary function near s=0 through condition A1. It is not fitted to data and the central theorem holds for every allowed value.
  • gamma2 = user-specified, 0 <= gamma2 < 1/2
    Controls the tail decay of the boundary function through condition A2; gamma2 > 0 makes the interval width shrink to zero and ensures the sequential test stops under alternatives.
  • l_m = user-specified, e.g. 1, log(m), sqrt(m), with l_m/m -> 0
    Starting time for the supremum in Theorem 4.1. It does not affect the limit distribution but changes finite-sample behavior and is chosen by hand.
assumptions (5)
  • domain assumption Assumption A: X_t are iid with finite mean and finite positive variance, and the variance estimator sigmahat_t^2 is strongly consistent and positive a.s.
    This is the statistical model stated before Definition 2.1 and used throughout Theorems 2.2, 3.1 and 4.1.
  • ad hoc to paper Assumption B: rho is continuous on its positivity region and satisfies the growth conditions A1 near 0 and A2 near infinity.
    These conditions are introduced specifically to make the weighted supremum finite and to allow the tail arguments in the proof of Theorem 4.1.
  • standard math The partial sum process satisfies a functional central limit theorem.
    Used in equation (4) of the proof to obtain convergence on compact intervals; for iid data it is the classical Donsker theorem.
  • standard math The generalized Hajek-Renyi inequalities in equations (5) and (7) hold for the partial sums.
    These inequalities control the supremum over small and large time scales and are proved in the paper for iid data using Hajek-Renyi and Frank's extension.
  • standard math Wiener-process law of the iterated logarithm and Brownian-bridge time-inversion identities.
    Used in equations (9), (10) and Proposition 4.3 via Csorgo and Revesz results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A new and flexible class of sharp asymptotic time-uniform confidence sequences." pith.science (2026). https://pith.science/paper/H4VS5ECZ

@misc{pith2026250210380,
  author       = {Pith},
  title        = {Pith review of: A new and flexible class of sharp asymptotic time-uniform confidence sequences},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H4VS5ECZ}},
  note         = {Machine review of arXiv:2502.10380}
}
read the original abstract

Confidence sequences are anytime-valid analogues of classical confidence intervals that do not suffer from multiplicity issues under optional continuation of the data collection. As in classical statistics, asymptotic confidence sequences are a nonparametric tool showing under which high-level assumptions asymptotic coverage is achieved so that they also give a certain robustness guarantee against distributional deviations. In this paper, we propose a new flexible class of confidence sequences yielding sharp asymptotic time-uniform confidence sequences under mild assumptions. Furthermore, we highlight the connection to corresponding sequential testing problems and detail the underlying limit theorem.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 11 canonical work pages

  1. [8]

    Safe testing. J. R. Stat. Soc. Ser. B Stat. Methodol. qkae011. http://dx.doi.org/10.1093/jrsssb/qkae011. Hájek, J., Rényi, A.,

  2. [12]

    Sequential change point tests based on U-statistics. Scand. J. Stat. 49 (3), 1184–1214. http://dx.doi.org/10.1111/sjos.12558. Kirch, C., Tadjuidje Kamgaing, J.,

  3. [18]

    Boundary crossing probabilities for the Wiener process and sample sums. Ann. Math. Stat. 41 (5), 1410–1429. http: //dx.doi.org/10.1214/aoms/1177696787. Rom, D.M., Holland, B.,

  4. [21]

    Time-uniform central limit theory and asymptotic confidence sequences. Ann. Statist. 52 (6), 2613–2640. http://dx.doi.org/10.1214/24-AOS2408. Statistics and Probability Letters 226 (2025) 110462 6

  5. [284]

    Frank, O.,

    http://dx.doi.org/10.1016/C2013-0-10553-3, 666546. Frank, O.,

  6. [1955]

    Acta Math

    Generalization of an inequality of Kolmogorov. Acta Math. Acad. Sci. Hung. 6 (3), 281–283. http://dx.doi.org/10.1007/BF02024392. Hlávka, Z., Hušková, M., Kirch, C., Meintanis, S.G.,

  7. [1966]

    Generalization of an inequality of Hájek and Rényi. Scand. Actuar. J. 1966 (1–2), 85–89. http://dx.doi.org/10.1080/03461238.1966.10405700. Franke, J., Hefter, M., Herzwurm, A., Ritter, K., Schwaar, S.,

  8. [1970]

    Statistical methods related to the law of the iterated logarithm. Ann. Math. Stat. 41 (5), 1397–1409. http://dx.doi.org/10.1214/aoms/ 1177696786. Robbins, H., Siegmund, D.,

Show all 21 references
  1. [1976]

    On confidence sequences. Ann. Statist. 4 (2), 265–280. http://dx.doi.org/10.1214/aos/1176343406. Lan, K.K.G., DeMets, D.L.,

  2. [1983]

    Biometrika 70 (3), 659–663

    Discrete sequential boundaries for clinical trials. Biometrika 70 (3), 659–663. http://dx.doi.org/10.2307/2336502. Lei, L., Fithian, W.,

  3. [1995]

    A new closed multiple testing procedure for hierarchical families of hypotheses. J. Statist. Plann. Inference 46 (3), 265–275. http://dx.doi.org/10.1016/0378-3758(94)00116-D. Stöhr, C.,

  4. [1996]

    Econometrica 64 (5), 1045–1065

    Monitoring structural change. Econometrica 64 (5), 1045–1065. http://dx.doi.org/10.2307/2171955. Csörgő, M., Révész, P.,

  5. [2004]

    Monitoring changes in linear models. J. Statist. Plann. Inference 126 (1), 225–251. http: //dx.doi.org/10.1016/j.jspi.2003.07.014. Hušková, M., Koubková, A.,

  6. [2006]

    Change-point monitoring in linear models. Econom. J. 9 (3), 373–403. http://dx.doi.org/10.1111/j.1368- 423X.2006.00190.x. Aue, A., Kirch, C.,

  7. [2012]

    TEST 21 (4), 605–634

    Monitoring changes in the error distribution of autoregressive models based on Fourier methods. TEST 21 (4), 605–634. http://dx.doi.org/10.1007/s11749-011-0265-z. Horváth, L., Hušková, M., Kokoszka, P., Steinebach, J.,

  8. [2015]

    On the use of estimating functions in monitoring time series for change points. J. Statist. Plann. Inference 161, 25–49. http://dx.doi.org/10.1016/j.jspi.2014.12.009. Koubková, A.,

  9. [2019]

    Sequential Change Point Procedures Based on U-Statistics and the Detection of Covariance Changes in Functional Data (Ph.D. thesis). Otto-von-Guericke-Universität Magdeburg, http://dx.doi.org/10.25673/13826. Waudby-Smith, I., Arbour, D., Sinha, R., Kennedy, E.H., Ramdas, A.,

  10. [2021]

    http: //dx.doi.org/10.48550/arXiv.2101.07380, arXiv preprint arXiv:2101.07380

    Sequential causal inference in a single world of connected units. http: //dx.doi.org/10.48550/arXiv.2101.07380, arXiv preprint arXiv:2101.07380. Chochola, O.,

  11. [2022]

    Adaptive quantile computation for Brownian bridge in change-point analysis. Comput. Statist. Data Anal. 167, 107375. http://dx.doi.org/10.1016/j.csda.2021.107375. Grünwald, P., de Heide, R., Koolen, W.,

  12. [2023]

    Game-theoretic statistics and safe anytime-valid inference. Statist. Sci. 38 (4), 576–601. http://dx.doi.org/ 10.1214/23-STS894. Robbins, H.,

  13. [2024]

    http://dx.doi.org/10.48550/arXiv.2212.14411, arXiv Preprint arXiv:2212.14411v5

    Near-optimal non-parametric sequential tests and confidence sequences with possibly dependent observations. http://dx.doi.org/10.48550/arXiv.2212.14411, arXiv Preprint arXiv:2212.14411v5. Bibaut, A., Petersen, M., Vlassis, N., Dimakopoulou, M., van der Laan, M.,

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.