Pith. sign in

REVIEW 3 minor 57 references

Beyond First-order Asymptotics in Sequential Mean Testing

T0 review · 0 major / 3 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read The stopping time of the KL_inf sequential test for bounded means, centered and scaled by sqrt(log(1/alpha)), converges in distribution to a Gaussian with explicit variance.

desk verdict This paper derives a CLT for the KL_inf statistic and transfers it to obtain an explicit Gaussian limit for the stopping time of the sequential test. read the letter →

arxiv 2606.04520 v1 pith:27W4PAOQ submitted 2026-06-03 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH
keywords sequentialhypothesistestingmeanboundeddistributionsKL_infcentrallimittheoremstoppingtimesasymptoticspower-onetests
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Sequential testing of a mean for bounded random variables needs a level-alpha power-one procedure whose expected stopping time is as small as possible. The KL_inf-based test is already known to match the information-theoretic lower bound exactly in the small-alpha regime. The paper first derives a central limit theorem for the KL_inf statistic that describes its fluctuations around the deterministic limit, then transfers this result to obtain a second-order normal limit for the stopping time itself. A reader would care because the result supplies an explicit variance that quantifies how much the duration of the optimal test varies around its leading-term expectation.

What carries the argument

The KL_inf statistic, whose own central limit theorem is established first and then mapped to the stopping time through the first-order optimality relation.

What would settle it

Monte Carlo simulations on a bounded distribution in which the empirical distribution of the centered and scaled stopping times fails to approach the predicted normal as alpha tends to zero.

Watch

Extended reading notes

Core claim

We prove a central limit theorem for the stopping time of the KL_inf-based sequential test. After appropriate centering, the stopping time scaled by sqrt(log(1/alpha)) converges in distribution to a normal random variable whose variance is given explicitly. The proof proceeds by establishing a novel CLT for the KL_inf statistic that characterizes its fluctuations around its deterministic limit, then using this intermediate result to obtain the limit for the stopping time.

Load-bearing premise

The KL_inf statistic itself obeys a central limit theorem under the bounded-distribution setting.

Editorial extensions

If this is right

  • The duration of the asymptotically optimal test admits an explicit second-order normal approximation.
  • The variance in the limit depends only on the underlying distribution and the information quantities already appearing in the first-order analysis.
  • The same two-step argument (CLT for the statistic, then transfer to stopping time) applies uniformly to any bounded distribution.
  • Numerical experiments confirm that the Gaussian approximation becomes accurate for small alpha.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The explicit variance could be used to construct refined thresholds that achieve closer-to-nominal error rates in moderate sample sizes.
  • Analogous second-order CLTs may exist for other information-based sequential tests whose first-order optimality is already known.
  • The variance expression offers a way to compare the variability of different asymptotically optimal procedures on the same problem.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The paper studies a KL_inf-based sequential test for the mean of bounded distributions that attains the information-theoretic lower bound on expected stopping time with exact constants as alpha to 0. It establishes a novel CLT for the KL_inf statistic around its deterministic limit and transfers the result to show that the stopping time, after appropriate centering and scaling by sqrt(log(1/alpha)), converges in distribution to a Gaussian with explicit variance. Numerical experiments are included to support the theory.

Significance. If the CLTs hold under the stated conditions, the work supplies a second-order characterization of an asymptotically optimal sequential test, including an explicit limiting variance. This is a meaningful advance beyond first-order asymptotics in sequential analysis. The two-step strategy (CLT for the statistic followed by continuous mapping to the stopping time) and the provision of numerical corroboration are strengths.

minor comments (3)
  1. [Abstract and Section 3] The abstract states that proofs exist for the two CLTs but does not list the precise technical conditions (e.g., moment assumptions beyond boundedness or handling of the boundary case mu = 0). These should be stated explicitly in the main text, preferably in the statement of the main theorems.
  2. [Theorem 4.1] The explicit variance in the limiting Gaussian for the stopping time is a key contribution; its derivation should be highlighted with a dedicated display equation rather than left implicit in the proof.
  3. [Section 5] In the numerical experiments, the number of Monte Carlo replications and the specific bounded distributions used (e.g., support bounds) should be reported so that the observed agreement with the CLT can be assessed for variability.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive assessment of our manuscript, the accurate summary of its contributions, and the recommendation for minor revision. The report correctly identifies the two-step strategy (CLT for the KL_inf statistic followed by continuous mapping to the stopping time) and the value of the explicit limiting variance as strengths.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected

full rationale

The derivation consists of two explicit mathematical steps: establishing a novel CLT for the KL_inf statistic around its deterministic limit, followed by transferring the result to the stopping time via continuous mapping or inversion. The abstract and described structure contain no self-definitional equations, no fitted parameters renamed as predictions, and no load-bearing self-citations that reduce the central claim to prior author work by construction. The argument relies on standard tightness and Lipschitz properties for bounded random variables, which are independent of the target second-order result, rendering the chain self-contained.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the standard setup of power-one sequential tests for bounded random variables and on the first-order optimality of the KL_inf test; no free parameters or new entities are introduced.

assumptions (2)
  • domain assumption Random variables are bounded
    Stated as the problem setting for the mean-testing framework.
  • domain assumption KL_inf statistic admits a central limit theorem under the given conditions
    This is the novel intermediate result used to reach the stopping-time CLT.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Beyond First-order Asymptotics in Sequential Mean Testing." pith.science (2026). https://pith.science/paper/27W4PAOQ

@misc{pith2026260604520,
  author       = {Pith},
  title        = {Pith review of: Beyond First-order Asymptotics in Sequential Mean Testing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/27W4PAOQ}},
  note         = {Machine review of arXiv:2606.04520}
}
abstract

We revisit the problem of sequentially testing the mean of bounded distributions in a level-$\alpha$ power-one framework. We study a $\mathrm{KL_{inf}}$-based sequential test that is known to attain the information-theoretic lower bound on the expected stopping time with exact constants as $\alpha \to 0$. Going beyond first-order asymptotics, we establish a central limit theorem (CLT) for the stopping time of this test. Our analysis proceeds in two steps. First, we prove a novel CLT for the $\mathrm{KL_{inf}}$ statistic itself, characterizing its fluctuations around its deterministic limit. We then leverage this result to show that the stopping time, centered appropriately and scaled by $\sqrt{\log(1/\alpha)}$, converges in distribution to a Gaussian limit with an explicit variance. This yields a second-order characterization of an asymptotically optimal sequential test for bounded distributions. Finally, we present numerical experiments that corroborate our theoretical findings.

Figures

Figures reproduced from arXiv: 2606.04520 by the authors.

Figure 2
Figure 2. Histogram of the statistic √ n(KLinf(ˆqn, mo) − KLinf(q, mo)) when q ∼ Bernoulli(0.6). The orange curve is the density of N (0, σ2 (q, mo)). defined in (3). We set mo = 0.2 and consider the data￾generating distribution q ∼ Bernoulli(0.6). We consider two values of α: 10−4 , 10−8 . For each confidence level α, we simulate 5000 independent sample paths, compute τα along each path, and form the centered-and-scaled stat… view at source ↗
Figure 1
Figure 1. Histogram of the statistic √ n(KLinf(ˆqn, mo) − KLinf(q, mo)) when q ∼ Beta(3, 2). The orange curve is the density of N (0, σ2 (q, mo)). Experiment 2 (CLT for the stopping time). We now study the asymptotic normality of the stopping rule τα 1 0 1 n = 1000 0.0 0.5 1.0 Density 1 0 1 n = 2000 0.0 0.5 1.0 1.5 n (KLinf(qn, m0) KLinf(q, m0)) (0, 2 (q, m0)) [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 5
Figure 5. Histogram of the statistic p log(1/α)(τα/ log(1/α) − 1/KLinf(ˆq, mo)). The orange curve is the density of N [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Histogram of the statistic p log(1/α)(τα/log(1/α) − 1/KLinf(q, mo)) with α = 10−4 on left and α = 10−8 on right. The orange curve is the density of N [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 4 canonical work pages

  1. [1]

    Langley , title =

    P. Langley , title =. Proceedings of the 17th International Conference on Machine Learning (ICML 2000) , address =. 2000 , pages =

  2. [2]

    Proceedings of Thirty Fourth Conference on Learning Theory , pages =

    Regret Minimization in Heavy-Tailed Bandits , author =. Proceedings of Thirty Fourth Conference on Learning Theory , pages =. 2021 , volume =

  3. [3]

    Operations Research , volume=

    The fragility of optimized bandit algorithms , author=. Operations Research , volume=. 2025 , publisher=

  4. [5]

    The Annals of Statistics , volume=

    The expected sample size of some tests of power one , author=. The Annals of Statistics , volume=. 1974 , publisher=

  5. [6]

    , author=

    Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards. , author=. J. Mach. Learn. Res. , volume=

  6. [7]

    Advances in Neural Information Processing Systems , volume=

    Optimal best-arm identification methods for tail-risk measures , author=. Advances in Neural Information Processing Systems , volume=

  7. [8]

    Algorithmic Learning Theory , pages=

    Optimal -Correct Best-Arm Selection for Heavy-Tailed Distributions , author=. Algorithmic Learning Theory , pages=. 2020 , organization=

  8. [9]

    Advances in Applied Mathematics , volume=

    Optimal adaptive policies for sequential allocation problems , author=. Advances in Applied Mathematics , volume=. 1996 , publisher=

Show all 57 references
  1. [10]

    Advances in Applied Mathematics , volume=

    Asymptotically efficient adaptive allocation rules , author=. Advances in Applied Mathematics , volume=. 1985 , publisher=

  2. [11]

    T. M. Mitchell. The Need for Biases in Learning Generalizations. 1980

  3. [12]

    M. J. Kearns , title =

  4. [13]

    Machine Learning: An Artificial Intelligence Approach, Vol. I. 1983

  5. [14]

    R. O. Duda and P. E. Hart and D. G. Stork. Pattern Classification. 2000

  6. [15]

    Suppressed for Anonymity , author=

  7. [16]

    Newell and P

    A. Newell and P. S. Rosenbloom. Mechanisms of Skill Acquisition and the Law of Practice. Cognitive Skills and Their Acquisition. 1981

  8. [17]

    A. L. Samuel. Some Studies in Machine Learning Using the Game of Checkers. IBM Journal of Research and Development. 1959

  9. [19]

    Breakthroughs in statistics: Foundations and basic theory , pages=

    Sequential tests of statistical hypotheses , author=. Breakthroughs in statistics: Foundations and basic theory , pages=. 1992 , publisher=

  10. [20]

    2004 , publisher=

    Sequential analysis , author=. 2004 , publisher=

  11. [21]

    The Annals of Mathematical Statistics , pages=

    Optimum character of the sequential probability ratio test , author=. The Annals of Mathematical Statistics , pages=. 1948 , publisher=

  12. [22]

    Proceedings of the National Academy of Sciences , volume=

    Iterated logarithm inequalities , author=. Proceedings of the National Academy of Sciences , volume=

  13. [23]

    2013 , publisher=

    Sequential Analysis: Tests and Confidence Intervals , author=. 2013 , publisher=

  14. [24]

    Sequential Design of Experiments , author=. Ann. Math. Statist. , volume=

  15. [25]

    , author=

    An Asymptotically Optimal Bandit Algorithm for Bounded Support Models. , author=. COLT , pages=

  16. [26]

    2009 , publisher=

    Stopped random walks , author=. 2009 , publisher=

  17. [27]

    Calcutta Statistical Association Bulletin , volume=

    Asymptotic Normality of Sequential Stopping Times with Applications: Confidence Intervals for an Exponential Mean , author=. Calcutta Statistical Association Bulletin , volume=. 2020 , publisher=

  18. [28]

    2003 , publisher=

    Applied probability and queues , author=. 2003 , publisher=

  19. [29]

    2023 , school =

    Bandits with Heavy Tails: Algorithms Analysis and Optimality , author =. 2023 , school =

  20. [30]

    Advances in Neural Information Processing Systems , volume=

    Top two algorithms revisited , author=. Advances in Neural Information Processing Systems , volume=

  21. [31]

    Mathematical Proceedings of the Cambridge Philosophical Society , volume=

    Large-sample theory of sequential estimation , author=. Mathematical Proceedings of the Cambridge Philosophical Society , volume=. 1952 , organization=

  22. [32]

    Asymptotically optimal and computationally efficient average treatment effect estimation in A/B testing , author=

  23. [35]

    2017 , publisher=

    Probability and measure , author=. 2017 , publisher=

  24. [36]

    Bandits with Heavy Tails: Algorithms Analysis and Optimality

    Agrawal, S. Bandits with Heavy Tails: Algorithms Analysis and Optimality. PhD thesis, Tata Institute of Fundamental Research, 2023. URL http://hdl.handle.net/10603/478863

  25. [37]

    and Ramdas, A

    Agrawal, S. and Ramdas, A. On stopping times of power-one sequential tests: Tight lower and upper bounds. arXiv preprint arXiv:2504.19952, 2025

  26. [38]

    Optimal -correct best-arm selection for heavy-tailed distributions

    Agrawal, S., Juneja, S., and Glynn, P. Optimal -correct best-arm selection for heavy-tailed distributions. In Algorithmic Learning Theory, pp.\ 61--110. PMLR, 2020

  27. [39]

    K., and Koolen, W

    Agrawal, S., Juneja, S. K., and Koolen, W. M. Regret minimization in heavy-tailed bandits. In Proceedings of Thirty Fourth Conference on Learning Theory, volume 134 of Proceedings of Machine Learning Research, pp.\ 26--62. PMLR, 15--19 Aug 2021 a

  28. [40]

    M., and Juneja, S

    Agrawal, S., Koolen, W. M., and Juneja, S. Optimal best-arm identification methods for tail-risk measures. Advances in Neural Information Processing Systems, 34: 0 25578--25590, 2021 b

  29. [41]

    Anscombe, F. J. Large-sample theory of sequential estimation. In Mathematical Proceedings of the Cambridge Philosophical Society, volume 48, pp.\ 600--607. Cambridge University Press, 1952

  30. [42]

    Applied probability and queues

    Asmussen, S. Applied probability and queues. Springer, 2003

  31. [43]

    Probability and measure

    Billingsley, P. Probability and measure. John Wiley & Sons, 2017

  32. [44]

    Burnetas, A. N. and Katehakis, M. N. Optimal adaptive policies for sequential allocation problems. Advances in Applied Mathematics, 17 0 (2): 0 122--142, 1996

  33. [45]

    Sequential design of experiments

    Chernoff, H. Sequential design of experiments. Ann. Math. Statist., 30 0 (4): 0 755--770, 1959

  34. [46]

    Darling, D. A. and Robbins, H. Iterated logarithm inequalities. Proceedings of the National Academy of Sciences, 57 0 (5): 0 1188--1192, 1967

  35. [47]

    Deep, V., Bassamboo, A., and Juneja, S. K. Asymptotically optimal and computationally efficient average treatment effect estimation in a/b testing. 2024

  36. [48]

    Asymptotic optimality theory of confidence intervals of the mean

    Deep, V., Bassamboo, A., and Juneja, S. Asymptotic optimality theory of confidence intervals of the mean. arXiv preprint arXiv:2501.19126, 2025

  37. [49]

    and Glynn, P

    Fan, L. and Glynn, P. W. The fragility of optimized bandit algorithms. Operations Research, 73 0 (6): 0 3173--3198, 2025

  38. [50]

    Stopped random walks

    Gut, A. Stopped random walks. Springer, 2009

  39. [51]

    and Takemura, A

    Honda, J. and Takemura, A. An asymptotically optimal bandit algorithm for bounded support models. In COLT, pp.\ 67--79, 2010

  40. [52]

    and Takemura, A

    Honda, J. and Takemura, A. Non-asymptotic analysis of a new bandit algorithm for semi-bounded rewards. J. Mach. Learn. Res., 16: 0 3721--3756, 2015

  41. [53]

    Top two algorithms revisited

    Jourdan, M., Degenne, R., Baudry, D., de Heide, R., and Kaufmann, E. Top two algorithms revisited. Advances in Neural Information Processing Systems, 35: 0 26791--26803, 2022

  42. [54]

    Lai, T. L. and Robbins, H. Asymptotically efficient adaptive allocation rules. Advances in Applied Mathematics, 6 0 (1): 0 4--22, 1985

  43. [55]

    Asymptotic normality of sequential stopping times with applications: Confidence intervals for an exponential mean

    Mukhopadhyay, N. Asymptotic normality of sequential stopping times with applications: Confidence intervals for an exponential mean. Calcutta Statistical Association Bulletin, 72 0 (1): 0 17--34, 2020

  44. [56]

    and Agrawal, S

    Panda, S. and Agrawal, S. Regret tail characterization of optimal bandit algorithms with generic rewards. arXiv preprint arXiv:2604.14876, 2026

  45. [57]

    and Siegmund, D

    Robbins, H. and Siegmund, D. The expected sample size of some tests of power one. The Annals of Statistics, 2 0 (3): 0 415--436, 1974

  46. [58]

    Sequential Analysis: Tests and Confidence Intervals

    Siegmund, D. Sequential Analysis: Tests and Confidence Intervals. Springer Science & Business Media, 2013

  47. [59]

    Sequential tests of statistical hypotheses

    Wald, A. Sequential tests of statistical hypotheses. In Breakthroughs in statistics: Foundations and basic theory, pp.\ 256--298. Springer, 1992

  48. [60]

    and Wolfowitz, J

    Wald, A. and Wolfowitz, J. Optimum character of the sequential probability ratio test. The Annals of Mathematical Statistics, pp.\ 326--339, 1948

  49. [61]

    Almost sure null bankruptcy of testing-by-betting strategies

    Wang, H., Agrawal, S., and Ramdas, A. Almost sure null bankruptcy of testing-by-betting strategies. arXiv preprint arXiv:2602.08888, 2026

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.