Pith. sign in

REVIEW 3 cited by

Inference with the Upper Confidence Bound Algorithm

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04595 v1 pith:CFH7U4E4 submitted 2024-08-08 stat.ML cs.AIcs.LGcs.SYeess.SYmath.STstat.TH

classification stat.MLcs.AIcs.LGcs.SYeess.SYmath.STstat.TH
keywords algorithmstabilitywhenarmsnumberboundconfidencediscuss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

In this paper, we discuss the asymptotic behavior of the Upper Confidence Bound (UCB) algorithm in the context of multiarmed bandit problems and discuss its implication in downstream inferential tasks. While inferential tasks become challenging when data is collected in a sequential manner, we argue that this problem can be alleviated when the sequential algorithm at hand satisfies certain stability property. This notion of stability is motivated from the seminal work of Lai and Wei (1982). Our first main result shows that such a stability property is always satisfied for the UCB algorithm, and as a result the sample means for each arm are asymptotically normal. Next, we examine the stability properties of the UCB algorithm when the number of arms $K$ is allowed to grow with the number of arm pulls $T$. We show that in such a case the arms are stable when $\frac{\log K}{\log T} \rightarrow 0$, and the number of near-optimal arms are large.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Simulation-Based Inference for Adaptive Experiments

    stat.ME 2025-06 conditional novelty 7.0 of 10

    Simulation with optimism resimulates an adaptive experiment under the null with positively biased nuisance means, yielding asymptotically valid tests and narrower confidence intervals after bandit designs.

  2. Stabilizing Bandits using Regularization: Precise Regret and A Quantitative Central Limit Theorem

    stat.ML 2026-03 conditional novelty 6.0 of 10

    Log-barrier regularized stochastic mirror descent yields Lai–Wei stable bandit sampling, valid Wald intervals, near-optimal regret up to logs, and asymptotic normality under o(√T) corruption.

  3. Asymptotic Theory and Sequential Testing for Adaptive Bandits

    stat.ME 2026-02 conditional novelty 6.0 of 10

    An urn-based bandit allocation yields reward estimators whose functional central limit theorem reduces to standard Brownian motion after an information-time transformation, so classical group-sequential boundaries rem...

Pith tools