REVIEW 3 cited by
Inference with the Upper Confidence Bound Algorithm
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
In this paper, we discuss the asymptotic behavior of the Upper Confidence Bound (UCB) algorithm in the context of multiarmed bandit problems and discuss its implication in downstream inferential tasks. While inferential tasks become challenging when data is collected in a sequential manner, we argue that this problem can be alleviated when the sequential algorithm at hand satisfies certain stability property. This notion of stability is motivated from the seminal work of Lai and Wei (1982). Our first main result shows that such a stability property is always satisfied for the UCB algorithm, and as a result the sample means for each arm are asymptotically normal. Next, we examine the stability properties of the UCB algorithm when the number of arms $K$ is allowed to grow with the number of arm pulls $T$. We show that in such a case the arms are stable when $\frac{\log K}{\log T} \rightarrow 0$, and the number of near-optimal arms are large.
Forward citations
Cited by 3 Pith papers
-
Simulation-Based Inference for Adaptive Experiments
Simulation with optimism resimulates an adaptive experiment under the null with positively biased nuisance means, yielding asymptotically valid tests and narrower confidence intervals after bandit designs.
-
Stabilizing Bandits using Regularization: Precise Regret and A Quantitative Central Limit Theorem
Log-barrier regularized stochastic mirror descent yields Lai–Wei stable bandit sampling, valid Wald intervals, near-optimal regret up to logs, and asymptotic normality under o(√T) corruption.
-
Asymptotic Theory and Sequential Testing for Adaptive Bandits
An urn-based bandit allocation yields reward estimators whose functional central limit theorem reduces to standard Brownian motion after an information-time transformation, so classical group-sequential boundaries rem...
Discussion (0). Continue with ORCID to comment.