REVIEW 1 cited by
Upper Confidence Bounds for Combining Stochastic Bandits
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We provide a simple method to combine stochastic bandit algorithms. Our approach is based on a "meta-UCB" procedure that treats each of $N$ individual bandit algorithms as arms in a higher-level $N$-armed bandit problem that we solve with a variant of the classic UCB algorithm. Our final regret depends only on the regret of the base algorithm with the best regret in hindsight. This approach provides an easy and intuitive alternative strategy to the CORRAL algorithm for adversarial bandits, without requiring the stability conditions imposed by CORRAL on the base algorithms. Our results match lower bounds in several settings, and we provide empirical validation of our algorithm on misspecified linear bandit and model selection problems.
Forward citations
Cited by 1 Pith paper
-
Offline-to-online hyperparameter transfer for stochastic bandits
Offline data from a distribution of bandit tasks provably identifies near-optimal algorithm hyperparameters, with inter-task sample complexity depending on a new piecewise-complexity measure QD.
Discussion (0). Continue with ORCID to comment.