REVIEW 2 cited by
Ensemble sampling for linear bandits: small ensembles suffice
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We provide the first useful and rigorous analysis of ensemble sampling for the stochastic linear bandit setting. In particular, we show that, under standard assumptions, for a $d$-dimensional stochastic linear bandit with an interaction horizon $T$, ensemble sampling with an ensemble of size of order $d \log T$ incurs regret at most of the order $(d \log T)^{5/2} \sqrt{T}$. Ours is the first result in any structured setting not to require the size of the ensemble to scale linearly with $T$ -- which defeats the purpose of ensemble sampling -- while obtaining near $\smash{\sqrt{T}}$ order regret. Our result is also the first to allow for infinite action sets.
Forward citations
Cited by 2 Pith papers
-
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
BLCE-G and BLCE achieve minimax-optimal regret for linear contextual bandits with only O(log log T) parameter updates and reduced computational cost by avoiding near G-optimal design.
-
Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning
A quantile-of-means ensemble method achieves minimax optimal variance-dependent regret bounds for finite-horizon MDPs without count-based uncertainty estimates.
Discussion (0). Sign in to comment.