Pith. sign in

REVIEW

Ensemble sampling for linear bandits: small ensembles suffice

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.08376 v4 pith:UIYE6NYC submitted 2023-11-14 stat.ML cs.LG

classification stat.MLcs.LG
keywords ensemblesamplingfirstlinearorderbanditregretresult
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We provide the first useful and rigorous analysis of ensemble sampling for the stochastic linear bandit setting. In particular, we show that, under standard assumptions, for a $d$-dimensional stochastic linear bandit with an interaction horizon $T$, ensemble sampling with an ensemble of size of order $d \log T$ incurs regret at most of the order $(d \log T)^{5/2} \sqrt{T}$. Ours is the first result in any structured setting not to require the size of the ensemble to scale linearly with $T$ -- which defeats the purpose of ensemble sampling -- while obtaining near $\smash{\sqrt{T}}$ order regret. Our result is also the first to allow for infinite action sets.

Discussion (0). Continue with ORCID to comment.

Pith tools