An explore-exploit algorithm achieves O(OPT^{2/3}) regret when combining multiple MTS heuristics with bandit access, and this is tight up to log factors.
Better best of both worlds bounds for bandits with switching costs
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning-Augmented Algorithms for MTS with Bandit Access to Multiple Predictors
An explore-exploit algorithm achieves O(OPT^{2/3}) regret when combining multiple MTS heuristics with bandit access, and this is tight up to log factors.