A new two-phase UCB algorithm, Safe-LUCB, achieves Õ(√T) regret with a known positive safety gap and Õ(T^{2/3}) regret in the worst case, while keeping every action safe with high probability.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits.Theoretical Computer Science, 410(19):1876–1902, 2009
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Linear Stochastic Bandits Under Safety Constraints
A new two-phase UCB algorithm, Safe-LUCB, achieves Õ(√T) regret with a known positive safety gap and Õ(T^{2/3}) regret in the worst case, while keeping every action safe with high probability.