A constant alpha-divergence bound between the true and approximate posterior does not guarantee sub-linear regret in Thompson sampling; adding forced exploration restores sub-linear regret for alpha<=0.
Thompson sampling for complex online problems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Thompson Sampling with Approximate Inference
A constant alpha-divergence bound between the true and approximate posterior does not guarantee sub-linear regret in Thompson sampling; adding forced exploration restores sub-linear regret for alpha<=0.