The softmax gradient bandit converges almost surely to the optimal action for any constant learning rate, removing the small-learning-rate restriction of prior work.
Analysis of thompson sampling for the multi-armed bandit problem
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Small steps no more: Global convergence of stochastic gradient bandits for arbitrary learning rates
The softmax gradient bandit converges almost surely to the optimal action for any constant learning rate, removing the small-learning-rate restriction of prior work.