The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.
Learning, Regret Minimization, and Equilibria, page 79–102
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.GT 1years
2025 1verdicts
REJECT 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Decision-Making Under Complete Uncertainty: You Will Regret Not Being Greedy
The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.