The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.
Greedy when sure and conservative when uncertain about the opponents
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.GT 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Decision-Making Under Complete Uncertainty: You Will Regret Not Being Greedy
The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.