The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.
Cormen, Charles E
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.GT 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Decision-Making Under Complete Uncertainty: You Will Regret Not Being Greedy
The greedy strategy of choosing the highest observed average rating has minimax-optimal worst-case regret 1/8 in a two-product, two-rating, one-observation game, and its regret tends to zero with more observations, but the claim that Thompson Sampling retains positive regret is flawed.