A UCB-based algorithm for learning matching equilibria with bandit feedback claims an O~(sqrt(T mk pa)) regret bound, but the proof's final step undercounts the number of pairs matched per round.
Uncoupled learning dynamics with o(\log t) swap regret in multiplayer games
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning in Matching Games with Bandit Feedback
A UCB-based algorithm for learning matching equilibria with bandit feedback claims an O~(sqrt(T mk pa)) regret bound, but the proof's final step undercounts the number of pairs matched per round.