REVIEW 2 cited by
$O(T^{-1})$ Convergence of Optimistic-Follow-the-Regularized-Leader in Two-Player Zero-Sum Markov Games
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We prove that optimistic-follow-the-regularized-leader (OFTRL), together with smooth value updates, finds an $O(T^{-1})$-approximate Nash equilibrium in $T$ iterations for two-player zero-sum Markov games with full information. This improves the $\tilde{O}(T^{-5/6})$ convergence rate recently shown in the paper Zhang et al (2022). The refined analysis hinges on two essential ingredients. First, the sum of the regrets of the two players, though not necessarily non-negative as in normal-form games, is approximately non-negative in Markov games. This property allows us to bound the second-order path lengths of the learning dynamics. Second, we prove a tighter algebraic inequality regarding the weights deployed by OFTRL that shaves an extra $\log T$ factor. This crucial improvement enables the inductive analysis that leads to the final $O(T^{-1})$ rate.
Forward citations
Cited by 2 Pith papers
-
Solving Zero-Sum Convex Markov Games
Independent policy-gradient algorithms provably compute approximate Nash equilibria in two-player zero-sum convex Markov games.
-
Minimax-Optimal Multi-Agent Robust Reinforcement Learning
Robust Q-FTRL achieves ε-robust CCE in R-contaminated Markov games with H^3 S Σ_i A_i min{H,1/R}/ε^2 samples up to logs, matching a new lower bound; two-player zero-sum gives NE.
Discussion (0). Continue with ORCID to comment.