Pith. sign in

$O(T^{-1})$ Convergence of Optimistic-Follow-the-Regularized-Leader in Two-Player Zero-Sum Markov Games

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We prove that optimistic-follow-the-regularized-leader (OFTRL), together with smooth value updates, finds an $O(T^{-1})$-approximate Nash equilibrium in $T$ iterations for two-player zero-sum Markov games with full information. This improves the $\tilde{O}(T^{-5/6})$ convergence rate recently shown in the paper Zhang et al (2022). The refined analysis hinges on two essential ingredients. First, the sum of the regrets of the two players, though not necessarily non-negative as in normal-form games, is approximately non-negative in Markov games. This property allows us to bound the second-order path lengths of the learning dynamics. Second, we prove a tighter algebraic inequality regarding the weights deployed by OFTRL that shaves an extra $\log T$ factor. This crucial improvement enables the inductive analysis that leads to the final $O(T^{-1})$ rate.

fields

cs.GT 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Solving Zero-Sum Convex Markov Games

cs.GT · 2025-06-19 · conditional · novelty 7.0

Independent policy-gradient algorithms provably compute approximate Nash equilibria in two-player zero-sum convex Markov games.

citing papers explorer

Showing 1 of 1 citing paper.

  • Solving Zero-Sum Convex Markov Games cs.GT · 2025-06-19 · conditional · none · ref 127 · internal anchor

    Independent policy-gradient algorithms provably compute approximate Nash equilibria in two-player zero-sum convex Markov games.