REVIEW 4 major objections 6 minor 66 references
Rapid Learning in Constrained Minimax Games with Negative Momentum
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Negative momentum, borrowed from unconstrained games, carries over to constrained minimax solvers and yields the first algorithm reported to surpass CFR+ across multiple game types, with MoCFR+ reaching about $10^9$ times lower…
desk verdict Empirically promising negative-momentum acceleration of RM+/CFR+, but the main theorem has a parameter-range error and the proof does not cover the algorithms that produce the headline results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the momentum buffer kept in the dual space, $\mu_t = \beta\mu_{t-1} - F(z_t)$ with the negative coefficient $\beta<0$, together with its "Restarting Aggregated Momentum" (RAM) extension that sums buffer snapshots over an interval of length $k$ and periodically re-anchors an attachment point $L_{\text{att}}$ in the FTRL form $L_t = L_{t-1} + F(z_t) - \beta(L_{\text{att}} - L_{t-1})$. The negative sign turns the usual momentum extrapolation into a friction-like drag that damps oscillation in the strategy trajectory. The machinery is transferred to regret matching through the known FTRL/OMD-to-RM+ correspondence, where the attachment regret vector $R_{\text{att}}$ plays the anchor role, producing Algorithm 2 (MoRM+).
What would settle it
Run MoCFR+ and CFR+ on the same four games with identical alternation, averaging, and hyperparameter-selection protocols; if the exploitability gap does not remain roughly $10^9$ across seeds, the headline claim fails. A sharper check is to write the MoRM+ update from Algorithm 2 as an instance of online linear optimization: if the negative-momentum term $-\beta(R_{\text{att}} - R_t)$ cannot be expressed as a valid loss sequence under the FTRL/OMD-to-RM+ correspondence, then the theoretical bridge that justifies MoRM+ is broken.
Extended reading notes
Core claim
The central claim is that negative momentum, stored as a dual-space buffer $\mu_t = \beta\mu_{t-1} - F(z_t)$ with $\beta<0$, is a universal acceleration mechanism for constrained zero-sum solvers. Plugged into mirror descent or FTRL, it yields updates $z_{t+1} = \arg\min_z\{\eta\langle z, -\mu_t\rangle + D_\psi(z,z_t)\}$; with negative entropy as the regularizer and an infinitely large buffer, the algorithm converges exponentially to the regularized equilibrium $z^*$ of a modified game (Theorem 2), that point is an $O(-\beta/\eta)$-Nash equilibrium of the original game (Theorem 3), and with the RAM buffer re-anchored every $k$ steps the algorithm converges to the set of Nash equilibria (Theorem 4). The same buffer is grafted onto the regret update of RM+, giving $R_{t+1} = [R_t + r(x_t) - \beta(R_{\text{att}} - R_t)]_+$, which defines MoRM+ and, through regret decomposition, MoCFR+; the experiments claim these variants dominate their base algorithms and state-of-the-art baselines, with MoCFR+ reaching roughly nine orders of magnitude lower exploitability than CFR+.
Load-bearing premise
The load-bearing premise is that the negative-momentum trick, proven to converge in one particular regularized mirror-descent setting, still works when grafted into the different regret-matching update that the headline experiments use; no theorem in the paper covers that grafted algorithm.
Editorial extensions
If this is right
- MoCFR+ reaches exploitability about $10^9$ times lower than CFR+ and outperforms PCFR+ on all four tested extensive-form games, a claimed first for any algorithm across those game types.
- The negative-momentum buffer applies uniformly to OMD and FTRL instantiations such as MoMWU and DMoGDA, so existing game solvers can be upgraded by altering only the buffer/loss update.
- With negative entropy regularizer and an infinite buffer, convergence is exponential to an $O(-\beta/\eta)$-equilibrium; with finite $k$ under RAM, the algorithm provably reaches the exact Nash set.
- Larger $|\beta|$ accelerates convergence but biases the regularized equilibrium away from Nash, exposing an explicit speed-accuracy tradeoff controlled by $\beta$ and $k$.
Reading between the lines
- The paper proves convergence only for the entropy-regularized mirror-descent variant; whether the FTRL/OMD-to-RM+ correspondence survives with the negative-momentum term is left open, so the impressive MoCFR+ numbers currently rest on an analogy rather than a theorem.
- An adaptive rule for the re-anchoring interval $k$ — for example, triggering a reset when exploitability stalls — could make the method free of per-game tuning, since the paper's own ablation shows a stair-step pattern when $k$ is too large.
- The same dual-space buffer could be applied to Monte Carlo variants of CFR or to optimistic/predictive updates, where stochastic gradients would test whether the friction-like damping survives noise; that extension is not explored in the paper.
- A protocol-stricter comparison — simultaneous updates, untuned $\beta$, and identical averaging schemes across algorithms — would clarify how much of the reported $10^9$ advantage comes from the negative momentum itself rather than from alternation and averaging choices.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a negative-momentum mechanism for constrained minimax games, instantiated as MoMD/MoFTRL (with entropy regularization), MoRM+, and their extensive-form counterparts MoCFR+ and DMoGDA/DMoMWU. The theoretical claims are exponential convergence of the entropy-regularized momentum algorithm to an approximate equilibrium (Theorems 2 and 3) and convergence to Nash equilibria with a sufficiently large restart interval (Theorem 4). The experimental sections report large exploitability reductions on normal-form and extensive-form benchmarks, most notably that MoCFR+ attains roughly 10^9 lower exploitability than CFR+ and outperforms PCFR+ across Kuhn Poker, Leduc Poker, Goofspiel, and Liar's Dice.
Significance. If the theoretical results were correct and the empirical claims reproducible, this would be a meaningful contribution: it extends negative-momentum ideas from unconstrained bilinear games to constrained decision sets, proposes a simple restarting-buffer mechanism, and demonstrates strong empirical performance across several standard game benchmarks. The paper ships code, uses standard benchmarks, and includes ablation studies for the restart interval. However, the central theorem as stated contains a parameter-range error, the proof of Theorem 4 is only a sketch, and the headline empirical algorithm MoCFR+ has no convergence theorem. The theory therefore does not currently support the paper's central claims as cleanly as the text suggests.
major comments (4)
- [Appendix D, proof of Theorem 2] The final step of the proof is algebraically incorrect. The proof establishes Dψ(z*,z_{t+1}) ≤ (1+β/2)DKL(z*,z_t) + C·DKL(z_{t+1},z_t) with C = -1 - (3/2)β - 4η²/β, and contraction requires C ≤ 0. The stated condition η ≤ sqrt(-(1+3β/2)β/2) does not imply this. For example, β = -0.5 and η = 0.2 are allowed by the stated bound (which equals 0.25), yet C = -1 + 0.75 + 0.32 = 0.07 > 0. A correct condition is η ≤ sqrt(-(1+3β/2)β/4), a factor √2 smaller. Since Theorem 3 inherits this condition from Theorem 2, the theoretical guarantee as stated is invalid and must be corrected.
- [Theorems 2-3 vs. Table 1] Even with the corrected bound, none of the MoMWU hyperparameters in Table 1 satisfy the theorem's conditions: for the size-25 random game, β = -0.02 and η = 7, while the corrected bound is about 0.0696; for the 3×3 game, β = -0.06 and η = 1, while the corrected bound is about 0.117. In addition, Theorems 2 and 3 assume k = ∞, whereas all Table 1 experiments use finite k. Thus the sentence 'The following theoretical analysis supports the experimental results above' is not justified by the stated theorems, and the paper should either relegate the theory to a qualitative role or select experimental hyperparameters that fall inside the proven regime.
- [Appendix D, proof of Theorem 4] Theorem 4 asserts convergence to the set of Nash equilibria with a finite restart interval, but its proof is a sketch. It invokes Lemma 7 and Lemma 8 as 'adapted from Abe et al. (2023)' without stating or proving them for the present algorithm, and it does not show that the periodic attachment update in Algorithm 1 produces the sequence z_att^{(n)} for which Lemma 7's strict decrease holds. The compactness and continuity argument is plausible but incomplete. The paper needs a full proof of these lemmas in the current setting, or a clear statement that Theorem 4 is conjectural.
- [Algorithm 2 and Eq. (21)] The headline algorithm MoCFR+ (and its normal-form counterpart MoRM+) has no convergence guarantee in the paper. The RM+/OMD correspondence in Eq. (16) and Lemma 5 is imported from Farina et al. (2023) for the unmodified regret-matching update; the paper does not prove that this correspondence remains valid after the negative-momentum term -β(R_att - R_t) is inserted in Algorithm 2 and Eq. (21). Since the paper's main empirical claim concerns MoCFR+, the connection between the proven entropy-regularized MoMD analysis and the actual headline algorithm is an analogy rather than a theorem. The empirical results can stand on their own, but the paper should not describe them as supported by the theoretical analysis.
minor comments (6)
- [Introduction, after Eq. (8)] The text refers to 'Figure 5(a)' when discussing the divergence of GDAm with non-negative momentum; this should be Figure 1(a), since Figure 5 is the ablation figure in Appendix C.
- [Appendix D] There are typos: 'Yong's inequality' should be 'Young's inequality', and in the proof of Theorem 4 there is a stray parenthesis in 'each z(n)_att )'.
- [Appendix C] The text contains 'learing rate' instead of 'learning rate' and 'The experiment results' instead of 'The experimental results.'
- [Preliminaries] The phrase 'norm-form' should be 'normal-form games'.
- [Table 1] Use consistent notation '3×3' instead of '3*3' for the matrix game.
- [Experiments, EFG section] The strong claim that this is 'the first instance where an algorithm surpasses CFR+ performance across various types of games' should be tempered or substantiated with a more precise comparison to existing strong CFR variants, since the claim goes beyond the experiments reported here.
Circularity Check
No circular derivation: the convergence theorems and the MoCFR+ extension are independent of the experimental claims, though Theorem 2 has a separate algebraic gap.
full rationale
The paper's theoretical results (Theorems 2-4) are genuine Lyapunov-style convergence proofs for the entropy-regularized MoMWU/RAM algorithm to a regularized equilibrium defined by an anchor point; this is self-referential in the sense that the anchor is algorithm-dependent, but that is a standard regularized-saddle-point fixed-point analysis, not a fit or a parameter renamed as a prediction. The proof of Theorem 2 derives the contraction inequality from first-order optimality, Pinsker's inequality, and Young's inequality; it does not assume its own conclusion. The anchor z_att is either fixed (k=∞) or updated deterministically, and the modified game in Eq. (15) is a definition of the target equilibrium, not an input fitted to the experimental output. The extension to MoRM+/MoCFR+ is presented as an algorithmic analogy built on the external FTRL/OMD-to-RM+ correspondence of Farina et al. 2023 (Eq. 16, Lemma 5, Algorithm 2, Eq. 21); no theorem in the paper covers MoRM+/MoCFR+, so the empirical speedup is explicitly an experimental observation rather than a derived prediction of the entropy-regularizer theory. That is a correctness or support gap, not circularity. There is no load-bearing self-citation chain, no fitted parameter relabeled as a prediction, and no imported uniqueness theorem used to force the choice of algorithm. The algebraic error in Theorem 2's stated η bound, and the fact that the Table 1 and Table 2 hyperparameters violate even the corrected bound, are serious correctness defects but do not make the derivation circular. Therefore no significant circularity is present.
Assumptions & free parameters
free parameters (3)
- negative momentum coefficient beta =
per game, e.g., -0.04 (3x3), -0.02 (random NFG), -0.2 (Kuhn)
- restart interval k =
per game, e.g., 50 (3x3), 100 (random NFG), 5-100 across EFGs
- learning rate eta =
per game, e.g., 1.0 (3x3), 7.0 (random 25), 2.0 (Kuhn)
assumptions (4)
- standard math F is monotone in the bilinear game, so the inner product bound used after Eq. (32) holds.
- domain assumption The regularized equilibrium z* of the modified game Eq. (15) satisfies the variational inequality invoked below Eq. (33).
- ad hoc to paper Lemmas 7 and 8, adapted from Abe et al. (2023), hold for the periodic attachment updates in Algorithm 1.
- ad hoc to paper The RM+/OMD correspondence (Eq. 16 and Lemma 5 from Farina et al. 2023) remains valid when the negative-momentum term is added in Algorithm 2.
Cite this review
Pith. "Pith review of Rapid Learning in Constrained Minimax Games with Negative Momentum." pith.science (2026). https://pith.science/paper/ZYUK5Y47
@misc{pith2026250100533,
author = {Pith},
title = {Pith review of: Rapid Learning in Constrained Minimax Games with Negative Momentum},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZYUK5Y47}},
note = {Machine review of arXiv:2501.00533}
}
read the original abstract
In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the constrained setting and provides a universal enhancement to the classic game-solver algorithms. Additionally, we provide theoretical guarantee of convergence for our momentum-augmented algorithms with entropy regularizer. We then extend these algorithms to their extensive-form counterparts. Experimental results on both Normal Form Games (NFGs) and Extensive Form Games (EFGs) demonstrate that our momentum techniques can significantly improve algorithm performance, surpassing both their original versions and the SOTA baselines by a large margin.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Abe, K.; Ariu, K.; Sakamoto, M.; and Iwasaki, A. 2024. Adaptively Perturbed Mirror Descent for Learning in Games. arXiv:2305.16610
work page Pith review arXiv 2024
-
[4]
Abe, K.; Ariu, K.; Sakamoto, M.; Toyoshima, K.; and Iwasaki, A. 2023. Last-Iterate Convergence with Full and Noisy Feedback in Two-Player Zero-Sum Games. arXiv:2208.09855
work page Pith review arXiv 2023
-
[5]
Abernethy, J. D.; Hazan, E.; and Rakhlin, A. 2008. Competing in the Dark: An Efficient Algorithm for Bandit Linear Optimization. In Annual Conference on Learning Theory, 263--274. Omnipress
work page 2008
-
[6]
Balduzzi, D.; Racaniere, S.; Martens, J.; Foerster, J.; Tuyls, K.; and Graepel, T. 2018. The mechanics of n-player differentiable games. In International Conference on Machine Learning, 354--363. PMLR
work page 2018
-
[7]
Berry, M.; and Shukla, P. 2016. Curl force dynamics: symmetries, chaos and constants of motion. New Journal of Physics, 18(6): 063018
work page 2016
-
[8]
Brown, N.; and Sandholm, T. 2018. Superhuman AI for heads-up no-limit poker: Libratus beats top professionals. Science, 359(6374): 418--424
2018
Show all 66 references
-
[9]
Burch, N.; Moravcik, M.; and Schmid, M. 2019. Revisiting CFR+ and alternating updates. Journal of Artificial Intelligence Research, 64: 429--443
2019
-
[10]
Cai, Y.; Farina, G.; Grand-Clément, J.; Kroer, C.; Lee, C.-W.; Luo, H.; and Zheng, W. 2023. Last-Iterate Convergence Properties of Regret-Matching Algorithms in Games. arXiv:2311.00676
2023 arXiv
-
[11]
U.; Fleuret, F.; and Jaggi, M
Chavdarova, T.; Pagliardini, M.; Stich, S. U.; Fleuret, F.; and Jaggi, M. 2021. Taming GANs with Lookahead-Minmax. arXiv:2006.14567
2021 arXiv
-
[12]
M.; Gidel, G.; Tracey, B.; Tuyls, K.; Omidshafiei, S.; Balduzzi, D.; and Jaderberg, M
Czarnecki, W. M.; Gidel, G.; Tracey, B.; Tuyls, K.; Omidshafiei, S.; Balduzzi, D.; and Jaderberg, M. 2020. Real world games look like spinning tops. Advances in Neural Information Processing Systems, 33: 17443--17454
2020
-
[13]
J.; and Golowich, N
Daskalakis, C.; Foster, D. J.; and Golowich, N. 2020. Independent Policy Gradient Methods for Competitive Reinforcement Learning. In Advances in Neural Information Processing Systems
2020
-
[14]
Daskalakis, C.; and Panageas, I. 2019. Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization. In Innovations in Theoretical Computer Science Conference, volume 124 of LIPIcs, 27:1--27:18. Schloss Dagstuhl - Leibniz-Zentrum f \" u r Informatik
2019
-
[15]
S.; Chen, J.; Li, L.; Xiao, L.; and Zhou, D
Du, S. S.; Chen, J.; Li, L.; Xiao, L.; and Zhou, D. 2017. Stochastic variance reduction methods for policy evaluation. In International Conference on Machine Learning, 1049--1058. PMLR
2017
-
[16]
Farina, G.; Grand-Clément, J.; Kroer, C.; Lee, C.-W.; and Luo, H. 2023. Regret Matching+: (In)Stability and Fast Convergence in Games. arXiv:2305.14709
2023 arXiv
-
[17]
Farina, G.; Kroer, C.; Brown, N.; and Sandholm, T. 2019. Stable-Predictive Optimistic Counterfactual Regret Minimization. In International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 1853--1862. PMLR
2019
-
[18]
Farina, G.; Kroer, C.; and Sandholm, T. 2019 a . Online convex optimization for sequential decision processes and extensive-form games. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, 1917--1925
2019
-
[19]
Farina, G.; Kroer, C.; and Sandholm, T. 2019 b . Optimistic Regret Minimization for Extensive-Form Games via Dilated Distance-Generating Functions. In Advances in Neural Information Processing Systems, 5222--5232
2019
-
[20]
Farina, G.; Kroer, C.; and Sandholm, T. 2021 a . Better Regularization for Sequential Decision Spaces: Fast Convergence Rates for Nash, Correlated, and Team Equilibria. In EC '21: The 22nd ACM Conference on Economics and Computation , 432. ACM
2021
-
[21]
Farina, G.; Kroer, C.; and Sandholm, T. 2021 b . Faster game solving via predictive blackwell approachability: Connecting regret matching and mirror descent. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 5363--5371
2021
-
[22]
Fiez, T.; and Ratliff, L. J. 2021. Local convergence analysis of gradient descent ascent with finite timescale separation. In International Conference on Learning Representation
2021
-
[23]
N.; Chen, R
Foerster, J. N.; Chen, R. Y.; Al-Shedivat, M.; Whiteson, S.; Abbeel, P.; and Mordatch, I. 2017. Learning with Opponent-Learning Awareness. In Adaptive Agents and Multi-Agent Systems
2017
-
[24]
A.; Pezeshki, M.; Le Priol, R.; Huang, G.; Lacoste-Julien, S.; and Mitliagkas, I
Gidel, G.; Hemmat, R. A.; Pezeshki, M.; Le Priol, R.; Huang, G.; Lacoste-Julien, S.; and Mitliagkas, I. 2019. Negative momentum for improved game dynamics. In The 22nd International Conference on Artificial Intelligence and Statistics, 1802--1811. PMLR
2019
-
[25]
Golowich, N.; Pattathil, S.; and Daskalakis, C. 2020. Tight last-iterate convergence rates for no-regret learning in multi-player games. Advances in Neural Information Processing Systems, 33: 20766--20778
2020
-
[26]
J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A
Goodfellow, I. J.; Pouget - Abadie, J.; Mirza, M.; Xu, B.; Warde - Farley, D.; Ozair, S.; Courville, A. C.; and Bengio, Y. 2020. Generative adversarial networks. Communications of the ACM, 63(11): 139--144
2020
-
[27]
Grand-Cl \'e ment, J.; and Kroer, C. 2024. Solving optimization problems with Blackwell approachability. Mathematics of Operations Research, 49(2): 697--728
2024
-
[28]
Gulrajani, I.; Ahmed, F.; Arjovsky, M.; Dumoulin, V.; and Courville, A. C. 2017. Improved Training of Wasserstein GANs. In Advances in Neural Information Processing Systems, 5767--5777
2017
-
[29]
Hart, S.; and Mas-Colell, A. 2000. A simple adaptive procedure leading to correlated equilibrium. Econometrica, 68(5): 1127--1150
2000
-
[30]
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017. GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. In Advances in Neural Information Processing Systems, 6626--6637
2017
-
[31]
Hoda, S.; Gilpin, A.; Pena, J.; and Sandholm, T. 2010. Smoothing techniques for computing Nash equilibria of sequential games. Mathematics of Operations Research, 35(2): 494--512
2010
-
[32]
Huang, K.; and Zhang, S. 2022. New first-order algorithms for stochastic variational inequalities. SIAM Journal on Optimization, 32(4): 2745--2772
2022
-
[33]
Korpelevich, G. M. 1976. The extragradient method for finding saddle points and other problems. Matecon, 12: 747--756
1976
-
[34]
Kroer, C.; Peysakhovich, A.; Sodomka, E.; and Stier-Moses, N. E. 2019. Computing large market equilibria using abstractions. In ACM Conference on Economics and Computation, 745--746
2019
-
[35]
Kroer, C.; Waugh, K.; K l n c -Karzan, F.; and Sandholm, T. 2020. Faster algorithms for extensive-form game solving via improved smoothing functions. Mathematical Programming, 179(1-2): 385--417
2020
-
[36]
Kuhn, H. W. 1950. A simplified two-person poker. Contributions to the Theory of Games, 1: 97--103
1950
-
[37]
D.; Saeta, B.; Bradbury, J.; Ding, D.; Borgeaud, S.; Lai, M.; Schrittwieser, J.; Anthony, T.; Hughes, E.; Danihelka, I.; and Ryan-Davis, J
Lanctot, M.; Lockhart, E.; Lespiau, J.-B.; Zambaldi, V.; Upadhyay, S.; Pérolat, J.; Srinivasan, S.; Timbers, F.; Tuyls, K.; Omidshafiei, S.; Hennes, D.; Morrill, D.; Muller, P.; Ewalds, T.; Faulkner, R.; Kramár, J.; Vylder, B. D.; Saeta, B.; Bradbury, J.; Ding, D.; Borgeaud, S...
2020 arXiv
-
[38]
Lanctot, M.; Waugh, K.; Zinkevich, M.; and Bowling, M. H. 2009. Monte Carlo Sampling for Regret Minimization in Extensive Games. In Advances in Neural Information Processing Systems, 1078--1086. Curran Associates, Inc
2009
-
[39]
Lee, C.; Kroer, C.; and Luo, H. 2021. Last-iterate Convergence in Extensive-Form Games. In Advances in Neural Information Processing Systems, 14293--14305
2021
-
[40]
Liang, T.; and Stokes, J. 2019. Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks. In International Conference on Artificial Intelligence and Statistics, 907--915. PMLR
2019
-
[41]
Lis \' y , V.; Lanctot, M.; and Bowling, M. H. 2015. Online Monte Carlo Counterfactual Regret Minimization for Search in Imperfect Information Games. In International Conference on Autonomous Agents and Multiagent Systems, 27--36. ACM
2015
-
[42]
Liu, M.; Ozdaglar, A.; Yu, T.; and Zhang, K. 2023. The Power of Regularization in Solving Extensive-Form Games. arXiv:2206.09495
2023 arXiv
-
[43]
Liu, W.; Jiang, H.; Li, B.; and Li, H. 2022. Equivalence Analysis between Counterfactual Regret Minimization and Online Mirror Descent. In International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, 13717--13745. PMLR
2022
-
[44]
Lockhart, E.; Lanctot, M.; P \' e rolat, J.; Lespiau, J.; Morrill, D.; Timbers, F.; and Tuyls, K. 2019. Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent. In International Joint Conference on Artificial Intelligence, 464--470
2019
-
[45]
P.; Acuna, D.; Vicol, P.; and Duvenaud, D
Lorraine, J. P.; Acuna, D.; Vicol, P.; and Duvenaud, D. 2022. Complex momentum for optimization in games. In International Conference on Artificial Intelligence and Statistics, 7742--7765. PMLR
2022
-
[46]
Madras, D.; Creager, E.; Pitassi, T.; and Zemel, R. S. 2018. Learning Adversarially Fair and Transferable Representations. In International Conference on Machine Learning, 3381--3390. PMLR
2018
-
[47]
Mertikopoulos, P.; Lecouat, B.; Zenati, H.; Foo, C.; Chandrasekhar, V.; and Piliouras, G. 2019. Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile. In International Conference on Learning Representations
2019
-
[48]
M.; Nowozin, S.; and Geiger, A
Mescheder, L. M.; Nowozin, S.; and Geiger, A. 2017. The Numerics of GANs. In Advances in Neural Information Processing Systems, 1825--1835
2017
-
[49]
Morav c \' k, M.; Schmid, M.; Burch, N.; Lis \`y , V.; Morrill, D.; Bard, N.; Davis, T.; Waugh, K.; Johanson, M.; and Bowling, M. 2017. Deepstack: Expert-level artificial intelligence in heads-up no-limit poker. Science, 356(6337): 508--513
2017
-
[50]
Orabona, F. 2023. A Modern Introduction to Online Learning. arXiv:1912.13213
2023 arXiv
-
[51]
Peng, W.; Dai, Y.-H.; Zhang, H.; and Cheng, L. 2020. Training GANs with centripetal acceleration. Optimization Methods and Software, 35(5): 955--973
2020
-
[52]
P \'e rolat, J.; Munos, R.; Lespiau, J.-B.; Omidshafiei, S.; Rowland, M.; Ortega, P.; Burch, N.; Anthony, T.; Balduzzi, D.; De Vylder, B.; et al. 2021. From poincar \'e recurrence to convergence in imperfect information games: Finding equilibrium via regularization. In Interna...
2021
-
[53]
Polyak, B. T. 1964. Some methods of speeding up the convergence of iteration methods. Ussr computational mathematics and mathematical physics, 4(5): 1--17
1964
-
[54]
Ross, S. M. 1971. Goofspiel—the game of pure strategy. Journal of Applied Probability, 8(3): 621--625
1971
-
[55]
Sch \" a fer, F.; and Anandkumar, A. 2019. Competitive Gradient Descent. In Advances in Neural Information Processing Systems, 7623--7633
2019
-
[56]
S.; Su, W
Shi, B.; Du, S. S.; Su, W. J.; and Jordan, M. I. 2019. Acceleration via Symplectic Discretization of High-Resolution Differential Equations. In Advances in Neural Information Processing Systems
2019
-
[57]
Sinha, A.; Namkoong, H.; and Duchi, J. C. 2018. Certifying Some Distributional Robustness with Principled Adversarial Training. In International Conference on Learning Representations
2018
-
[58]
Z.; Loizou, N.; Lanctot, M.; Mitliagkas, I.; Brown, N.; and Kroer, C
Sokota, S.; D'Orazio, R.; Kolter, J. Z.; Loizou, N.; Lanctot, M.; Mitliagkas, I.; Brown, N.; and Kroer, C. 2023. A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum Games. In International Conference on Learning Representations,
2023
-
[59]
P.; Larson, B.; Piccione, C.; Burch, N.; Billings, D.; and Rayner, C
Southey, F.; Bowling, M. P.; Larson, B.; Piccione, C.; Burch, N.; Billings, D.; and Rayner, C. 2012. Bayes' Bluff: Opponent Modelling in Poker. arXiv:1207.1411
2012 arXiv
-
[60]
Tammelin, O. 2014. Solving Large Imperfect Information Games Using CFR+. arXiv:1407.5042
2014 arXiv
-
[61]
Neumann, J
v. Neumann, J. 1928. Zur theorie der gesellschaftsspiele. Mathematische annalen, 100(1): 295--320
1928
-
[62]
Vlatakis - Gkaragkounis, E.; Flokas, L.; and Piliouras, G. 2019. Poincar \' e Recurrence, Cycles and Spurious Equilibria in Gradient-Descent-Ascent for Non-Convex Non-Concave Zero-Sum Games. In Advances in Neural Information Processing Systems, 10450--10461
2019
-
[63]
K.; Jagota, A
Warmuth, M. K.; Jagota, A. K.; et al. 1997. Continuous and discrete-time nonlinear gradient descent: Relative loss bounds and convergence. In International Symposium on Artificial Intelligence and Mathematics, volume 326. Citeseer
1997
-
[64]
Wei, C.; Lee, C.; Zhang, M.; and Luo, H. 2021. Linear Last-iterate Convergence in Constrained Saddle-point Optimization. In International Conference on Learning Representations
2021
-
[65]
Zhang, G.; and Wang, Y. 2021. On the suboptimality of negative momentum for minimax optimization. In International Conference on Artificial Intelligence and Statistics, 2098--2106. PMLR
2021
-
[66]
Zinkevich, M.; Johanson, M.; Bowling, M.; and Piccione, C. 2007. Regret minimization in games with incomplete information. Advances in Neural Information Processing Systems, 20
2007
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.