Pith. sign in

REVIEW 2 major objections 3 minor 2 cited by

The paper proves that for any C² payoff on a torus, the mean-field Langevin descent-ascent dynamics converges exponentially fast to the unique mixed Nash equilibrium, provided the initial strategies are sufficiently close in Wasserstein dis

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 05:36 UTC pith:WKYHCY3W

load-bearing objection Real local stability theorem for MFL-DA, but the finite-particle claim in the title/abstract is not delivered anywhere in the paper. the 2 major comments →

arxiv 2602.01564 v2 pith:WKYHCY3W submitted 2026-02-02 cs.LG math.APmath.OCmath.PR

Local exponential stability of mean-field Langevin descent-ascent and associated particle system

classification cs.LG math.APmath.OCmath.PR MSC 35B3549Q2291A10
keywords mean-field Langevin descent-ascentmixed Nash equilibriumWasserstein gradient flowexponential stabilityspectral gapzero-sum gamesentropic regularizationinteracting particle system
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper answers a local version of an open problem about the behavior of mean-field Langevin descent-ascent (MFL-DA), a coupled optimization flow on probability measures used for entropically regularized two-player zero-sum games. The authors prove that the unique mixed Nash equilibrium is locally exponentially stable: if both players start within a small Wasserstein neighborhood of the equilibrium, their strategy distributions converge to it at a quantitative exponential rate, and the densities converge in C^{1,α}. The key is to show that, near the equilibrium, the effective payoff becomes strongly convex–concave in the Wasserstein sense, even when the original payoff is nonconvex–nonconcave. This is established by a spectral-gap analysis of the linearized operator at the equilibrium. If true, the result gives a stable basin of attraction for the equilibrium and supplies the quantitative convergence rate that was missing.

Core claim

Theorem 1.1 asserts that, on the flat torus, for any C² payoff f, the unique mixed Nash equilibrium (μ*, ν*) of the entropy-regularized game is locally exponentially attracting for the MFL-DA flow. There exist δ, λ, C>0 such that if W₂(μ₀,μ*) + W₂(ν₀,ν*) < δ, then W₂(μ_t,μ*) + W₂(ν_t,ν*) ≤ C e^{-λt} times the initial distance, and the densities converge in C^{1,α} for every α∈(0,1). The proof shows that the second variation of the free energy at the equilibrium is uniformly coercive in the Wasserstein geometry, with a coercivity constant λ_gap > 0 coming from the spectral gap of the linearized operator; this local coercivity extends to a neighborhood, yielding a local evolution-variational i

What carries the argument

The key machinery is the linearized operator L acting on gradient vector fields, LΦ = -ρ*^{-1}∇·(ρ*∇Φ) + ∇²V_{ν*}·Φ, defined at the equilibrium density ρ* with effective potential V_{ν*}. A spectral gap argument shows its smallest eigenvalue λ_gap is strictly positive, which upgrades the merely non-negative second variation of the entropy-regularized payoff to a uniform coercivity on the Wasserstein tangent space. This coercivity, combined with a continuity argument (Lemma 3.3), yields a locally strong displacement convex–concave structure, and the resulting evolution variational inequalities imply exponential contraction.

Load-bearing premise

The proof requires the mixed Nash equilibrium to have a smooth, uniformly positive density, so that the linearized operator has a positive spectral gap; if the equilibrium density is not uniformly bounded away from zero, the local coercivity that drives exponential convergence is no longer guaranteed.

What would settle it

Compute the spectral gap λ_gap for a specific C² payoff on the torus and simulate MFL-DA from a small random perturbation of the equilibrium; if the observed rate is not at least close to λ_gap, or if trajectories escape for arbitrarily small perturbations, the theorem's conclusion would be refuted. A more direct test: construct a C² payoff whose equilibrium density vanishes somewhere; if the dynamics still shows exponential convergence, the uniform-positivity assumption is not necessary.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Local convergence with quantitative rate: any trajectory starting within a small Wasserstein neighbourhood of the equilibrium converges exponentially, and the rate λ can be chosen arbitrarily close to the spectral gap λ_gap.
  • Stronger notions of convergence: since C⁰ convergence of densities controls the Nikaido–Isoda error, the result also gives convergence of the game-theoretic payoff gap, not just Wasserstein distance.
  • Robustness to geometry: the proof extends to any compact Riemannian manifold without boundary, with the same coercivity identity, so the stability is not an artifact of the flat torus.
  • Global result under convexity: if the payoff is λ-convex–concave, the local argument becomes global, giving exponential convergence from arbitrary initial data.
  • Characterization of the rate: λ_gap is explicitly given as a Rayleigh quotient involving ∇²φ and the Hessian of the effective potential, so the convergence rate is computable in principle.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The spectral-gap mechanism suggests a definition of a 'condition number' for mean-field zero-sum games, analogous to strong-convexity constants; games with larger λ_gap should admit faster particle-based algorithms.
  • The same coercivity framework could be applied to discrete-time and particle approximations, potentially yielding non-asymptotic convergence rates uniform in the number of particles, provided the mean-field limit is stable.
  • The local convex-concave structure implies that heuristics from finite-dimensional GDA—such as momentum or adaptive step sizes—may transfer to the mean-field setting near equilibrium, offering a design principle for practical algorithms.
  • If the equilibrium density degenerates (approaches zero), the spectral gap may shrink to zero; exploring such boundary cases could reveal sharp thresholds for local stability.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper proves that the mean-field Langevin descent-ascent (MFL-DA) dynamics, on the flat torus and for a C^2 payoff f, is locally exponentially stable around its unique mixed Nash equilibrium: if the initial pair (μ0,ν0) is sufficiently close to (μ*,ν*) in W2, then the W2 distance of the solution decays exponentially at a quantitative rate, and the densities converge in C^{1,α}. The proof proceeds by establishing a positive spectral gap λ_gap for the linearized Wasserstein Hessian at equilibrium (Lemma 3.1), extending this to a local convex–concave coercivity in a W^{2,p} neighborhood (Lemma 3.3), deriving a local contraction estimate (Proposition 3.4), and bootstrapping this via smoothing and interpolation lemmas to obtain the final theorem. The abstract and title additionally claim that the finite-N particle system inherits this stability, but no such result is proved in the manuscript.

Significance. If the PDE-level theorem is correct—and the presented proofs are detailed and internally consistent—it gives a substantive positive answer to the local-stability and quantitative-rate parts of the Wang–Chizat open problem. The key mechanism, extracting a coercivity estimate for the entropy near equilibrium via spectral analysis of the linearized operator, is a useful and potentially transferable idea. The full appendix proofs are a strength: Lemmas 3.1, 3.3, 3.6, 3.7 and Proposition 3.4 are proved in detail, and the bootstrap argument in Theorem 1.1 is explicit. However, the advertised finite-N inheritance claim is not delivered anywhere in the paper, which significantly affects the paper's stated contribution as written.

major comments (2)
  1. [Abstract and Section 4] The abstract states: 'We further show that the finite-N particle system inherits this stability up to times exponential in N, with an N-independent exponential rate modulo a finite-particle error floor.' The title similarly advertises 'and associated particle system'. However, no theorem, proposition, or lemma in the main text or in Appendices A–B states or proves any finite-N inheritance result. Section 4 explicitly lists as the third open question 'whether the local stability proved here at the PDE level can be transferred to the finite-particle dynamics, yielding convergence rates that are uniform in the number of particles.' This is an explicit admission that the advertised finite-N result is not delivered. This is not a technical flaw in the PDE theorem, but it is a substantial overstatement of the paper's contributions. The authors must either supply a proof of the finite-N stabili
  2. [Theorem 1.1 / proof of (1.5)] The proof of Theorem 1.1 uses Lemma 3.5 with λ' possibly negative and then bootstraps from t≥1. In the displayed line before (1.5), the authors write W2^2(t) < e^{-2λ'} (initial) e^{-2λ(t-1)}, which is consistent only if the initial W2^2 is smaller than e^{2λ'} δ; this is stated earlier as W2^2(μ0,μ*)+W2^2(ν0,ν*) < e^{2λ'}δ. The argument is coherent, but the constant C in (1.5) depends on the possibly large factor e^{-2λ'}. The theorem requires only existence of some C, so this is not an error; however, the presentation could mislead a reader into thinking C is uniform in f. Please clarify that C is allowed to depend on f (including through λ'), as the abstract already implies.
minor comments (3)
  1. [Lemma 3.5 and footnote 7] The assumption '∇²yy f(x,y) ≤ −λ'I' combined with footnote 7 ('∇²yy f + λ'I is nonpositive definite') is confusing when λ' is allowed to be negative. Since the lemma is used for a possibly negative λ', the intended interpretation is that λ'-displacement concavity is meant in the sense of Definition A.25. Please state the hypothesis in the language of Definition A.25 or add a sentence clarifying the sign convention for negative λ'.
  2. [Remark 3.8] The claim W2(μt,μ*)+W2(νt,ν*) = O(e^{-λt}) for all λ ∈ (0, λ_gap) is stated as a remark following the theorem. This is a consequence of (1.5) with C depending on λ, but the asymptotic notation O(e^{-λt}) usually hides constants that can depend on λ; please make this dependence explicit to avoid ambiguity.
  3. [Section 3.1, Eq. (3.1)] The operator L is introduced formally as an object on gradient vector fields. In Remark B.1, the authors discuss a spectral decomposition and an L²μ*-orthogonal Hodge projection, but this is not used in the proof. This remark is interesting but somewhat speculative; consider shortening or marking it as a heuristic aside.

Circularity Check

0 steps flagged

No circular derivation: the mean-field stability proof is self-contained; minor self-citations are re-proved. The abstract's finite-N claim is unsupported and explicitly left open in the Conclusion, which is a claims mismatch, not circularity.

full rationale

The central chain (Lemma 3.1 -> Lemma 3.3 -> Proposition 3.4 -> Theorem 1.1) is internally consistent and does not reduce to its inputs. lambda_gap is defined as a variational Rayleigh quotient in (B.7) and proved positive by compactness/ellipticity; it is not fitted to the convergence data, and the exponential rate is then derived rather than assumed. The local convex-concavity lemma is obtained by continuity of the second variation, and the bootstrap uses external smoothing/interpolation facts proved in Appendix B.3. The only self-citations, mainly 'c.f. [CJS25]' in Lemmas A.20-A.21 and Lemma 3.6, are not load-bearing because the cited statements are re-proved in the appendices; hence they do not create circularity. The abstract's sentence 'We further show that the finite-N particle system inherits this stability up to times exponential in N...' has no matching theorem, and Section 4 explicitly lists as open 'whether the local stability proved here at the PDE level can be transferred to the finite-particle dynamics, yielding convergence rates that are uniform in the number of particles.' This is an overstatement of delivered results, not a circularity. Appendix A.10 similarly notes that R^n requires extra coercivity assumptions; this is an honest limitation. Overall no circular step is present; the low score reflects only the minor self-citation burden and the flagged claims-content mismatch.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The central argument rests on standard optimal transport, spectral, and PDE tools; no fitted constants or invented entities. The main external dependence is the authors' own CJS25 preprint for several technical stability lemmas, though proofs are largely reproduced here.

axioms (4)
  • domain assumption Unique mixed Nash equilibrium exists and its densities are proportional to e^{-V}, with C², uniformly positive densities on the torus.
    Used in Lemma 3.1 and Appendix B.1; follows from f∈C² on the compact torus and [DEFRG20], but is the base state around which all coercivity is proved.
  • standard math Second variation formula (B.4) and the integration-by-parts identity (B.6) for Gibbs measures.
    Invoked in Lemma 3.1; derived in Lemmas A.21 and A.23 using [Vil08, CMV03, BGL13].
  • standard math Spectral compactness / Poincaré inequality for the weighted Laplacian; Holley–Stroock perturbation bound.
    Used to upgrade non-negativity of the Hessian to a uniform spectral gap λ_gap > 0; Remark 3.2 gives the Poincaré-inequality route via [BGL13].
  • domain assumption Sobolev stability of optimal transport maps (Lemma A.20) and parabolic smoothing/interpolation (Lemmas 3.6–3.7).
    Needed to bootstrap W₂-smallness at t=0 into W^{1,p}-smallness at t≥1; proofs rely on the torus geometry and uniform density bounds.

pith-pipeline@v1.3.0-alltime-deepseek · 36233 in / 13532 out tokens · 142805 ms · 2026-08-03T05:36:33.784825+00:00 · methodology

0 comments
Cite this review

Pith. "Pith review of Local exponential stability of mean-field Langevin descent-ascent and associated particle system." pith.science (2026). https://pith.science/paper/WKYHCY3W

@misc{pith2026260201564,
  author       = {Pith},
  title        = {Pith review of: Local exponential stability of mean-field Langevin descent-ascent and associated particle system},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WKYHCY3W}},
  note         = {Machine review of arXiv:2602.01564}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We study the mean-field Langevin descent-ascent (MFL-DA), a coupled optimization dynamics on the space of probability measures for entropically regularized two-player zero-sum games, together with its associated interacting particle system. For general nonconvex-nonconcave payoffs, Wang and Chizat (COLT 2024) asked whether the original single-timescale MFL-DA converges to the mixed Nash equilibrium and, if so, at what rate. We prove a local affirmative answer in Wasserstein space: if the initial datum is sufficiently close to the mixed Nash equilibrium, then the mean-field dynamics converges to it exponentially fast at a quantitative rate. We further show that the finite-$N$ particle system inherits this stability up to times exponential in $N$, with an $N$-independent exponential rate modulo a finite-particle error floor. Combined with the recent counterexample of Mourrat and Pillaud-Vivien for MFL-DA, which shows that global convergence cannot hold in general, our theorem completes the positive local counterpart of the Wang-Chizat question: the mixed Nash equilibrium has a robust basin of attraction, stable under both the mean-field flow and its finite-particle approximation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PL conditions do not guarantee convergence of gradient descent-ascent dynamics

    math.OC 2026-02 unverdicted novelty 8.0

    There exists a function satisfying two-sided PL conditions for which GDA dynamics fail to converge by circling the saddle point.

  2. Oscillating solutions to the mean-field Langevin descent-ascent flow

    math.OC 2026-04 unverdicted novelty 7.0

    For certain double-well payoff functions with strong coupling and small noise, the mean-field Langevin descent-ascent flow stays near a limit cycle and fails to converge.

Reference graph

Works this paper leans on

56 extracted references · 5 linked inside Pith · cited by 2 Pith papers

  1. [1]

    Wasserstein generative adversarial networks

    Martin Arjovsky, Soumith Chintala, and L \'e on Bottou. Wasserstein generative adversarial networks. In International conference on machine learning , pages 214--223. PMLR, 2017

  2. [2]

    u rich. Birkh \

    Luigi Ambrosio, Nicola Gigli, and Giuseppe Savar \'e . Gradient Flows: In Metric Spaces and in the Space of Probability Measures . Lectures in Mathematics. ETH Z \"u rich. Birkh \"a user Basel, 2 edition, 2008

  3. [3]

    Convergence of two-timescale gradient descent ascent dynamics: finite-dimensional and mean-field perspectives

    Jing An and Jianfeng Lu. Convergence of two-timescale gradient descent ascent dynamics: finite-dimensional and mean-field perspectives. arXiv preprint arXiv:2501.17122 , 2025

  4. [4]

    Analysis and geometry of Markov diffusion operators , volume 348

    Dominique Bakry, Ivan Gentil, and Michel Ledoux. Analysis and geometry of Markov diffusion operators , volume 348. Springer Science & Business Media, 2013

  5. [5]

    V.S. Borkar. Stochastic Approximation: A Dynamical Systems Viewpoint . Cambridge University Press, 2008

  6. [6]

    Functional Analysis, Sobolev Spaces and Partial Differential Equations

    Haim Brezis. Functional Analysis, Sobolev Spaces and Partial Differential Equations . Universitext. Springer, New York, 2011

  7. [7]

    On the global convergence of gradient descent for over-parameterized models using optimal transport

    Lenaic Chizat and Francis Bach. On the global convergence of gradient descent for over-parameterized models using optimal transport. Advances in neural information processing systems , 31, 2018

  8. [8]

    Mean-field langevin dynamics: Exponential convergence and annealing

    L \'e na \" c Chizat. Mean-field langevin dynamics: Exponential convergence and annealing. Transactions on Machine Learning Research , 2022

  9. [9]

    Coupled wasserstein gradient flows for min-max and cooperative games

    Lauren Conger, Franca Hoffmann, Eric Mazumdar, and Lillian J Ratliff. Coupled wasserstein gradient flows for min-max and cooperative games. arXiv preprint arXiv:2411.07403 , 2024

  10. [10]

    Wasserstein- ojasiewicz inequalities and asymptotics of mckean-vlasov equation

    Beomjun Choi, Seunghoon Jeong, and Geuntaek Seo. Wasserstein- ojasiewicz inequalities and asymptotics of mckean-vlasov equation. arXiv preprint arXiv:2511.23361 , 2025

  11. [11]

    Carrillo, Robert J

    Jos\'e A. Carrillo, Robert J. McCann, and C\'edric Villani. Kinetic equilibration rates for granular media and related equations: entropy dissipation and mass transportation estimates. Rev. Mat. Iberoamericana , 19(3):971--1018, 2003

  12. [12]

    On the convergence of min-max langevin dynamics and algorithm

    Yang Cai, Siddharth Mitra, Xiuyuan Wang, and Andre Wibisono. On the convergence of min-max langevin dynamics and algorithm. In Proceedings of Thirty Eighth Conference on Learning Theory , volume 291 of PMLR , pages 677--754, 2025

  13. [13]

    A mean-field analysis of two-player zero-sum games

    Carlos Domingo-Enrich, Kilian Fatras, Nicolas Roux, and Gauthier Gidel. A mean-field analysis of two-player zero-sum games. Advances in Neural Information Processing Systems , 33:2020--2031, 2020

  14. [14]

    Convergence rates of two-time-scale gradient descent-ascent dynamics for solving nonconvex min-max problems

    Thinh Doan. Convergence rates of two-time-scale gradient descent-ascent dynamics for solving nonconvex min-max problems. In Learning for Dynamics and Control Conference , pages 192--206. PMLR, 2022

  15. [15]

    Reducing exit-times of diffusions with repulsive interactions

    Paul-Eric Chaudru de Raynal, Manh Hong Duong, Pierre Monmarch \'e , Milica Toma s evi \'c , and Julian Tugaut. Reducing exit-times of diffusions with repulsive interactions. ESAIM: Probability and Statistics , 27:723--748, 2023

  16. [16]

    The complexity of constrained min-max optimization

    Constantinos Daskalakis, Stratis Skoulakis, and Manolis Zampetakis. The complexity of constrained min-max optimization. In Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing , pages 1466--1478, 2021

  17. [17]

    Quantitative harris-type theorems for diffusions and mckean--vlasov processes

    Andreas Eberle, Arnaud Guillin, and Raphael Zimmer. Quantitative harris-type theorems for diffusions and mckean--vlasov processes. Transactions of the American Mathematical Society , 371(10):7135--7173, 2019

  18. [18]

    Lawrence C. Evans. Partial Differential Equations , volume 19 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 2 edition, 2010

  19. [19]

    Do gans always have nash equilibria? In International Conference on Machine Learning , pages 3029--3039

    Farzan Farnia and Asuman Ozdaglar. Do gans always have nash equilibria? In International Conference on Machine Learning , pages 3029--3039. PMLR, 2020

  20. [20]

    A certain class of diffusion processes associated with nonlinear parabolic equations

    Tadahisa Funaki. A certain class of diffusion processes associated with nonlinear parabolic equations. Zeitschrift f \"u r Wahrscheinlichkeitstheorie und Verwandte Gebiete , 67(3):331--348, 1984

  21. [21]

    I. L. Glicksberg. A further generalization of the kakutani fixed point theorem, with application to nash equilibrium points. Proceedings of the American Mathematical Society , 3(1):170--174, 1952

  22. [22]

    Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

    Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems 27 , pages 2672--2680, 2014

  23. [23]

    Gilbarg and N

    D. Gilbarg and N. S. Trudinger. Elliptic partial differential equations of second order . Classics in Mathematics. Springer-Verlag, Berlin, 2001. Reprint of the 1998 edition

  24. [24]

    Exponential ergodicity in relative entropy and l^ 2 -wasserstein distance for non-equilibrium partially dissipative kinetic sdes

    Xing Huang, Eva Kopfer, Pierre Monmarch \'e , and Panpan Ren. Exponential ergodicity in relative entropy and l^ 2 -wasserstein distance for non-equilibrium partially dissipative kinetic sdes. arXiv preprint arXiv:2507.02518 , 2025

  25. [25]

    Finding mixed N ash equilibria of generative adversarial networks

    Ya-Ping Hsieh, Chen Liu, and Volkan Cevher. Finding mixed N ash equilibria of generative adversarial networks. In Proceedings of the 36th International Conference on Machine Learning , volume 97 of PMLR , pages 2810--2819, 2019

  26. [26]

    Mean-field langevin dynamics and energy landscape of neural networks

    Kaitong Hu, Zhenjie Ren, David S i s ka, and ukasz Szpruch. Mean-field langevin dynamics and energy landscape of neural networks. Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , 57(4):2043--2065, 2021

  27. [27]

    On gradient descent-ascent flows in metric spaces

    Noboru Isobe and Sho Shimoyama. On gradient descent-ascent flows in metric spaces. arXiv preprint arXiv:2506.20258 , 2025

  28. [28]

    What is local optimality in nonconvex-nonconcave minimax optimization? In International conference on machine learning , pages 4880--4889

    Chi Jin, Praneeth Netrapalli, and Michael Jordan. What is local optimality in nonconvex-nonconcave minimax optimization? In International conference on machine learning , pages 4880--4889. PMLR, 2020

  29. [29]

    Symmetric mean-field langevin dynamics for minimax optimization

    Minyoung Kim, Juho Lee, and Taiji Suzuki. Symmetric mean-field langevin dynamics for minimax optimization. In International Conference on Learning Representations , 2024

  30. [30]

    N. V. Krylov. Lectures on Elliptic and Parabolic Equations in Sobolev Spaces , volume 96 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 2008

  31. [31]

    Transformers learn nonlinear features in context: nonconvex mean-field dynamics on the attention landscape

    Juno Kim and Taiji Suzuki. Transformers learn nonlinear features in context: nonconvex mean-field dynamics on the attention landscape. In Proceedings of the 41st International Conference on Machine Learning , pages 24527--24561, 2024

  32. [32]

    Second order parabolic differential equations

    Gary M Lieberman. Second order parabolic differential equations . World scientific, 1996

  33. [33]

    Convergence of mean-field L angevin stochastic descent-ascent for distributional minimax optimization

    Zhangyi Liu, Feng Liu, Rui Gao, and Shuang Li. Convergence of mean-field L angevin stochastic descent-ascent for distributional minimax optimization. In Proceedings of the 42nd International Conference on Machine Learning , volume 267 of PMLR , pages 38869--38893, 13--19 Jul 2025

  34. [34]

    Convergence of time-averaged mean field gradient descent dynamics for continuous multi-player zero-sum games

    Yulong Lu and Pierre Monmarch \'e . Convergence of time-averaged mean field gradient descent dynamics for continuous multi-player zero-sum games. arXiv preprint arXiv:2505.07642 , 2025

  35. [35]

    Some geometric calculations on W asserstein space

    John Lott. Some geometric calculations on W asserstein space. Communications in Mathematical Physics , 277:423--437, 2008

  36. [36]

    Two-scale gradient descent ascent dynamics finds mixed N ash equilibria of continuous games: A mean-field perspective

    Yulong Lu. Two-scale gradient descent ascent dynamics finds mixed N ash equilibria of continuous games: A mean-field perspective. In International Conference on Machine Learning , pages 22790--22811. PMLR, 2023

  37. [37]

    A convexity principle for interacting gases

    Robert J McCann. A convexity principle for interacting gases. Advances in mathematics , 128(1):153--179, 1997

  38. [38]

    S. Mei, A. Montanari, and P.-M. Nguyen. A mean field view of the landscape of two-layer neural networks. Proceedings of the National Academy of Sciences , 115(33):E7665--E7671, 2018

  39. [39]

    The numerics of gans

    Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. The numerics of gans. Advances in neural information processing systems , 30, 2017

  40. [40]

    Local convergence rates for wasserstein gradient flows and mckean-vlasov equations with multiple stationary solutions

    Pierre Monmarch \'e and Julien Reygner. Local convergence rates for wasserstein gradient flows and mckean-vlasov equations with multiple stationary solutions. Probability Theory and Related Fields , 2025

  41. [41]

    Learning in games via reinforcement and regularization

    Panayotis Mertikopoulos and William H Sandholm. Learning in games via reinforcement and regularization. Mathematics of Operations Research , 41(4):1297--1324, 2016

  42. [42]

    On finding local nash equilibria (and only local nash equilibria) in zero-sum games

    Eric Mazumdar, S Shankar Sastry, and Michael I Jordan. On finding local nash equilibria (and only local nash equilibria) in zero-sum games. ACM/IMS Journal of Data Science , 2(2):1--26, 2025

  43. [43]

    Provably convergent quasistatic dynamics for mean-field two-player zero-sum games

    Chao Ma and Lexing Ying. Provably convergent quasistatic dynamics for mean-field two-player zero-sum games. arXiv preprint arXiv:2202.10947 , 2022

  44. [44]

    On elliptic partial differential equations

    Louis Nirenberg. On elliptic partial differential equations. Annali della Scuola Normale Superiore di Pisa-Scienze Fisiche e Matematiche , 13(2):115--162, 1959

  45. [45]

    Convex analysis of the mean field langevin dynamics

    Atsushi Nitanda, Denny Wu, and Taiji Suzuki. Convex analysis of the mean field langevin dynamics. In International Conference on Artificial Intelligence and Statistics , pages 9741--9757. PMLR, 2022

  46. [46]

    The geometry of dissipative evolution equations: the porous medium equation

    Felix Otto. The geometry of dissipative evolution equations: the porous medium equation. Comm. Partial Differential Equations , 26(1-2):101--174, 2001

  47. [47]

    Comparison between w2 distance and \.h - 1 norm, and localization of wasserstein distance

    R \'e mi Peyre. Comparison between w2 distance and \.h - 1 norm, and localization of wasserstein distance. ESAIM: Control, Optimisation and Calculus of Variations , 24(4):1489--1501, 2018

  48. [48]

    The Fokker-Planck Equation: Methods of Solution and Applications

    Hannes Risken. The Fokker-Planck Equation: Methods of Solution and Applications . Springer Series in Synergetics. Springer, Berlin; New York, 2nd edition, 1989

  49. [49]

    Population games and evolutionary dynamics

    William H Sandholm. Population games and evolutionary dynamics . MIT press, 2010

  50. [50]

    Convergence of mean-field langevin dynamics: time-space discretization, stochastic gradient, and variance reduction

    Taiji Suzuki, Denny Wu, and Atsushi Nitanda. Convergence of mean-field langevin dynamics: time-space discretization, stochastic gradient, and variance reduction. Advances in Neural Information Processing Systems , 36:15545--15577, 2023

  51. [51]

    Well-posedness of multidimensional diffusion processes with weakly differentiable coefficients

    Dario Trevisan. Well-posedness of multidimensional diffusion processes with weakly differentiable coefficients. Electronic Journal of Probability , 21(22):1--41, 2016

  52. [52]

    C. Villani. Optimal Transport: Old and New . Grundlehren der mathematischen Wissenschaften. Springer Berlin Heidelberg, 2008

  53. [53]

    Open problem: Convergence of single-timescale mean-field langevin descent-ascent for two-player zero-sum games

    Guillaume Wang and L \'e na\" i c Chizat. Open problem: Convergence of single-timescale mean-field langevin descent-ascent for two-player zero-sum games. In Proceedings of Thirty Seventh Conference on Learning Theory , volume 247 of PMLR , pages 5345--5350, 2024

  54. [54]

    An exponentially converging particle method for the mixed nash equilibrium of continuous games

    Guillaume Wang and L \'e na \" c Chizat. An exponentially converging particle method for the mixed nash equilibrium of continuous games. Open Journal of Mathematical Optimization , 6:1--66, January 2025

  55. [55]

    Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems

    Junchi Yang, Negar Kiyavash, and Niao He. Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems. Advances in neural information processing systems , 33:1153--1165, 2020

  56. [56]

    Optimal epoch stochastic gradient descent ascent methods for min-max optimization

    Yan Yan, Yi Xu, Qihang Lin, Wei Liu, and Tianbao Yang. Optimal epoch stochastic gradient descent ascent methods for min-max optimization. Advances in Neural Information Processing Systems , 33:5789--5800, 2020