REVIEW 3 major objections 2 minor 88 references
Solving Infinite-Player Games with Player-to-Strategy Networks
T0 review · 3 major / 2 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Neural networks that map players to strategies can find approximate Nash equilibria in games with infinitely many players.
desk verdict Useful representation and training rule for continuum-player games, but the empirical convergence claim is undermined by an unstated grid resolution and an untested blind spot for opposing-interest games. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The two load-bearing objects are the Player-to-Strategy Network (P2SN), a neural network that maps a player's feature vector (plus optional observation and noise inputs) to that player's strategy, and the Shared-Parameter Simultaneous Gradient (SPSG), the functional derivative in Equation 5. The P2SN makes the infinite strategy profile a finite set of shared parameters; a learnable Fourier feature map on player coordinates lets it represent high-frequency structure on low-dimensional player spaces that plain networks underfit. SPSG estimates the functional derivative by sampling one player per gradient call, forming the hybrid profile $s[i\mapsto r(i)]$, and backpropagating through the sampled player's output; because parameters are shared, the update is a stochastic estimate of the whole-profile derivative rather than a per-player gradient. For non-differentiable utilities the paper also adopts pseudo-gradient estimators, which perturb strategies with Gaussian noise and differentiate a smoothed utility, and noise-injected network inputs to represent mixed strategies over continuous action sets.
What would settle it
Run SPSG on a two-player zero-sum game (or a continuum analogue, such as two symmetric populations with opposing payoffs) with a single shared network: footnote 10 predicts the parameters do not move and regret never decreases. More directly, take one of the test games with a known closed-form equilibrium, compute the learned profile's regret on a much finer grid or exactly, and check whether the reported near-zero mean regret survives; if a narrow deviation outside the evaluation grid yields high utility, the observed convergence is an artifact of discretization.
Extended reading notes
Core claim
The paper's central claim is that a strategy profile for countably or uncountably many players can be carried by a single Player-to-Strategy Network, and that the update rule SPSG drives this network toward approximate Nash equilibrium. SPSG replaces the per-player gradient with the functional derivative $v(s) = \frac{d}{dr}\int_{i\sim\mu} u(s[i\mapsto r(i)], i)\,|_{r=s}$, effectively differentiating the mean utility through only one player at a time while holding the rest of the profile fixed; a Monte Carlo sample of players gives an unbiased estimator of the integral, and differentiating through the hybrid profile $s[i\mapsto r(i)]$ gives the gradient. The paper reports convergence of mean regret toward zero in one- and two-dimensional Ising games, distance-based variants of each, Cournot competition with a global demand function, Cournot competition with a local demand function, and a crowding game, with eight trials per experiment and training times of 9 to 14 minutes per run on one GPU.
Load-bearing premise
The load-bearing premise is that repeatedly sampling one player, differentiating that player's utility through the shared network, and stepping along the averaged result moves the whole infinite-player profile toward low regret—an assumption that is not proven, and that the paper's own footnote 10 shows fails when shared parameters face exactly opposing pulls.
Editorial extensions
If this is right
- If the reported convergence holds, continuum-player games can be solved without symmetry or mean-field assumptions, so boundary effects and heterogeneous player features are handled automatically.
- The method applies in principle to games with infinitely many players, states, and actions, including mixed strategies on continuous action spaces and utility functions with discontinuities.
- Because SPSG generalizes simultaneous gradient ascent, existing variants such as optimistic updates and pseudo-gradients can be swapped in as the optimizer inside the same P2SN representation.
- For games whose utilities decompose into pairwise or higher-order interactions as in Equation 6, unbiased utility estimates can be obtained by sampling the interacting players, making the method feasible for interaction-driven models.
- The paper's regret curves constitute evidence that approximate Nash equilibria are reachable in the tested families, including Ising games, Cournot competition, local Cournot competition, and crowding games.
Reading between the lines
- Because the convergence is empirical and the paper itself notes that shared parameters freeze in a two-player zero-sum game, an implicit scope restriction is that the games must not have exactly opposing players pulling shared parameters in opposite directions; a natural extension would partition parameters per role or use symplectic or optimistic updates for adversarial components.
- The regret certification used in the experiments evaluates best responses on a finite grid, which lower-bounds true regret; if a narrow profitable deviation falls between grid points, reported convergence could overstate equilibrium quality, so a direct test is to compare learned profiles against known analytic equilibria in the same games.
- If SPSG is an unbiased functional gradient, then the method should extend beyond pairwise aggregative utilities to games with three-way interactions and to continuous-time dynamics; testing on a game with a known closed-form asymmetric equilibrium would separate representation error from optimization error.
- A practical consequence left implicit by the authors is that the same player-to-strategy network could be reused for equilibrium selection and mechanism design: once a social planner can cheaply evaluate equilibria of a continuum game, they could search over game parameters to steer the equilibrium.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a Player-to-Strategy Network (P2SN), a neural network mapping player features to strategies, and Shared-Parameter Simultaneous Gradient (SPSG), an equilibrium-seeking algorithm that trains the network by estimating the gradient of the integrated simultaneous-gradient vector field. The method is tested on five infinite-player games (Ising, distance-based Ising, Cournot, local Cournot, crowding), and the authors report that regrets decrease toward zero, which they interpret as convergence to approximate Nash equilibria. The paper claims the approach handles infinitely many players, states, and actions, with no symmetry assumptions.
Significance. If the claims hold, the paper offers a genuinely new way to compute approximate Nash equilibria in continuum-player games without symmetry or mean-field structure, using a finite-dimensional neural network representation. The idea of treating the strategy profile as a function learned by a network is natural and useful, and the experiments span several nontrivial games. The paper is clearly written and provides a useful formalization of regret and exploitability in this setting. However, because the empirical evaluation relies on a lower-bound regret estimator with unspecified resolution and no external validation, and because the central unbiasedness claim is not justified, the significance of the contribution is not yet established.
major comments (3)
- [§6] The reported regrets are computed by discretizing the player space and each strategy space into N points, but the value of N is never stated. Since a finite-grid best response is a lower bound on the true supremum over the full strategy set, and a finite player grid can miss high-regret players, the plotted regrets in Figures 3, 6, 8, 10, 12, 14, and 15 are lower bounds on the true mean regret. Without specifying N or demonstrating that the regrets are insensitive to an increase in N, the central empirical claim of convergence to approximate Nash equilibria is not established. This is particularly concerning for the high-frequency bias fields in §6.2 and §6.5, where a coarse grid could miss narrow basins of large unilateral gains.
- [§5, Eq. (5)] The paper claims that the player-sampling procedure yields an unbiased estimator of the SPSG because 'the integral and derivative commute with each other' and offers no justification. For a functional derivative of an integral with respect to a function, this interchange requires regularity conditions that are not stated and may fail for the discontinuous utility functions the paper claims to handle. Additionally, footnote 10 states that in a two-player zero-sum game with shared parameters the parameters do not change at all, so the method excludes adversarial settings; this directly contradicts the unqualified abstract claim that the approach can handle infinite-player games. The authors should state sufficient conditions for Eq. (5) and restrict the applicability claims accordingly.
- [§6] The experimental validation contains no comparison against known equilibria or existing mean-field-game solvers. The only evidence is the internal regret metric, which uses the same finite-grid approximation that the method itself is being judged against. For example, in the 1D Ising game of §6.1, the best response is analytically computable (s(i) = sign(b(i) + ∫ s(j) dν(j))), so the learned profile could be checked directly against an exact best response. Without such an independent check, the plotted decreasing regrets do not demonstrate that the learned strategies are close to Nash equilibria in the true infinite-player game.
minor comments (2)
- [§6] The symbol N is used for two different things: the number of discretization points for the player and strategy grids, and the number of Monte Carlo utility samples (N = 200). The grid sizes for the player and strategy discretizations should be given explicitly and consistently.
- [§5, footnote 11] Footnote 11 assumes 0 < μ(I) < ∞, which excludes the counting measure on a countably infinite player set. The abstract's claim that the method handles countably infinite players should be qualified, since the estimator as written applies to normalizable measures.
Circularity Check
No significant circularity: SPSG and the regret certificate are independently defined, and the minor self-citations are not load-bearing.
full rationale
I walked the derivation chain from Equation (5) through the SPSG update and the evaluation protocol. Equation (5) is a definition of the functional derivative of aggregate utility with respect to the strategy profile; it does not presuppose the Nash equilibrium condition, and regret E(s) is defined independently via best responses to the current profile. No fitted parameter is later repackaged as a prediction: the paper states that game constants 'were chosen (before training)', and the networks are not trained to match any target equilibrium profile. The self-citations to Martin and Sandholm (2023, 2024) concern randomized policy networks and joint-perturbation pseudo-gradients; neither component is required by the differentiable experiments in Section 6, so these citations are not load-bearing for the central convergence observation. The finite-grid regret estimate in Section 6 is a lower bound on true regret and may overstate convergence, but that is a validation weakness rather than circularity, because the observed decrease is not forced by construction. The paper itself concedes that convergence theory is left to future work and notes a shared-parameter freeze in opposing-player games; those are limitations, not circular steps. Overall, no circular step is present; the score of 2 reflects only the existence of minor self-citations that do not carry the paper's central claim.
Assumptions & free parameters
free parameters (3)
- learning rate and optimizer schedule
- Fourier feature dimension N=64 =
64
- Fourier initialization scale sigma=64 =
64
assumptions (5)
- domain assumption Pure-strategy Nash equilibria exist in all test games (Glicksberg and Dasgupta-Maskin conditions).
- domain assumption The P2SN is expressive enough to approximate the equilibrium strategy profile (universal approximation).
- domain assumption Integration and differentiation commute, and the one-player-per-iteration Monte Carlo estimator of Eq. (5) is unbiased.
- domain assumption The measure mu is normalizable, with 0 < mu(I) < infinity.
- ad hoc to paper Finite-grid best responses and N=200 Monte Carlo utility samples give a faithful approximation of true regret.
Cite this review
Pith. "Pith review of Solving Infinite-Player Games with Player-to-Strategy Networks." pith.science (2026). https://pith.science/paper/HUPPV3MZ
@misc{pith2026250109330,
author = {Pith},
title = {Pith review of: Solving Infinite-Player Games with Player-to-Strategy Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/HUPPV3MZ}},
note = {Machine review of arXiv:2501.09330}
}
read the original abstract
We present a new approach to solving games with a countably or uncountably infinite number of players. Such games are often used to model multiagent systems with a large number of agents. The latter are frequently encountered in economics, financial markets, crowd dynamics, congestion analysis, epidemiology, and population ecology, among other fields. Our two primary contributions are as follows. First, we present a way to represent strategy profiles for an infinite number of players, which we name a Player-to-Strategy Network (P2SN). Such a network maps players to strategies, and exploits the generalization capabilities of neural networks to learn across an infinite number of inputs (players) simultaneously. Second, we present an algorithm, which we name Shared-Parameter Simultaneous Gradient (SPSG), for training such a network, with the goal of finding an approximate Nash equilibrium. This algorithm generalizes simultaneous gradient ascent and its variants, which are classical equilibrium-seeking dynamics used for multiagent reinforcement learning. We test our approach on infinite-player games and observe its convergence to approximate Nash equilibria. Our method can handle games with infinitely many states, infinitely many players, infinitely many actions (and mixed strategies on them), and discontinuous utility functions.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Robert J. Aumann. Markets with a continuum of traders. Econometrica, 32 0 (1): 0 39--50, 1964
1964
-
[2]
The mechanics of n-player differentiable games
David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel. The mechanics of n-player differentiable games. In International Conference on Machine Learning (ICML), 2018
2018
-
[3]
Berahas, Liyuan Cao, Krzysztof Choromanski, and Katya Scheinberg
Albert S. Berahas, Liyuan Cao, Krzysztof Choromanski, and Katya Scheinberg. A theoretical and empirical comparison of gradient approximations in derivative-free optimization. Foundations of Computational Mathematics, 22 0 (2): 0 507--560, 2022
2022
-
[4]
Learning equilibria in symmetric auction games using artificial neural networks
Martin Bichler, Maximilian Fichtl, Stefan Heidekr \"u ger, Nils Kohring, and Paul Sutterer. Learning equilibria in symmetric auction games using artificial neural networks. Nature Machine Intelligence, 3 0 (8): 0 687--695, 2021
2021
-
[5]
JAX : Composable transformations of Python + NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, et al. JAX : Composable transformations of Python + NumPy programs, 2018
2018
-
[6]
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm. Superhuman AI for multiplayer poker. Science, 365 0 (6456): 0 885--890, 2019
2019
-
[7]
Peter E. Caines. Mean field games. In John Baillieul and Tariq Samad, editors, Encyclopedia of Systems and Control. Springer London, Gewerbestrasse 11, 6330 Cham, Switzerland, 2013
2013
-
[8]
Caines, Minyi Huang, and Roland P
Peter E. Caines, Minyi Huang, and Roland P. Malham \'e . Mean field games. In Tamer Ba s ar and Georges Zaccour, editors, Handbook of Dynamic Game Theory. Springer International Publishing, Gewerbestrasse 11, 6330 Cham, Switzerland, 2018
2018
Show all 88 references
-
[9]
The dyson and coulomb games
Ren \'e Carmona, Mark Cerenzia, and Aaron Zeff Palmer. The dyson and coulomb games. In Annales Henri Poincar \'e , volume 21, pages 2897--2949. Springer, 2020
2020
-
[10]
Bertrand and Cournot mean field games
Patrick Chan and Ronnie Sircar. Bertrand and Cournot mean field games. Applied Mathematics & Optimization, 71 0 (3): 0 533--569, 2015
2015
-
[11]
Principes de la th \'e orie des richesses
Antoine Augustin Cournot. Principes de la th \'e orie des richesses . Hachette, 1863
-
[12]
Approximation by superpositions of a sigmoidal function
George Cybenko. Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems, 2 0 (4): 0 303--314, 1989
1989
-
[13]
The existence of equilibrium in discontinuous economic games 1: Theory
Partha Dasgupta and Eric Maskin. The existence of equilibrium in discontinuous economic games 1: Theory. Review of Economic Studies, 53: 0 1--26, 1986
1986
-
[14]
Training GAN s with optimism
Constantinos Daskalakis, Andrew Ilyas, Vasilis Syrgkanis, and Haoyang Zeng. Training GAN s with optimism. In International Conference on Learning Representations (ICLR), 2018
2018
-
[15]
A social equilibrium existence theorem
Gerard Debreu. A social equilibrium existence theorem. Proceedings of the National Academy of Sciences, 38 0 (10), 1952
1952
-
[16]
The D eep M ind JAX E cosystem, 2020
DeepMind, Igor Babuschkin, Kate Baumli, Alison Bell, Surya Bhupatiraju, Jake Bruce, Peter Buchlovsky, David Budden, Trevor Cai, Aidan Clark, Ivo Danihelka, Antoine Dedieu, Claudio Fantacci, Jonathan Godwin, Chris Jones, Ross Hemsley, Tom Hennigan, Matteo Hessel, Shaobo Hou, St...
2020
-
[17]
Duchi, Michael I
John C. Duchi, Michael I. Jordan, Martin J. Wainwright, and Andre Wibisono. Optimal rates for zero-order convex optimization: the power of two function evaluations. IEEE Transactions on Information Theory, 61 0 (5), 2015
2015
-
[18]
Beitrag zur theorie des ferromagnetismus
Ising Ernst. Beitrag zur theorie des ferromagnetismus. Zeitschrift f \"u r Physik A Hadrons and Nuclei , 31 0 (1): 0 253--258, 1925
1925
-
[19]
Fixed point and minimax theorems in locally convex topological linear spaces
Ky Fan. Fixed point and minimax theorems in locally convex topological linear spaces. Proceedings of the National Academy of Sciences, 38 0 (2): 0 121--126, 1952
1952
-
[20]
Feldman, Inwon C
William M. Feldman, Inwon C. Kim, and Aaron Zeff Palmer. The sharp interface limit of an Ising game. ESAIM: Control, Optimisation and Calculus of Variations, 30: 0 35, 2024
2024
-
[21]
Ising model versus normal form game
Serge Galam and Bernard Walliser. Ising model versus normal form game. Physica A: Statistical Mechanics and its Applications, 389 0 (3): 0 481--489, 2010
2010
-
[22]
Computing equilibria by incorporating qualitative models
Sam Ganzfried and Tuomas Sandholm. Computing equilibria by incorporating qualitative models. In International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS), 2010
2010
-
[23]
A further generalization of the K akutani fixed point theorem, with application to N ash equilibrium points
Irving Leonard Glicksberg. A further generalization of the K akutani fixed point theorem, with application to N ash equilibrium points. Proceedings of the American Mathematical Society, 3 0 (1): 0 170--174, 1952
1952
-
[24]
Monte Carlo sampling methods using Markov chains and their applications
Wilfred Keith Hastings. Monte Carlo sampling methods using Markov chains and their applications. Biometrika, 57 0 (1): 0 97--109, 1970
1970
-
[25]
Delving deep into rectifiers: Surpassing human-level performance on I mage N et classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on I mage N et classification. In International Conference on Computer Vision (ICCV), 2015
2015
-
[26]
Flax : A neural network library and ecosystem for JAX , 2023
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Z ee. Flax : A neural network library and ecosystem for JAX , 2023. URL http://github.com/google/flax
2023
-
[27]
Approximation capabilities of multilayer feedforward networks
Kurt Hornik. Approximation capabilities of multilayer feedforward networks. Neural networks, 4 0 (2): 0 251--257, 1991
1991
-
[28]
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White. Multilayer feedforward networks are universal approximators. Neural networks, 2 0 (5): 0 359--366, 1989
1989
-
[29]
The limits of min-max optimization algorithms: Convergence to spurious non-critical sets
Ya-Ping Hsieh, Panayotis Mertikopoulos, and Volkan Cevher. The limits of min-max optimization algorithms: Convergence to spurious non-critical sets. In International Conference on Machine Learning (ICML), 2021
2021
-
[30]
On the convergence of single-call stochastic extra-gradient methods
Yu-Guan Hsieh, Franck Iutzeler, J \'e r \^o me Malick, and Panayotis Mertikopoulos. On the convergence of single-call stochastic extra-gradient methods. In Conference on Neural Information Processing Systems (NeurIPS), 2019
2019
-
[31]
John D. Hunter. Matplotlib : A 2D graphics environment. Computing in Science & Engineering, 9 0 (3): 0 90--95, 2007
2007
-
[32]
Self-gated rectified linear unit for performance improvement of deep neural networks
Israt Jahan, Md Faisal Ahmed, Md Osman Ali, and Yeong Min Jang. Self-gated rectified linear unit for performance improvement of deep neural networks. ICT Express, 9 0 (3): 0 320--325, 2023
2023
-
[33]
Ali Khan
M. Ali Khan. Equilibrium points of nonatomic games over a nonreflexive Banach space. Journal of Approximation Theory, 43 0 (4): 0 370--376, 1985
1985
-
[34]
Ali Khan
M. Ali Khan. Equilibrium points of nonatomic games over a Banach space. Transactions of the American Mathematical Society, 293 0 (2): 0 737--749, 1986
1986
-
[35]
Ali Khan and Nikolaos S
M. Ali Khan and Nikolaos S. Papageorgiou. On Cournot -- Nash equilibria in generalized qualitative games with an atomless measure space of agents. Proceedings of the American Mathematical Society, 100 0 (3): 0 505--510, 1987 a
1987
-
[36]
Ali Khan and Nikolaos S
M. Ali Khan and Nikolaos S. Papageorgiou. On Cournot -- Nash equilibria in generalized qualitative games with a continuum of players. Nonlinear Analysis: Theory, Methods & Applications, 11 0 (6): 0 741--756, 1987 b
1987
-
[37]
Ali Khan and Yeneng Sun
M. Ali Khan and Yeneng Sun. Non-cooperative games with many players. Handbook of Game Theory with Economic Applications, 3: 0 1761--1808, 2002
2002
-
[38]
Yannelis
Taesung Kim, Karel Prikry, and Nicholas C. Yannelis. Equilibria in abstract economies with a measure space of agents and with an infinite dimensional strategy space. Journal of Approximation Theory, 56 0 (3): 0 256--266, 1989
1989
-
[39]
Auction Theory
Vijay Krishna. Auction Theory. Academic Press, 2002
2002
-
[40]
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien P \'e rolat, David Silver, and Thore Graepel. A unified game-theoretic approach to multiagent reinforcement learning. In Conference on Neural Information Processing Systems (NeurIPS), pag...
2017
-
[41]
Non-differentiable supervised learning with evolution strategies and hybrid methods
Karel Lenc, Erich Elsen, Tom Schaul, and Karen Simonyan. Non-differentiable supervised learning with evolution strategies and hybrid methods. arXiv:1906.03139, 0, 2019
1906 arXiv
-
[42]
Beitr s ge zum verst s ndnis der magnetischen eigenschaften in festen k s rpern
Wilhelm Lenz. Beitr s ge zum verst s ndnis der magnetischen eigenschaften in festen k s rpern. Physikalische Zeitschrift, 21 0 (1): 0 613--615, 1920
1920
-
[43]
Andrey Leonidov, Alexey Savvateev, and Andrew G. Semenov. QRE in the Ising game. In CEUR Workshop, 2020
2020
-
[44]
Andrey Leonidov, Alexey Savvateev, and Andrew G. Semenov. Ising game on graphs. Chaos, Solitons & Fractals, 180: 0 114540, 2024
2024
-
[45]
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno, Vladimir Ya Lin, Allan Pinkus, and Shimon Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural networks, 6 0 (6), 1993
1993
-
[46]
Differentiable game mechanics
Alistair Letcher, David Balduzzi, S \'e bastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel. Differentiable game mechanics. Journal of Machine Learning Research, 20 0 (84): 0 1--40, 2019 a
2019
-
[47]
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rockt \"a schel, and Shimon Whiteson. Stable opponent shaping in differentiable games. In International Conference on Learning Representations (ICLR), 2019 b
2019
-
[48]
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien P \'e rolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls. Computing approximate equilibria in sequential adversarial games by exploitability descent. In Proceedings of the International Joint Conference on Artifi...
2019
-
[49]
Finding mixed-strategy equilibria of continuous-action games without gradients using randomized policy networks
Carlos Martin and Tuomas Sandholm. Finding mixed-strategy equilibria of continuous-action games without gradients using randomized policy networks. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2023
2023
-
[50]
Joint-perturbation simultaneous pseudo-gradient
Carlos Martin and Tuomas Sandholm. Joint-perturbation simultaneous pseudo-gradient. arXiv:2408.09306, 0, 2024
2024
-
[51]
Ratliff, and S
Eric Mazumdar, Lillian J. Ratliff, and S. Shankar Sastry. On gradient-based learning in continuous games. SIAM Journal on Mathematics of Data Science, 2 0 (1): 0 103--131, 2020
2020
-
[52]
Mazumdar, Michael I
Eric V. Mazumdar, Michael I. Jordan, and S. Shankar Sastry. On finding local N ash equilibria (and only local N ash equilibria) in zero-sum games. arXiv:1901.00838, 0, 2019
1901 arXiv
-
[53]
Learning in games with continuous action sets and unknown payoff functions
Panayotis Mertikopoulos and Zhengyuan Zhou. Learning in games with continuous action sets and unknown payoff functions. Mathematical Programming, 173: 0 465--507, 2019
2019
-
[54]
Equation of state calculations by fast computing machines
Nicholas Metropolis, Arianna Rosenbluth, Marshall Rosenbluth, Augusta Teller, and Edward Teller. Equation of state calculations by fast computing machines. The Journal of Chemical Physics, 21 0 (6): 0 1087--1092, 1953
1953
-
[55]
Generic Uniqueness of Equilibria in Nonatomic Congestion Games
Igal Milchtaich. Generic Uniqueness of Equilibria in Nonatomic Congestion Games. Hebrew University of Jerusalem, 1996
1996
-
[56]
Generic uniqueness of equilibrium in large crowding games
Igal Milchtaich. Generic uniqueness of equilibrium in large crowding games. Mathematics of Operations Research, 25 0 (3): 0 349--364, 2000
2000
-
[57]
Topological conditions for uniqueness of equilibrium in networks
Igal Milchtaich. Topological conditions for uniqueness of equilibrium in networks. Mathematics of Operations Research, 30 0 (1): 0 225--244, 2005
2005
-
[58]
Learning equilibria in mean-field games: Introducing mean-field PSRO
Paul Muller, Mark Rowland, Romuald Elie, Georgios Piliouras, Julien Perolat, Mathieu Lauriere, Raphael Marinier, Olivier Pietquin, and Karl Tuyls. Learning equilibria in mean-field games: Introducing mean-field PSRO . In Autonomous Agents and Multi-Agent Systems, page 926–934,...
2022
-
[59]
Equilibrium points in n-person games
John Nash. Equilibrium points in n-person games. Proceedings of the National Academy of Sciences, 36: 0 48--49, 1950
1950
-
[60]
Non-cooperative games
John Nash. Non-cooperative games. Annals of Mathematics, 54: 0 289--295, 1951
1951
-
[61]
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny. Random gradient-free minimization of convex functions. Foundations of Computational Mathematics, 17 0 (2): 0 527--566, 2017
2017
-
[62]
Perfectly competitive markets as the limits of Cournot markets
William Novshek. Perfectly competitive markets as the limits of Cournot markets. Journal of Economic Theory, 35 0 (1): 0 72--82, 1985
1985
-
[63]
Graphon games
Francesca Parise and Asuman Ozdaglar. Graphon games. In Proceedings of the ACM Conference on Economics and Computation (EC), page 457–458, New York, NY, USA, 2019. Association for Computing Machinery. ISBN 9781450367929. doi:10.1145/3328526.3329638. URL https://doi.org/10.1145...
2019
-
[64]
Graphon games: A statistical framework for network games and interventions
Francesca Parise and Asuman Ozdaglar. Graphon games: A statistical framework for network games and interventions. Econometrica, 91 0 (1): 0 191--225, 2023
2023
-
[65]
Scaling mean field games by online mirror descent
Julien P\' e rolat, Sarah Perrin, Romuald Elie, Mathieu Lauri\` e re, Georgios Piliouras, Matthieu Geist, Karl Tuyls, and Olivier Pietquin. Scaling mean field games by online mirror descent. In Autonomous Agents and Multi-Agent Systems, page 1028–1037, Richland, SC, 2022. Inte...
2022
-
[66]
Fictitious play for mean field games: Continuous time analysis and applications
Sarah Perrin, Julien P \'e rolat, Mathieu Lauri \`e re, Matthieu Geist, Romuald Elie, and Olivier Pietquin. Fictitious play for mean field games: Continuous time analysis and applications. Conference on Neural Information Processing Systems (NeurIPS), 33: 0 13199--13213, 2020
2020
-
[67]
Approximation theory of the MLP model in neural networks
Allan Pinkus. Approximation theory of the MLP model in neural networks. Acta numerica, 8: 0 143--195, 1999
1999
-
[68]
A modification of the A rrow- H urwicz method for search of saddle points
Leonid Denisovich Popov. A modification of the A rrow- H urwicz method for search of saddle points. Mathematical notes of the Academy of Sciences of the USSR, 28: 0 845--848, 1980
1980
-
[69]
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville. On the spectral bias of neural networks. In International Conference on Machine Learning (ICML), pages 5301--5310. PMLR, 2019
-
[70]
Kali P. Rath. A direct proof of the existence of pure strategy equilibria in games with a continuum of players. Economic Theory, 2: 0 427--433, 1992
1992
-
[71]
The unreasonable effectiveness of quasirandom sequences, 2018
Martin Roberts. The unreasonable effectiveness of quasirandom sequences, 2018
2018
-
[72]
Ben Rosen
J. Ben Rosen. Existence and uniqueness of equilibrium points for concave n-person games. Econometrica, 33 0 (3): 0 520--534, 1965
1965
-
[73]
An automatic method for finding the greatest or least value of a function
HoHo Rosenbrock. An automatic method for finding the greatest or least value of a function. The computer journal, 3 0 (3): 0 175--184, 1960
1960
-
[74]
Evolution strategies as a scalable alternative to reinforcement learning, 2017
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. Evolution strategies as a scalable alternative to reinforcement learning, 2017
2017
-
[75]
Equilibrium points of nonatomic games
David Schmeidler. Equilibrium points of nonatomic games. Journal of Statistical Physics, 7: 0 295--300, 1973
1973
-
[76]
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir. An optimal algorithm for bandit and zero-order convex optimization with two-point feedback. Journal of Machine Learning Research, 18 0 (52): 0 1--11, 2017
2017
-
[77]
Fourier features let networks learn high frequency functions in low dimensional domains
Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. Conference on Neural Information Pr...
2020
-
[78]
Approximate exploitability: Learning a best response
Finbarr Timbers, Nolan Bard, Edward Lockhart, Marc Lanctot, Martin Schmid, Neil Burch, Julian Schrittwieser, Thomas Hubert, and Michael Bowling. Approximate exploitability: Learning a best response. In Proceedings of the International Joint Conference on Artificial Intelligenc...
2022
-
[79]
Munchausen reinforcement learning
Nino Vieillard, Olivier Pietquin, and Matthieu Geist. Munchausen reinforcement learning. Conference on Neural Information Processing Systems (NeurIPS), 33: 0 4235--4246, 2020
2020
-
[80]
Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J. van der Walt , Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nels...
2020
-
[81]
Multi-agent reinforcement learning in O pen S piel: A reproduction report
Michael Walton and Viliam Lisy. Multi-agent reinforcement learning in O pen S piel: A reproduction report. arXiv:2103.00187, 0, 2021
2021 arXiv
-
[82]
Yongzhao Wang and Michael P. Wellman. Empirical game-theoretic analysis for mean field games. In Autonomous Agents and Multi-Agent Systems, page 1025–1033, Richland, SC, 2023. International Foundation for Autonomous Agents and Multiagent Systems. ISBN 9781450394321
2023
-
[83]
Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun, Jan Peters, and J\" u rgen Schmidhuber. Natural evolution strategies. Journal of Machine Learning Research, 15 0 (1): 0 949--980, 2014
2014
-
[84]
COLA : Consistent learning with opponent-learning awareness
Timon Willi, Alistair Letcher, Johannes Treutlein, and Jakob Foerster. COLA : Consistent learning with opponent-learning awareness. In International Conference on Machine Learning (ICML), 2022
2022
-
[85]
Open and closed loop Nash equilibria in games with a continuum of players
Agnieszka Wiszniewska-Matyszkiel. Open and closed loop Nash equilibria in games with a continuum of players. Journal of Optimization Theory and Applications, 160: 0 280--301, 2014
2014
-
[86]
Population-aware online mirror descent for mean-field games by deep reinforcement learning
Zida Wu, Mathieu Lauri\` e re, Samuel Jia Cong Chua, Matthieu Geist, Olivier Pietquin, and Ankur Mehta. Population-aware online mirror descent for mean-field games by deep reinforcement learning. In Autonomous Agents and Multi-Agent Systems, page 2561–2563, Richland, SC, 2024....
2024
-
[87]
C. Xin, G. Yang, and J. P. Huang. Ising game: Nonequilibrium steady states of resource-allocation systems. Physica A: Statistical Mechanics and its Applications, 471: 0 666--673, 2017
2017
-
[88]
Stochastic search using the natural gradient
Sun Yi, Daan Wierstra, Tom Schaul, and J\" u rgen Schmidhuber. Stochastic search using the natural gradient. In International Conference on Machine Learning (ICML), 2009
2009
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.