REVIEW 3 major objections 3 minor 32 references
Discrete-time average-cost mean-field games on Polish spaces
T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Mean-field equilibria exist under drift and minorization conditions
desk verdict A genuinely broader existence result for average-cost mean-field games, but the main theorem as stated has a hole involving the initial distribution μ0. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the average cost optimality equation written through the Bellman operator T_mu u(x) = min_{a in A} [ c(x,a,mu) + integral_X u(y) (p(dy|x,a,mu) - lambda(dy)) ]. The minorization condition makes T_mu a contraction on bounded continuous functions with modulus beta = 1 - lambda(X), so each mu has a unique fixed point h_mu that encodes the optimal average cost and the optimality condition. The drift inequality defines the compact set P_c(X) of measures with integral w dmu ≤ integral w dlambda / (1 - alpha), and the corresponding set Xi of joint state-action measures is compact and convex. A fixed point of the set-valued map Gamma(nu) = C(nu) ∩ B(nu), where C imposes invariant-measure consistency and B imposes optimality, yields the mean-field equilibrium via Kakutani's fixed point theorem.
What would settle it
Construct a weakly continuous, bounded-cost model on a non-compact state space with compact action set, such as a shifted exponential transition p(dy|x)=$e^{{-(y-x)}}$1_{y>=x} dy on X=[0,infty), which has no common sub-probability minorant, and check numerically whether a mean-field equilibrium still exists; existence there would show the conditions are not necessary, while nonexistence under the other assumptions would show the minorization condition is load-bearing. To test the proof itself, find a weakly convergent sequence mu_n -> mu for which the fixed points h_{mu_n} of the average cost optimality equation do not converge uniformly on a compact set, which would break Proposition 4.4 and the closed-graph step of the fixed-point argument.
Extended reading notes
Core claim
The paper's central claim is Theorem 2.2: under Assumption 1, consisting of a bounded continuous cost, a weakly continuous transition kernel, a compact action set, a uniform minorization condition p(·|x,a,mu) ≥ lambda(·), and a uniform drift inequality with a single moment function w and constant alpha < 1, the mean-field game admits a mean-field equilibrium (pi*, mu*). Theorem 3.3 then states that, if all N agents adopt pi*, the resulting N-tuple of policies is an epsilon-Nash equilibrium for the finite N-player game for every epsilon > 0 and all sufficiently large N. The equilibrium is obtained by solving the average cost optimality equation for each candidate mean-field measure mu, extracting an optimal policy, and then applying a fixed-point argument to make mu consistent with the invariant distribution of the optimally controlled process.
Load-bearing premise
The entire argument rests on one uniform bound: a single sub-probability measure lambda and a single moment function w must control the transition kernel for every state, action, and mean-field measure, and if that uniformity fails, the contraction property and the compactness of the candidate equilibrium set are lost.
Editorial extensions
If this is right
- For any average-cost mean-field game satisfying Assumption 1, an equilibrium policy and a consistent state distribution exist, covering non-compact Polish state spaces rather than only finite or compact ones.
- The mean-field equilibrium policy is asymptotically optimal for the finite-player game: for each epsilon > 0 there is a threshold N(epsilon) beyond which no agent can improve by more than epsilon through a unilateral deviation.
- The existence proof runs through the average cost optimality equation and a fixed-point argument, so the same dynamic-programming machinery can be reused in related infinite-horizon control problems.
- Because all agents are identical, the same equilibrium policy works simultaneously for every player, so the epsilon-Nash property does not require agent-specific policies or centralized coordination.
Reading between the lines
- The uniform minorization and drift assumptions are probably stronger than necessary; a natural test is whether the theorem survives with a state-dependent minorant or with the drift inequality required only on states reachable under equilibrium play.
- The epsilon-Nash proof in the paper uses the simplification that transitions do not depend on mu; if p depended on mu in a Lipschitz way, the same extended-state comparison would likely acquire an extra term proportional to the distance between mean-field measures, yielding a similar approximation bound with a modified error.
- The belief-state transformation suggested in the conclusion points to a partially observed version becoming a fully observed mean-field game on the space of posterior distributions, provided the filter transition inherits the minorization and drift conditions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops an existence theory for discrete-time average-cost mean-field games on Polish state and action spaces. The model is a tuple (X,A,p,c,mu0); for a fixed state-measure mu, a generic agent minimizes the limsup average cost J_mu(pi) under the transition kernel p(·|x,a,mu). A mean-field equilibrium (MFE) is a pair (pi,mu) such that pi is optimal for mu and mu is invariant under pi, with the additional clause mu0=mu inside the definition of Psi (Definition 2.1). Under Assumption 1 (bounded continuous cost, weakly continuous transition, compact action space, uniform minorization by a fixed sub-probability lambda, and a drift inequality with a moment function w), the paper proves existence of an MFE in Theorem 2.2 by combining the average-cost optimality equation with Kakutani's fixed point theorem. Under Assumption 2 (mu-independent transitions and uniform continuity of the cost in the measure variable), Theorem 3.3 shows that the repeated MFE policy is an epsilon-Nash equilibrium for the N-player game for all sufficiently large N; the proof couples invariant distributions of the N-player chain with the mean-field chain. Section 4 contains the fixed-point construction on the space Xi of state-action measures, and Section 5 contains the finite-player approximation argument.
Significance. If the stated results are corrected, the paper would make a useful contribution: it extends average-cost mean-field games beyond compact state spaces using drift and minorization conditions, and it gives a relatively clean approximate-Nash argument based on invariant measures rather than transient analysis. The dynamic-programming approach through the average-cost optimality equation and the fixed-point formulation on joint state-action measures is natural, and the closed-graph proof in Proposition 4.3 is careful and substantially self-contained. The paper also clearly delineates the assumptions that make the average-cost criterion tractable. However, the manuscript as written contains a load-bearing inconsistency concerning the prescribed initial distribution mu0, and a separate incorrect product drift inequality in Section 5, so the central theorems are not yet established in the form stated.
major comments (3)
- [Definition 2.1, Section 4, Proposition 4.2] The definition of Psi(mu) requires both optimality for mu and the condition mu0=mu. In the fixed-point proof, Gamma is defined on Xi and a fixed point nu is shown to satisfy nu in C(nu) ∩ B(nu); Proposition 4.2 then declares (pi,nu1) a mean-field equilibrium. The proof establishes (i) that nu1 is invariant under pi and (ii), via Theorem 4.1, that pi is optimal for nu1 when the initial state is distributed as nu1. It never verifies that the prescribed initial distribution mu0 equals nu1. This is not a mere presentation gap: the statement is false for arbitrary mu0. For example, let X={0,1}, A={a}, p(0|x,a,mu)=1 for all x,a,mu, lambda=0.5 delta_0, w≡1, alpha=0.75, c≡1, and mu0=delta_1. Assumption 1 holds, but the only invariant measure under any policy is delta_0, while Psi(delta_0) is empty because mu0=delta_0 fails; no pair satisfies Definition 2.1. To repair the claim, either remove the clause mu0=mu from Psi and prove, under Assumption 1, that the average cost is independent of the initial distribution (the minorization condition gives uniform ergodicity, so this should be possible), or state and prove Theorem 2.2 only for mu0 equal to the equilibrium invariant marginal produced by the fixed point. The current proof supports only the latter, and even then the statement must say so.
- [Theorem 4.1 and Section 2] Theorem 4.1 proves optimality of a policy for mu only when x(0) ~ mu_{pi,mu}, the invariant distribution of the policy. However, the average cost J_mu(pi) in Section 2 is defined with the prescribed initial distribution x(0) ~ mu0. The proof of Theorem 4.1 works with J_{mu,n}(pi,h_mu,mu_{pi,mu}) and therefore establishes that the policy achieves the optimal average cost from its own invariant distribution. The equality of the average cost from mu0 and from mu_{pi,mu} is neither stated nor proved. This is the technical point on which the previous comment turns; it should be addressed explicitly when Definition 2.1 is corrected.
- [Section 5, Eqs. (5.1)-(5.2)] The claimed product drift inequality (5.2) is not a consequence of Assumption 1(e). With w_N = prod_i w(y_i) and phat_N = p^N - lambda^N, the left-hand side equals prod_i (∫ w dp(·|x_i,a_i)) - (∫ w dlambda)^N. For C = ∫ w dlambda >0 and w(x_i)=L for all i, this is (alpha L + C)^N - C^N, which is larger than alpha^N L^N; for instance alpha=0.5, C=1, L=100, N=2 gives 2601 > 2501. Thus (5.2) is false as stated. The uniqueness of the invariant distribution of the N-player chain can be obtained from the product minorization (5.1) alone, since uniform minorization implies uniform ergodicity, so the proof is repairable by replacing (5.2) with a valid argument or removing it from the hypotheses used to invoke [30,12].
minor comments (3)
- [Proof of Proposition 4.4] In the displayed expression for u_{k+1}^{(n)} - u_{k+1}, the second minimand contains a spurious factor beta multiplying the integral of u_k; the operator T has no such factor.
- [Theorem 2.2 statement vs. Section 2] Theorem 2.2 states the game as (X,A,p,c), whereas the model introduced in Section 2 is (X,A,p,c,mu0). This ambiguity is directly connected to Major Comment 1 and should be resolved in a revision.
- [Definition 2.1] The clause 'and mu0=mu' in the definition of Psi(mu) is unusual for a stationary mean-field equilibrium and is never used in the fixed-point construction; please clarify whether the initial distribution is meant to constrain the equilibrium marginal or whether a stationary MFE without that constraint is intended.
Circularity Check
No circularity: the equilibrium existence proof is a self-contained contraction/fixed-point argument, and the approximate-Nash proof is an independent coupling estimate; the μ0=μ clause is a verification gap, not a circular reduction.
full rationale
Walking the derivation chain: Theorem 2.2 is proved in Section 4 by defining the contraction operator T_μ (with modulus β = 1 − λ(X) under Assumption 1), obtaining h_μ from the Banach fixed point theorem, and then forming Γ = C ∩ B on the compact convex set Ξ. A fixed point of Γ yields a pair (π, ν1) with ν1 invariant under π (C-condition) and π attaining the ACOE minimum (B-condition); Proposition 4.2 then identifies this as an equilibrium via Theorem 4.1. No parameter in the construction is fitted to the equilibrium it predicts, and no step assumes the target result. Theorem 3.3 is derived from the mean-field equilibrium policy through an independent coupling and Law-of-Large-Numbers estimate (Proposition 5.1). The only self-citations are [26] (the closed-graph proof is adapted rather than imported as a theorem) and [27] (a standard contraction fact); neither supplies the main conclusion. Thus there is no circular reduction. For completeness, the proof of Proposition 4.2 does not verify the μ0 = ν1 clause in Definition 2.1; Theorem 2.2 as stated can fail when μ0 is not the invariant marginal of the equilibrium policy (e.g., a deterministic transition to state 0 with μ0 = δ_1). That is a correctness gap, not a circularity, because the fixed-point argument does not presuppose the equilibrium it claims to produce.
Assumptions & free parameters
assumptions (6)
- domain assumption Weak continuity of p and bounded continuity of c (Assumption 1(a)-(b)).
- domain assumption Compactness of the action space A (Assumption 1(c)).
- domain assumption Uniform minorization p(·|x,a,mu) >= lambda(·) for a non-degenerate sub-probability measure lambda (Assumption 1(d)).
- domain assumption Uniform drift inequality with a continuous moment function w and constant alpha<1 (Assumption 1(e)).
- domain assumption Transition kernel p does not depend on mu in the finite-agent approximation (Assumption 2(a)).
- domain assumption Average cost is independent of the initial distribution under drift and minorization.
Cite this review
Pith. "Pith review of Discrete-time average-cost mean-field games on Polish spaces." pith.science (2026). https://pith.science/paper/6TIC2GT5
@misc{pith2026190808793,
author = {Pith},
title = {Pith review of: Discrete-time average-cost mean-field games on Polish spaces},
year = {2026},
howpublished = {\url{https://pith.science/paper/6TIC2GT5}},
note = {Machine review of arXiv:1908.08793}
}
read the original abstract
In stochastic dynamic games, when the number of players is sufficiently large and the interactions between agents depend on empirical state distribution, one way to approximate the original game is to introduce infinite-population limit of the problem. In the infinite population limit, a generic agent is faced with a \emph{so-called} mean-field game. In this paper, we study discrete-time mean-field games with average-cost criteria. Using average cost optimality equation and Kakutani's fixed point theorem, we establish the existence of Nash equilibria for mean-field games under drift and minorization conditions on the dynamics of each agent. Then, we show that the equilibrium policy in the mean-field game, when adopted by each agent, is an approximate Nash equilibrium for the corresponding finite-agent game with sufficiently many agents.
Reference graph
Works this paper leans on
-
[1]
Adlakha, S., Johari, R., and Weintraub, G. (2015). Equil ibria of dynamic games with many players: Existence, approximation, and market structure. Journal of Economic Theory , 156:269–316
work page 2015
-
[2]
Aliprantis, C. and Border, K. (2006). Infinite Dimensional Analysis . Berlin, Springer, 3rd ed
work page 2006
-
[3]
Bensoussan, A., Frehse, J., and Yam, P. (2013). Mean Field Games and Mean Field Type Control Theory . Springer, New York
work page 2013
-
[4]
Billingsley, P. (1999). Convergence of Probability Measures . New York: Wiley, 2nd edition
work page 1999
-
[5]
Biswas, A. (2015). Mean field games with ergodic cost for d iscrete time Markov processes. arXiv:1510.08968. 16 AUTHOR/Turk J Math
arXiv 2015
-
[6]
Bogachev, V. (2007). Measure Theory: Volume II . Springer
work page 2007
-
[7]
Cardaliaguet, P. (2011). Notes on Mean-field Games
work page 2011
-
[8]
Carmona, R. and Delarue, F. (2013). Probabilistic analy sis of mean-field games. SIAM Journal on Control and Optimization, 51(4):2705–2734
work page 2013
Show all 32 references
-
[9]
Elliot, R., Li, X., and Ni, Y. (2013). Discrete time mean- field stochastic linear-quadratic optimal control problem s. Automatica, 49:3222–3233
2013
-
[10]
Gomes, D., Mohr, J., and Souza, R. (2010). Discrete time , finite state space mean field games. Journal de Math´ ematiques Pures et Appliqu´ ees, 93:308–328
2010
-
[11]
and Sa´ ude, J
Gomes, D. and Sa´ ude, J. (2014). Mean field games models - a brief survey. Dynamic Games and Applications , 4(2):110–154
2014
-
[12]
and Hernandez-Lerma, O
Gordienko, E. and Hernandez-Lerma, O. (1995). Average cost Markov control processes with weighted norms: Existence of canonical policies. APPLICATIONES MATHEMATICAE, 23(2):199–218
1995
-
[13]
and Lasserre, J
Hern´ andez-Lerma, O. and Lasserre, J. (1996). Discrete-Time Markov Control Processes: Basic Optimality Criteria. Springer
1996
-
[14]
and Lasserre, J
Hern´ andez-Lerma, O. and Lasserre, J. (1999). Further Topics on Discrete-Time Markov Control Processes . Springer
1999
-
[15]
and Montes-De-Oca, R
Hern´ andez-Lerma, O. and Montes-De-Oca, R. and Cavazo s-Cadena, R. (1991). Recurrence conditions for Markov decision processes with Borel state space: a survey. Annals of Operations Research , 28(1):29–46
1991
-
[16]
Huang, M. (2010). Large-population LQG games involvin g major player: The Nash certainity equivalence principle. SIAM Journal on Control and Optimization , 48(5):3318–3353
2010
-
[17]
Huang, M., Caines, P., and Malham´ e, R. (2007). Large-p opulation cost coupled LQG problems with nonuniform agents: Individual-mass behavior and decentralized ǫ -Nash equilibria. IEEE Transactions on Automatic Control , 52(9):1560–1571
2007
-
[18]
Huang, M., Malham´ e, R., and Caines, P. (2006). Large po pulation stochastic dynamic games: Closed loop McKean- Vlasov sysyems and the Nash certainity equivalence princip le. Communications in Information Systems , 6:221–252
2006
-
[19]
Langen, H. (1981). Convergence of dynamic programming models. Mathematics of Operations Research , 6(4):493– 512
1981
-
[20]
and P.Lions (2007)
Lasry, J. and P.Lions (2007). Mean field games. Japanese Journal of Mathematics , 2:229–260
2007
-
[21]
and Ba¸ sar, T
Moon, J. and Ba¸ sar, T. (2015). Discrete-time decentra lized control using the risk-sensitive performance criter ion in the large population regime: a mean field approach. In ACC 2015 , Chicago
2015
-
[22]
and Ba¸ sar, T
Moon, J. and Ba¸ sar, T. (2016a). Discrete-time mean fiel d Stackelberg games with a large number of followers. In CDC 2016 , Las Vegas
2016
-
[23]
and Ba¸ sar, T
Moon, J. and Ba¸ sar, T. (2016b). Robust mean field games f or coupled Markov jump linear systems. International Journal of Control , 89(7):1367–1381
2016
-
[24]
and Nair, G
Nourian, M. and Nair, G. (2013). Linear-quadratic-Gau ssian mean field games under high rate quantization. In CDC 2013 , Florence
2013
-
[25]
Parthasarathy, K. (1967). Probability Measures on Metric Spaces . AMS Bookstore
1967
-
[26]
Saldi, N., Ba¸ sar, T., and Raginsky, M. (2018a). Markov –Nash equilibria in mean-field games with discounted cost. SIAM Journal on Control and Optimization , 56(6):4256–4287
2018
-
[27]
Saldi, N., Linder, T., and Y¨ uksel, S. (2018b). Finite approximations in discrete-time stochastic contro l: Quantized models and asymptotic optimality . Springer, Cham
2018
-
[28]
Serfozo, R. (1982). Convergence of Lebesgue integrals with varying measures. Sankhya Ser.A , pages 380–402
1982
-
[29]
Tembine, H., Zhu, Q., and Ba¸ sar, T. (2014). Risk-sensi tive mean field games. IEEE Transactions on Automatic Control, 59(4):835–850. 17 AUTHOR/Turk J Math
2014
-
[30]
Vega-Amaya, O. (2003). The average cost optimality equ ation: a fixed point approach. Bolet ´ ın de la Sociedad Matem´ atica Mexicana, 9(3):185–195
2003
-
[31]
Wiecek, P. (2019). Discrete-time ergodic mean-field ga mes with average reward on compact spaces. Dynamic Games and Applications , pages 1–35
2019
-
[32]
and Altman, E
Wiecek, P. and Altman, E. (2015). Stationary anonymous sequential games with undiscounted rewards. Journal of Optimization Theory and Applications , 166(2):686-710. 18
2015
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.