Pith. sign in

REVIEW 3 major objections 3 minor 32 references

Discrete-time average-cost mean-field games on Polish spaces

T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Mean-field equilibria exist under drift and minorization conditions

desk verdict A genuinely broader existence result for average-cost mean-field games, but the main theorem as stated has a hole involving the initial distribution μ0. read the letter →

arxiv 1908.08793 v1 pith:6TIC2GT5 submitted 2019-08-22 math.OC cs.MA

classification math.OCcs.MA MSC 91A1591A1091A1393E20
keywords mean-fieldgamesaveragecostapproximateNashequilibriumPolishspacesoptimalityequationKakutanifixedpointtheoremdriftconditionminorization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves that discrete-time mean-field games with average-cost payoffs have a mean-field equilibrium when the state and action spaces are Polish and the transition kernel satisfies a uniform minorization condition together with a drift inequality. A mean-field equilibrium is a pair of a policy and a state distribution that are consistent: the policy minimizes the infinite-horizon average cost against that distribution, and the distribution is the invariant law generated by the policy. The paper further proves that this equilibrium policy, used by every agent in an N-agent version of the game, is an epsilon-Nash equilibrium for every epsilon greater than zero once N is large enough. These results matter because they extend average-cost mean-field-game theory beyond compact or finite state spaces and give a dynamic-programming route, through the average cost optimality equation, to approximate equilibria in large anonymous games.

What carries the argument

The load-bearing object is the average cost optimality equation written through the Bellman operator T_mu u(x) = min_{a in A} [ c(x,a,mu) + integral_X u(y) (p(dy|x,a,mu) - lambda(dy)) ]. The minorization condition makes T_mu a contraction on bounded continuous functions with modulus beta = 1 - lambda(X), so each mu has a unique fixed point h_mu that encodes the optimal average cost and the optimality condition. The drift inequality defines the compact set P_c(X) of measures with integral w dmu ≤ integral w dlambda / (1 - alpha), and the corresponding set Xi of joint state-action measures is compact and convex. A fixed point of the set-valued map Gamma(nu) = C(nu) ∩ B(nu), where C imposes invariant-measure consistency and B imposes optimality, yields the mean-field equilibrium via Kakutani's fixed point theorem.

What would settle it

Construct a weakly continuous, bounded-cost model on a non-compact state space with compact action set, such as a shifted exponential transition p(dy|x)=$e^{{-(y-x)}}$1_{y>=x} dy on X=[0,infty), which has no common sub-probability minorant, and check numerically whether a mean-field equilibrium still exists; existence there would show the conditions are not necessary, while nonexistence under the other assumptions would show the minorization condition is load-bearing. To test the proof itself, find a weakly convergent sequence mu_n -> mu for which the fixed points h_{mu_n} of the average cost optimality equation do not converge uniformly on a compact set, which would break Proposition 4.4 and the closed-graph step of the fixed-point argument.

Watch

Extended reading notes

Core claim

The paper's central claim is Theorem 2.2: under Assumption 1, consisting of a bounded continuous cost, a weakly continuous transition kernel, a compact action set, a uniform minorization condition p(·|x,a,mu) ≥ lambda(·), and a uniform drift inequality with a single moment function w and constant alpha < 1, the mean-field game admits a mean-field equilibrium (pi*, mu*). Theorem 3.3 then states that, if all N agents adopt pi*, the resulting N-tuple of policies is an epsilon-Nash equilibrium for the finite N-player game for every epsilon > 0 and all sufficiently large N. The equilibrium is obtained by solving the average cost optimality equation for each candidate mean-field measure mu, extracting an optimal policy, and then applying a fixed-point argument to make mu consistent with the invariant distribution of the optimally controlled process.

Load-bearing premise

The entire argument rests on one uniform bound: a single sub-probability measure lambda and a single moment function w must control the transition kernel for every state, action, and mean-field measure, and if that uniformity fails, the contraction property and the compactness of the candidate equilibrium set are lost.

Editorial extensions

If this is right

  • For any average-cost mean-field game satisfying Assumption 1, an equilibrium policy and a consistent state distribution exist, covering non-compact Polish state spaces rather than only finite or compact ones.
  • The mean-field equilibrium policy is asymptotically optimal for the finite-player game: for each epsilon > 0 there is a threshold N(epsilon) beyond which no agent can improve by more than epsilon through a unilateral deviation.
  • The existence proof runs through the average cost optimality equation and a fixed-point argument, so the same dynamic-programming machinery can be reused in related infinite-horizon control problems.
  • Because all agents are identical, the same equilibrium policy works simultaneously for every player, so the epsilon-Nash property does not require agent-specific policies or centralized coordination.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The uniform minorization and drift assumptions are probably stronger than necessary; a natural test is whether the theorem survives with a state-dependent minorant or with the drift inequality required only on states reachable under equilibrium play.
  • The epsilon-Nash proof in the paper uses the simplification that transitions do not depend on mu; if p depended on mu in a Lipschitz way, the same extended-state comparison would likely acquire an extra term proportional to the distance between mean-field measures, yielding a similar approximation bound with a modified error.
  • The belief-state transformation suggested in the conclusion points to a partially observed version becoming a fully observed mean-field game on the space of posterior distributions, provided the filter transition inherits the minorization and drift conditions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper develops an existence theory for discrete-time average-cost mean-field games on Polish state and action spaces. The model is a tuple (X,A,p,c,mu0); for a fixed state-measure mu, a generic agent minimizes the limsup average cost J_mu(pi) under the transition kernel p(·|x,a,mu). A mean-field equilibrium (MFE) is a pair (pi,mu) such that pi is optimal for mu and mu is invariant under pi, with the additional clause mu0=mu inside the definition of Psi (Definition 2.1). Under Assumption 1 (bounded continuous cost, weakly continuous transition, compact action space, uniform minorization by a fixed sub-probability lambda, and a drift inequality with a moment function w), the paper proves existence of an MFE in Theorem 2.2 by combining the average-cost optimality equation with Kakutani's fixed point theorem. Under Assumption 2 (mu-independent transitions and uniform continuity of the cost in the measure variable), Theorem 3.3 shows that the repeated MFE policy is an epsilon-Nash equilibrium for the N-player game for all sufficiently large N; the proof couples invariant distributions of the N-player chain with the mean-field chain. Section 4 contains the fixed-point construction on the space Xi of state-action measures, and Section 5 contains the finite-player approximation argument.

Significance. If the stated results are corrected, the paper would make a useful contribution: it extends average-cost mean-field games beyond compact state spaces using drift and minorization conditions, and it gives a relatively clean approximate-Nash argument based on invariant measures rather than transient analysis. The dynamic-programming approach through the average-cost optimality equation and the fixed-point formulation on joint state-action measures is natural, and the closed-graph proof in Proposition 4.3 is careful and substantially self-contained. The paper also clearly delineates the assumptions that make the average-cost criterion tractable. However, the manuscript as written contains a load-bearing inconsistency concerning the prescribed initial distribution mu0, and a separate incorrect product drift inequality in Section 5, so the central theorems are not yet established in the form stated.

major comments (3)
  1. [Definition 2.1, Section 4, Proposition 4.2] The definition of Psi(mu) requires both optimality for mu and the condition mu0=mu. In the fixed-point proof, Gamma is defined on Xi and a fixed point nu is shown to satisfy nu in C(nu) ∩ B(nu); Proposition 4.2 then declares (pi,nu1) a mean-field equilibrium. The proof establishes (i) that nu1 is invariant under pi and (ii), via Theorem 4.1, that pi is optimal for nu1 when the initial state is distributed as nu1. It never verifies that the prescribed initial distribution mu0 equals nu1. This is not a mere presentation gap: the statement is false for arbitrary mu0. For example, let X={0,1}, A={a}, p(0|x,a,mu)=1 for all x,a,mu, lambda=0.5 delta_0, w≡1, alpha=0.75, c≡1, and mu0=delta_1. Assumption 1 holds, but the only invariant measure under any policy is delta_0, while Psi(delta_0) is empty because mu0=delta_0 fails; no pair satisfies Definition 2.1. To repair the claim, either remove the clause mu0=mu from Psi and prove, under Assumption 1, that the average cost is independent of the initial distribution (the minorization condition gives uniform ergodicity, so this should be possible), or state and prove Theorem 2.2 only for mu0 equal to the equilibrium invariant marginal produced by the fixed point. The current proof supports only the latter, and even then the statement must say so.
  2. [Theorem 4.1 and Section 2] Theorem 4.1 proves optimality of a policy for mu only when x(0) ~ mu_{pi,mu}, the invariant distribution of the policy. However, the average cost J_mu(pi) in Section 2 is defined with the prescribed initial distribution x(0) ~ mu0. The proof of Theorem 4.1 works with J_{mu,n}(pi,h_mu,mu_{pi,mu}) and therefore establishes that the policy achieves the optimal average cost from its own invariant distribution. The equality of the average cost from mu0 and from mu_{pi,mu} is neither stated nor proved. This is the technical point on which the previous comment turns; it should be addressed explicitly when Definition 2.1 is corrected.
  3. [Section 5, Eqs. (5.1)-(5.2)] The claimed product drift inequality (5.2) is not a consequence of Assumption 1(e). With w_N = prod_i w(y_i) and phat_N = p^N - lambda^N, the left-hand side equals prod_i (∫ w dp(·|x_i,a_i)) - (∫ w dlambda)^N. For C = ∫ w dlambda >0 and w(x_i)=L for all i, this is (alpha L + C)^N - C^N, which is larger than alpha^N L^N; for instance alpha=0.5, C=1, L=100, N=2 gives 2601 > 2501. Thus (5.2) is false as stated. The uniqueness of the invariant distribution of the N-player chain can be obtained from the product minorization (5.1) alone, since uniform minorization implies uniform ergodicity, so the proof is repairable by replacing (5.2) with a valid argument or removing it from the hypotheses used to invoke [30,12].
minor comments (3)
  1. [Proof of Proposition 4.4] In the displayed expression for u_{k+1}^{(n)} - u_{k+1}, the second minimand contains a spurious factor beta multiplying the integral of u_k; the operator T has no such factor.
  2. [Theorem 2.2 statement vs. Section 2] Theorem 2.2 states the game as (X,A,p,c), whereas the model introduced in Section 2 is (X,A,p,c,mu0). This ambiguity is directly connected to Major Comment 1 and should be resolved in a revision.
  3. [Definition 2.1] The clause 'and mu0=mu' in the definition of Psi(mu) is unusual for a stationary mean-field equilibrium and is never used in the fixed-point construction; please clarify whether the initial distribution is meant to constrain the equilibrium marginal or whether a stationary MFE without that constraint is intended.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the equilibrium existence proof is a self-contained contraction/fixed-point argument, and the approximate-Nash proof is an independent coupling estimate; the μ0=μ clause is a verification gap, not a circular reduction.

full rationale

Walking the derivation chain: Theorem 2.2 is proved in Section 4 by defining the contraction operator T_μ (with modulus β = 1 − λ(X) under Assumption 1), obtaining h_μ from the Banach fixed point theorem, and then forming Γ = C ∩ B on the compact convex set Ξ. A fixed point of Γ yields a pair (π, ν1) with ν1 invariant under π (C-condition) and π attaining the ACOE minimum (B-condition); Proposition 4.2 then identifies this as an equilibrium via Theorem 4.1. No parameter in the construction is fitted to the equilibrium it predicts, and no step assumes the target result. Theorem 3.3 is derived from the mean-field equilibrium policy through an independent coupling and Law-of-Large-Numbers estimate (Proposition 5.1). The only self-citations are [26] (the closed-graph proof is adapted rather than imported as a theorem) and [27] (a standard contraction fact); neither supplies the main conclusion. Thus there is no circular reduction. For completeness, the proof of Proposition 4.2 does not verify the μ0 = ν1 clause in Definition 2.1; Theorem 2.2 as stated can fail when μ0 is not the invariant marginal of the equilibrium policy (e.g., a deterministic transition to state 0 with μ0 = δ_1). That is a correctness gap, not a circularity, because the fixed-point argument does not presuppose the equilibrium it claims to produce.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

No free parameters or invented entities appear in the proof. The central claim rests on standard MDP theorems (Banach and Kakutani fixed points, measurable selection, Prokhorov tightness) and on the stated drift and minorization conditions, which are the main domain assumptions. These conditions are strong but common in average-cost stochastic control.

assumptions (6)
  • domain assumption Weak continuity of p and bounded continuity of c (Assumption 1(a)-(b)).
    Used to make the Bellman operator T_mu map Cb(X) into itself and to justify the closed graph of Gamma.
  • domain assumption Compactness of the action space A (Assumption 1(c)).
    Ensures the minimizers in the optimality equation exist for each x and that the set Xi is tight.
  • domain assumption Uniform minorization p(·|x,a,mu) >= lambda(·) for a non-degenerate sub-probability measure lambda (Assumption 1(d)).
    Gives the contraction property of T_mu with modulus beta=1-lambda(X), the uniqueness of invariant measures, and the average cost optimality equation.
  • domain assumption Uniform drift inequality with a continuous moment function w and constant alpha<1 (Assumption 1(e)).
    Ensures the state process does not escape to infinity, giving a compact set P_c(X) of admissible state measures and ergodic properties needed for average cost.
  • domain assumption Transition kernel p does not depend on mu in the finite-agent approximation (Assumption 2(a)).
    Used in Theorem 3.3 to factor the invariant distribution of the N-player process as a product of single-agent invariant distributions.
  • domain assumption Average cost is independent of the initial distribution under drift and minorization.
    The paper defines J_mu with the fixed initial distribution mu0 but proves optimality in Theorem 4.1 for the initial distribution mu_pi_mu. The independence is a standard consequence of uniform ergodicity, but it is not stated explicitly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Discrete-time average-cost mean-field games on Polish spaces." pith.science (2026). https://pith.science/paper/6TIC2GT5

@misc{pith2026190808793,
  author       = {Pith},
  title        = {Pith review of: Discrete-time average-cost mean-field games on Polish spaces},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6TIC2GT5}},
  note         = {Machine review of arXiv:1908.08793}
}
read the original abstract

In stochastic dynamic games, when the number of players is sufficiently large and the interactions between agents depend on empirical state distribution, one way to approximate the original game is to introduce infinite-population limit of the problem. In the infinite population limit, a generic agent is faced with a \emph{so-called} mean-field game. In this paper, we study discrete-time mean-field games with average-cost criteria. Using average cost optimality equation and Kakutani's fixed point theorem, we establish the existence of Nash equilibria for mean-field games under drift and minorization conditions on the dynamics of each agent. Then, we show that the equilibrium policy in the mean-field game, when adopted by each agent, is an approximate Nash equilibrium for the corresponding finite-agent game with sufficiently many agents.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 31 canonical work pages

  1. [1]

    Adlakha, S., Johari, R., and Weintraub, G. (2015). Equil ibria of dynamic games with many players: Existence, approximation, and market structure. Journal of Economic Theory , 156:269–316

  2. [2]

    and Border, K

    Aliprantis, C. and Border, K. (2006). Infinite Dimensional Analysis . Berlin, Springer, 3rd ed

  3. [3]

    Bensoussan, A., Frehse, J., and Yam, P. (2013). Mean Field Games and Mean Field Type Control Theory . Springer, New York

  4. [4]

    Billingsley, P. (1999). Convergence of Probability Measures . New York: Wiley, 2nd edition

  5. [5]

    Biswas, A. (2015). Mean field games with ergodic cost for d iscrete time Markov processes. arXiv:1510.08968. 16 AUTHOR/Turk J Math

  6. [6]

    Bogachev, V. (2007). Measure Theory: Volume II . Springer

  7. [7]

    Cardaliaguet, P. (2011). Notes on Mean-field Games

  8. [8]

    and Delarue, F

    Carmona, R. and Delarue, F. (2013). Probabilistic analy sis of mean-field games. SIAM Journal on Control and Optimization, 51(4):2705–2734

Show all 32 references
  1. [9]

    Elliot, R., Li, X., and Ni, Y. (2013). Discrete time mean- field stochastic linear-quadratic optimal control problem s. Automatica, 49:3222–3233

  2. [10]

    Gomes, D., Mohr, J., and Souza, R. (2010). Discrete time , finite state space mean field games. Journal de Math´ ematiques Pures et Appliqu´ ees, 93:308–328

  3. [11]

    and Sa´ ude, J

    Gomes, D. and Sa´ ude, J. (2014). Mean field games models - a brief survey. Dynamic Games and Applications , 4(2):110–154

  4. [12]

    and Hernandez-Lerma, O

    Gordienko, E. and Hernandez-Lerma, O. (1995). Average cost Markov control processes with weighted norms: Existence of canonical policies. APPLICATIONES MATHEMATICAE, 23(2):199–218

  5. [13]

    and Lasserre, J

    Hern´ andez-Lerma, O. and Lasserre, J. (1996). Discrete-Time Markov Control Processes: Basic Optimality Criteria. Springer

  6. [14]

    and Lasserre, J

    Hern´ andez-Lerma, O. and Lasserre, J. (1999). Further Topics on Discrete-Time Markov Control Processes . Springer

  7. [15]

    and Montes-De-Oca, R

    Hern´ andez-Lerma, O. and Montes-De-Oca, R. and Cavazo s-Cadena, R. (1991). Recurrence conditions for Markov decision processes with Borel state space: a survey. Annals of Operations Research , 28(1):29–46

  8. [16]

    Huang, M. (2010). Large-population LQG games involvin g major player: The Nash certainity equivalence principle. SIAM Journal on Control and Optimization , 48(5):3318–3353

  9. [17]

    Huang, M., Caines, P., and Malham´ e, R. (2007). Large-p opulation cost coupled LQG problems with nonuniform agents: Individual-mass behavior and decentralized ǫ -Nash equilibria. IEEE Transactions on Automatic Control , 52(9):1560–1571

  10. [18]

    Huang, M., Malham´ e, R., and Caines, P. (2006). Large po pulation stochastic dynamic games: Closed loop McKean- Vlasov sysyems and the Nash certainity equivalence princip le. Communications in Information Systems , 6:221–252

  11. [19]

    Langen, H. (1981). Convergence of dynamic programming models. Mathematics of Operations Research , 6(4):493– 512

  12. [20]

    and P.Lions (2007)

    Lasry, J. and P.Lions (2007). Mean field games. Japanese Journal of Mathematics , 2:229–260

  13. [21]

    and Ba¸ sar, T

    Moon, J. and Ba¸ sar, T. (2015). Discrete-time decentra lized control using the risk-sensitive performance criter ion in the large population regime: a mean field approach. In ACC 2015 , Chicago

  14. [22]

    and Ba¸ sar, T

    Moon, J. and Ba¸ sar, T. (2016a). Discrete-time mean fiel d Stackelberg games with a large number of followers. In CDC 2016 , Las Vegas

  15. [23]

    and Ba¸ sar, T

    Moon, J. and Ba¸ sar, T. (2016b). Robust mean field games f or coupled Markov jump linear systems. International Journal of Control , 89(7):1367–1381

  16. [24]

    and Nair, G

    Nourian, M. and Nair, G. (2013). Linear-quadratic-Gau ssian mean field games under high rate quantization. In CDC 2013 , Florence

  17. [25]

    Parthasarathy, K. (1967). Probability Measures on Metric Spaces . AMS Bookstore

  18. [26]

    Saldi, N., Ba¸ sar, T., and Raginsky, M. (2018a). Markov –Nash equilibria in mean-field games with discounted cost. SIAM Journal on Control and Optimization , 56(6):4256–4287

  19. [27]

    Saldi, N., Linder, T., and Y¨ uksel, S. (2018b). Finite approximations in discrete-time stochastic contro l: Quantized models and asymptotic optimality . Springer, Cham

  20. [28]

    Serfozo, R. (1982). Convergence of Lebesgue integrals with varying measures. Sankhya Ser.A , pages 380–402

  21. [29]

    Tembine, H., Zhu, Q., and Ba¸ sar, T. (2014). Risk-sensi tive mean field games. IEEE Transactions on Automatic Control, 59(4):835–850. 17 AUTHOR/Turk J Math

  22. [30]

    Vega-Amaya, O. (2003). The average cost optimality equ ation: a fixed point approach. Bolet ´ ın de la Sociedad Matem´ atica Mexicana, 9(3):185–195

  23. [31]

    Wiecek, P. (2019). Discrete-time ergodic mean-field ga mes with average reward on compact spaces. Dynamic Games and Applications , pages 1–35

  24. [32]

    and Altman, E

    Wiecek, P. and Altman, E. (2015). Stationary anonymous sequential games with undiscounted rewards. Journal of Optimization Theory and Applications , 166(2):686-710. 18

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.