{"id":"225f679f-be29-4540-9cf5-3c6f887fd3b0","arxiv_id":"1908.08793","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Discrete-time average-cost mean-field games on Polish spaces admit mean-field equilibria under drift and minorization conditions, and these equilibria yield approximate Nash equilibria for finite-player games.","lead":"This paper proves that discrete-time mean-field games with average-cost payoffs have equilibria on very general spaces, provided the random dynamics satisfy a drift and a minorization condition. It also shows that the equilibrium strategy is approximately optimal for large but finite player populations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The MFE definition imposes μ0=μ, but the fixed-point proof never verifies this; for a fixed μ0 the theorem is false as stated (two-state counterexample).","rationale":"The paper's core fixed-point strategy is largely sound: the contraction of T_μ, the compactness of Ξ, the closed-graph argument, and the approximation argument in Section 5 all appear to hold modulo minor typos (e.g., the stray β in Proposition 4.4 and the compact-convergence wording). The main flaw is in the formal statement of the existence theorem: the equilibrium definition carries the condition μ0=μ, but the proof produces a fixed point whose first marginal is an invariant measure and never connects it to the prescribed initial distribution μ0. Theorem 4.1 proves optimality only when the process starts from the invariant measure of the candidate policy. Because the paper does not state or prove initial-distribution independence for the average cost, the theorem is false as written for a fixed μ0 that differs from the equilibrium invariant measure; the two-state deterministic example makes this concrete. This is not an objection to the underlying mathematics—under Assumption 1 the average cost is in fact initial-independent, and deleting the μ0=μ clause would repair the statement—but it is a load-bearing presentation gap that must be fixed. The reader's verdict of CONDITIONAL remains appropriate: the result is likely correct in intent but requires a corrected definition or an added ergodicity lemma. I therefore leave the verdict unchanged, and I partially agree with the reader, who flagged the μ0=μ ambiguity in the rationale but did not make it the weakest assumption.","tokens_in":15291,"tokens_out":25801,"duration_ms":254692,"concrete_test":"Specialize Theorem 2.2 to the two-state deterministic model: X={0,1}, A={a}, p(0|x,a,μ)=1, c≡1, λ=0.5δ_0, w(x)=1, α=0.75, μ0=δ_1. Verify each part of Assumption 1 holds (minorization, drift, weak continuity, compact A). Then compute that the only invariant measure for the sole policy is δ_0, and that Definition 2.1 with μ0=δ_1 gives Ψ(δ_0)=∅, so no mean-field equilibrium exists. This would falsify the literal statement of Theorem 2.2. If the authors intend the theorem without the μ0=μ requirement, the same example should be rerun after adding a lemma that under Assumption 1 the average cost J_μ(π) is independent of the initial distribution for every stationary policy; if that lemma is true, the μ0=μ condition can be deleted and the existence result restored in the standard sense.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Definition 2.1 says π∈Ψ(μ) only if π is optimal for μ and μ0=μ. The fixed-point argument in Section 4 constructs Γ and obtains a fixed point ν∈Ξ, then Proposition 4.2 declares (π,ν1) a mean-field equilibrium from ν∈C(ν)∩B(ν). It never checks μ0=ν1. Theorem 4.1 only establishes optimality when the initial distribution is the invariant measure μ_{π,μ}, not the prescribed μ0. If μ0 is part of the model and is arbitrary, the theorem can fail. Example satisfying Assumption 1: X={0,1}, A={a}, p(0|x,a,μ)=1 for all x,a,μ, λ=0.5δ_0, w constant 1, α=0.75, c≡1. The unique invariant measure for the only policy is δ_0. If μ0=δ_1, then Ψ(δ_0)=∅ because no policy satisfies μ0=δ_0; hence no (π,μ) meets Definition 2.1. This contradicts Theorem 2.2 as stated. The paper silently relies on the ergodic fact that the average cost is independent of the initial distribution, but this is neither stated nor proved, and the μ0=μ clause makes it logically indispensable. The central claim is therefore not established for arbitrary μ0; it holds only if μ0 is chosen to equal the equilibrium invariant marginal, or if the μ0=μ condition is removed and initial-condition independence is proved under Assumption 1.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops an existence theory for discrete-time average-cost mean-field games on Polish state and action spaces. The model is a tuple (X,A,p,c,mu0); for a fixed state-measure mu, a generic agent minimizes the limsup average cost J_mu(pi) under the transition kernel p(·|x,a,mu). A mean-field equilibrium (MFE) is a pair (pi,mu) such that pi is optimal for mu and mu is invariant under pi, with the additional clause mu0=mu inside the definition of Psi (Definition 2.1). Under Assumption 1 (bounded continuous cost, weakly continuous transition, compact action space, uniform minorization by a fixed sub-probability lambda, and a drift inequality with a moment function w), the paper proves existence of an MFE in Theorem 2.2 by combining the average-cost optimality equation with Kakutani's fixed point theorem. Under Assumption 2 (mu-independent transitions and uniform continuity of the cost in the measure variable), Theorem 3.3 shows that the repeated MFE policy is an epsilon-Nash equilibrium for the N-player game for all sufficiently large N; the proof couples invariant distributions of the N-player chain with the mean-field chain. Section 4 contains the fixed-point construction on the space Xi of state-action measures, and Section 5 contains the finite-player approximation argument.","tokens_in":15564,"tokens_out":11274,"duration_ms":109121,"significance":"If the stated results are corrected, the paper would make a useful contribution: it extends average-cost mean-field games beyond compact state spaces using drift and minorization conditions, and it gives a relatively clean approximate-Nash argument based on invariant measures rather than transient analysis. The dynamic-programming approach through the average-cost optimality equation and the fixed-point formulation on joint state-action measures is natural, and the closed-graph proof in Proposition 4.3 is careful and substantially self-contained. The paper also clearly delineates the assumptions that make the average-cost criterion tractable. However, the manuscript as written contains a load-bearing inconsistency concerning the prescribed initial distribution mu0, and a separate incorrect product drift inequality in Section 5, so the central theorems are not yet established in the form stated.","major_comments":[{"comment":"The definition of Psi(mu) requires both optimality for mu and the condition mu0=mu. In the fixed-point proof, Gamma is defined on Xi and a fixed point nu is shown to satisfy nu in C(nu) ∩ B(nu); Proposition 4.2 then declares (pi,nu1) a mean-field equilibrium. The proof establishes (i) that nu1 is invariant under pi and (ii), via Theorem 4.1, that pi is optimal for nu1 when the initial state is distributed as nu1. It never verifies that the prescribed initial distribution mu0 equals nu1. This is not a mere presentation gap: the statement is false for arbitrary mu0. For example, let X={0,1}, A={a}, p(0|x,a,mu)=1 for all x,a,mu, lambda=0.5 delta_0, w≡1, alpha=0.75, c≡1, and mu0=delta_1. Assumption 1 holds, but the only invariant measure under any policy is delta_0, while Psi(delta_0) is empty because mu0=delta_0 fails; no pair satisfies Definition 2.1. To repair the claim, either remove the clause mu0=mu from Psi and prove, under Assumption 1, that the average cost is independent of the initial distribution (the minorization condition gives uniform ergodicity, so this should be possible), or state and prove Theorem 2.2 only for mu0 equal to the equilibrium invariant marginal produced by the fixed point. The current proof supports only the latter, and even then the statement must say so.","section":"Definition 2.1, Section 4, Proposition 4.2"},{"comment":"Theorem 4.1 proves optimality of a policy for mu only when x(0) ~ mu_{pi,mu}, the invariant distribution of the policy. However, the average cost J_mu(pi) in Section 2 is defined with the prescribed initial distribution x(0) ~ mu0. The proof of Theorem 4.1 works with J_{mu,n}(pi,h_mu,mu_{pi,mu}) and therefore establishes that the policy achieves the optimal average cost from its own invariant distribution. The equality of the average cost from mu0 and from mu_{pi,mu} is neither stated nor proved. This is the technical point on which the previous comment turns; it should be addressed explicitly when Definition 2.1 is corrected.","section":"Theorem 4.1 and Section 2"},{"comment":"The claimed product drift inequality (5.2) is not a consequence of Assumption 1(e). With w_N = prod_i w(y_i) and phat_N = p^N - lambda^N, the left-hand side equals prod_i (∫ w dp(·|x_i,a_i)) - (∫ w dlambda)^N. For C = ∫ w dlambda >0 and w(x_i)=L for all i, this is (alpha L + C)^N - C^N, which is larger than alpha^N L^N; for instance alpha=0.5, C=1, L=100, N=2 gives 2601 > 2501. Thus (5.2) is false as stated. The uniqueness of the invariant distribution of the N-player chain can be obtained from the product minorization (5.1) alone, since uniform minorization implies uniform ergodicity, so the proof is repairable by replacing (5.2) with a valid argument or removing it from the hypotheses used to invoke [30,12].","section":"Section 5, Eqs. (5.1)-(5.2)"}],"minor_comments":[{"comment":"In the displayed expression for u_{k+1}^{(n)} - u_{k+1}, the second minimand contains a spurious factor beta multiplying the integral of u_k; the operator T has no such factor.","section":"Proof of Proposition 4.4"},{"comment":"Theorem 2.2 states the game as (X,A,p,c), whereas the model introduced in Section 2 is (X,A,p,c,mu0). This ambiguity is directly connected to Major Comment 1 and should be resolved in a revision.","section":"Theorem 2.2 statement vs. Section 2"},{"comment":"The clause 'and mu0=mu' in the definition of Psi(mu) is unusual for a stationary mean-field equilibrium and is never used in the fixed-point construction; please clarify whether the initial distribution is meant to constrain the equilibrium marginal or whether a stationary MFE without that constraint is intended.","section":"Definition 2.1"}],"recommendation":"major_revision","confidential_remarks":"The main issue is the mismatch between the formal definition of an MFE and the fixed-point construction: the proof constructs an invariant measure and proves optimality from that invariant measure, while the definition imposes mu0 = mu. This is fixable by removing the mu0 = mu clause and adding a short uniform-ergodicity argument, or by restricting the theorem to mu0 equal to the equilibrium invariant marginal. The incorrect product drift inequality in Section 5 is also fixable, since product minorization alone suffices for the uniqueness claim used there. These corrections do not require new conceptual machinery, so I would not recommend rejection on these grounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core here is real: under drift and minorization, the paper extends discrete-time average-cost mean-field game existence to Polish, possibly non-compact state spaces with mean-field-dependent transitions, going beyond Wiecek's compact-space setting and Biswas's measure-independent dynamics. The proof strategy—average-cost optimality equation plus Kakutani—is coherent, and the approximate-Nash part for finite-player games is a reasonable adaptation of existing coupling arguments. The citation pattern looks fair; the self-citation to the author's discounted-cost paper is legitimate because the graph-closure argument is adapted from there.\n\nThe central problem is Definition 2.1. The set Ψ(μ) is defined to require both optimality for μ and μ0 = μ, so a mean-field equilibrium (π, μ) must have the given initial distribution equal to the invariant measure. The fixed-point proof never checks this. Theorem 4.1 only establishes optimality when the initial distribution is the invariant measure μ_{π,μ}, not the prescribed μ0. Proposition 4.2 then says the fixed point yields an equilibrium, but it silently treats 'optimal when x(0) ~ ν1' as satisfying the μ0 condition. That is not just a cosmetic gap: take X={0,1}, A={a}, p(0|x,a,μ)=1 for all x,a,μ, λ=0.5δ0, w constant, α=0.75, c≡1, and μ0=δ1. Assumption 1 holds, the only invariant measure is δ0, and Ψ(δ0) is empty because μ0≠δ0. So Theorem 2.2 is false as stated for arbitrary μ0.\n\nI suspect this is a typo in the definition, and the fix is likely easy: drop the μ0=μ condition and prove, or cite, that the average cost is independent of the initial distribution under Assumption 1. There is enough ergodicity in the drift/minorization assumptions that this should be true. Alternatively, state the theorem only for μ0 equal to the equilibrium invariant measure. But as written, the proof and the statement do not line up.\n\nMinor issues: there is a typo in Proposition 5.1 where the displayed ρ equality looks wrong, and Theorem 4.1's wording about initial distributions could confuse readers. Neither is load-bearing once the main definition is fixed.\n\nWho should read this: people working on discrete-time mean-field games or average-cost MDPs with ergodic conditions. The paper deserves a serious referee, but the referee should be asked to pin down the equilibrium definition and the role of μ0. I would send it out, not desk-reject it.","headline":"A genuinely broader existence result for average-cost mean-field games, but the main theorem as stated has a hole involving the initial distribution μ0.","tokens_in":16084,"tokens_out":3171,"would_cite":false,"duration_ms":36069,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A15","91A10","91A13","93E20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Mean-field equilibria exist under drift and minorization conditions","keywords":["mean-field games","average cost","approximate Nash equilibrium","Polish spaces","average cost optimality equation","Kakutani fixed point theorem","drift condition","minorization condition"],"falsifier":"Construct a weakly continuous, bounded-cost model on a non-compact state space with compact action set, such as a shifted exponential transition p(dy|x)=$e^{{-(y-x)}}$1_{y>=x} dy on X=[0,infty), which has no common sub-probability minorant, and check numerically whether a mean-field equilibrium still exists; existence there would show the conditions are not necessary, while nonexistence under the other assumptions would show the minorization condition is load-bearing. To test the proof itself, find a weakly convergent sequence mu_n -> mu for which the fixed points h_{mu_n} of the average cost optimality equation do not converge uniformly on a compact set, which would break Proposition 4.4 and the closed-graph step of the fixed-point argument.","tokens_in":15062,"feed_emoji":"🎲","tokens_out":6432,"duration_ms":62232,"temperature":0.7,"pith_summary":"This paper proves that discrete-time mean-field games with average-cost payoffs have a mean-field equilibrium when the state and action spaces are Polish and the transition kernel satisfies a uniform minorization condition together with a drift inequality. A mean-field equilibrium is a pair of a policy and a state distribution that are consistent: the policy minimizes the infinite-horizon average cost against that distribution, and the distribution is the invariant law generated by the policy. The paper further proves that this equilibrium policy, used by every agent in an N-agent version of the game, is an epsilon-Nash equilibrium for every epsilon greater than zero once N is large enough. These results matter because they extend average-cost mean-field-game theory beyond compact or finite state spaces and give a dynamic-programming route, through the average cost optimality equation, to approximate equilibria in large anonymous games.","feed_headline":"Mean-field equilibria exist under drift and minorization conditions","feed_subtitle":"A uniform drift and minorization bound lets one policy solve large-player average-cost games approximately.","key_machinery":"The load-bearing object is the average cost optimality equation written through the Bellman operator T_mu u(x) = min_{a in A} [ c(x,a,mu) + integral_X u(y) (p(dy|x,a,mu) - lambda(dy)) ]. The minorization condition makes T_mu a contraction on bounded continuous functions with modulus beta = 1 - lambda(X), so each mu has a unique fixed point h_mu that encodes the optimal average cost and the optimality condition. The drift inequality defines the compact set P_c(X) of measures with integral w dmu ≤ integral w dlambda / (1 - alpha), and the corresponding set Xi of joint state-action measures is compact and convex. A fixed point of the set-valued map Gamma(nu) = C(nu) ∩ B(nu), where C imposes invariant-measure consistency and B imposes optimality, yields the mean-field equilibrium via Kakutani's fixed point theorem.","core_discovery":"The paper's central claim is Theorem 2.2: under Assumption 1, consisting of a bounded continuous cost, a weakly continuous transition kernel, a compact action set, a uniform minorization condition p(·|x,a,mu) ≥ lambda(·), and a uniform drift inequality with a single moment function w and constant alpha < 1, the mean-field game admits a mean-field equilibrium (pi*, mu*). Theorem 3.3 then states that, if all N agents adopt pi*, the resulting N-tuple of policies is an epsilon-Nash equilibrium for the finite N-player game for every epsilon > 0 and all sufficiently large N. The equilibrium is obtained by solving the average cost optimality equation for each candidate mean-field measure mu, extracting an optimal policy, and then applying a fixed-point argument to make mu consistent with the invariant distribution of the optimally controlled process.","pith_inferences":["The uniform minorization and drift assumptions are probably stronger than necessary; a natural test is whether the theorem survives with a state-dependent minorant or with the drift inequality required only on states reachable under equilibrium play.","The epsilon-Nash proof in the paper uses the simplification that transitions do not depend on mu; if p depended on mu in a Lipschitz way, the same extended-state comparison would likely acquire an extra term proportional to the distance between mean-field measures, yielding a similar approximation bound with a modified error.","The belief-state transformation suggested in the conclusion points to a partially observed version becoming a fully observed mean-field game on the space of posterior distributions, provided the filter transition inherits the minorization and drift conditions."],"forward_implications":["For any average-cost mean-field game satisfying Assumption 1, an equilibrium policy and a consistent state distribution exist, covering non-compact Polish state spaces rather than only finite or compact ones.","The mean-field equilibrium policy is asymptotically optimal for the finite-player game: for each epsilon > 0 there is a threshold N(epsilon) beyond which no agent can improve by more than epsilon through a unilateral deviation.","The existence proof runs through the average cost optimality equation and a fixed-point argument, so the same dynamic-programming machinery can be reused in related infinite-horizon control problems.","Because all agents are identical, the same equilibrium policy works simultaneously for every player, so the epsilon-Nash property does not require agent-specific policies or centralized coordination."],"supporting_citations":[{"why":"Supplies the average cost optimality equation, measurable selection, the Ionescu Tulcea theorem, and the compactness criterion for moment functions used throughout.","marker":"[13]"},{"why":"Provides the fixed-point approach to the average cost optimality equation and existence of a unique invariant measure under the minorization condition.","marker":"[30]"},{"why":"Gives the contraction theorem for the Bellman operator on bounded continuous functions.","marker":"[27]"},{"why":"Provides the fixed-point theorem for convex-valued correspondences used to close the existence argument.","marker":"[2]"},{"why":"Is the discounted-cost predecessor whose closed-graph proof is adapted here to the average-cost setting.","marker":"[26]"},{"why":"Supplies the continuous-convergence theorem for integrals with varying measures used in the closed-graph proof.","marker":"[19]"},{"why":"Closest prior average-cost result on compact spaces; its extended-state formulation is adapted for the finite-player approximation theorem.","marker":"[31]"},{"why":"Provides the unique invariant measure lemma used to connect optimal policies with their stationary state distributions.","marker":"[12]"}],"fun_headline_variants":["Mean-field equilibria exist under drift and minorization","Drift and minorization ensure mean-field equilibrium","Mean-field policy is epsilon-Nash for large finite games","Average-cost mean-field equilibria on Polish spaces","Discrete-time mean-field equilibrium from drift and minorization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire argument rests on one uniform bound: a single sub-probability measure lambda and a single moment function w must control the transition kernel for every state, action, and mean-field measure, and if that uniformity fails, the contraction property and the compactness of the candidate equilibrium set are lost.","fun_headline_variants_meta":{"raw":{"variants":["Mean-field equilibria exist under drift and minorization","Drift and minorization ensure mean-field equilibrium","Mean-field policy is epsilon-Nash for large finite games","Average-cost mean-field equilibria on Polish spaces","Discrete-time mean-field equilibrium from drift and minorization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001652,"raw_usage":{"total_tokens":6511,"prompt_tokens":843,"completion_tokens":5668,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":5592}},"tokens_in":459,"tokens_out":5668,"duration_ms":36785,"temperature":1.0,"reasoning_tokens":5592,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:46:39.622638+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a weakly continuous, bounded-cost model on a non-compact state space with compact action set, such as a shifted exponential transition p(dy|x)=$e^{{-(y-x)}}$1_{y>=x} dy on X=[0,infty), which has no common sub-probability minorant, and check numerically whether a mean-field equilibrium still exists; existence there would show the conditions are not necessary, while nonexistence under the other assumptions would show the minorization condition is load-bearing. To test the proof itself, find a weakly convergent sequence mu_n -> mu for which the fixed points h_{mu_n} of the average cost optimality equation do not converge uniformly on a compact set, which would break Proposition 4.4 and the closed-graph step of the fixed-point argument.","supporting_citations":[{"cited_title":"and Lasserre, J","cited_arxiv_id":null,"evidence_quote":"Supplies the average cost optimality equation, measurable selection, the Ionescu Tulcea theorem, and the compactness criterion for moment functions used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the fixed-point approach to the average cost optimality equation and existence of a unique invariant measure under the minorization condition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the contraction theorem for the Bellman operator on bounded continuous functions."},{"cited_title":"and Border, K","cited_arxiv_id":null,"evidence_quote":"Provides the fixed-point theorem for convex-valued correspondences used to close the existence argument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the discounted-cost predecessor whose closed-graph proof is adapted here to the average-cost setting."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the continuous-convergence theorem for integrals with varying measures used in the closed-graph proof."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior average-cost result on compact spaces; its extended-state formulation is adapted for the finite-player approximation theorem."},{"cited_title":"and Hernandez-Lerma, O","cited_arxiv_id":null,"evidence_quote":"Provides the unique invariant measure lemma used to connect optimal policies with their stationary state distributions."}],"review_version":1}