{"id":"b6849e20-39c4-4622-b747-f92bded9e520","arxiv_id":"2608.12775","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Existence of an approximate Stackelberg equilibrium is proved for a mean-field model of carbon emissions where a price-setting leader and heterogeneous production regions interact under a hard emission cap enforced by reflection.","lead":"This paper builds a two-level game model in which a central regulator sets product prices and regions choose production, while a hard emission cap is enforced by automatically reflecting the pollution stock back below the cap. It proves that, as the number of regions grows, an approximate Stackelberg equilibrium exists and illustrates the equilibrium with Monte Carlo simulations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Leader optimality hinges on t-Lipschitz regularity of the running-supremum density that Lemma 4.3 does not prove; Proposition 4.2's separate argument for this is not justified.","rationale":"The reader's weakest_assumption correctly identifies the running-supremum density regularity as load-bearing for the leader's optimality, but it slightly mislocates the gap: Lemma 4.3 only asserts local boundedness, while the t-Lipschitz property is needed later in Proposition 4.2 and is not established by the text. This is precisely where the proof of Corollary 4.4 depends on an unproved regularity assertion. I do not raise the omitted proof of Lemma 3.3 as the primary concern, because the best-response computation there is a pointwise quadratic optimization once the mean field m is fixed; that omission is not the most insecure step. The ANE proof in Theorem 3.6 is also compressed, but the stated convergence (3.14) and (3.16) is at least plausible and is given a partial appendix proof, so I would not place the decisive weight there. The paper has substantial structure and a plausible central theorem; the concern is a proof gap, not a demonstrated falsehood, so conditional acceptance remains the appropriate verdict if the t-regularity step can be completed. If it cannot, the leader-optimality verification and hence Theorem 4.5 are unproven.","tokens_in":29241,"tokens_out":27219,"duration_ms":296895,"concrete_test":"Independently prove the local Lipschitz continuity in t of p(s,u;t,z) from the Malliavin derivative (4.16), by computing for smooth bounded f the derivative d/dt E[f(sup_{r in [t,s]} U^{t,z}(r))] via Malliavin integration by parts. If the calculation cannot be closed without a t-Lipschitz bound on p that is itself the conclusion, or if it produces a bound involving E[K^t(s)^{-1} |W(t)|] that is not uniformly finite on compact subsets, the gap in Proposition 4.2 is confirmed; if the derivative is bounded directly, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 4.5 depends on Corollary 4.4, whose verification uses Proposition 4.2 to assert that Psi belongs to W^{1,2,1;infinity}_{loc}([0,T] x A) and then applies the generalized Ito formula. The proof of Proposition 4.2 needs t-regularity of the density p(s,u;t,z) of sup_{r in [t,s]} U^{t,z}(r): the term L4/h in (4.24) is controlled by asserting that its second term is locally bounded by the local Lipschitz continuity of p(s,u;t,z) w.r.t. t. However, Lemma 4.3 states and proves only local boundedness of p and of p^{Q1}, p^{Q2}; it does not establish t-Lipschitz continuity. The separate argument inside Proposition 4.2 claims that P(sup_{r in [t,s]} V^{t,z}(r) <= m) is locally Lipschitz in t 'by using local Lipschitz continuity of E_t', where E_t is the Girsanov density built from W and Z^{t,z}. This is not proved and is not an immediate consequence of the preceding time-change computation: E_t depends on t through W(t) and through the solution Z^{t,z}, and no estimate of its Lipschitz constant in t is supplied. If p is only locally bounded, the L4/h estimate does not close, Psi may fail to be in the stated Sobolev space, and the verification argument for the leader's optimality in Corollary 4.4 loses its foundation. The existence of an (epsilon,0)-Stackelberg equilibrium is therefore not established by the proof as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies a finite-horizon Stackelberg differential game with one regulator (leader) and n heterogeneous regions (followers), where the average cumulative emission process is constrained by a deterministic cap Z through a reflected SDE with a local-time abatement process A. The leader chooses a price-adjustment process u; the followers choose production processes q_i to maximize profits. For fixed u, the followers' problem is solved as an MFG, yielding an epsilon(n)-Nash equilibrium. The leader's HJB equation is transformed via a Cole-Hopf change of variables into a linear Robin problem, solved probabilistically by a reflected geometric Brownian motion; the authors use Malliavin calculus for the density of the running supremum and a generalized Ito formula for verification, obtaining a semi-explicit optimal strategy u*. The paper concludes that the resulting strategies form an (epsilon,0)-Stackelberg equilibrium and presents Monte Carlo sensitivity analysis.","tokens_in":29567,"tokens_out":9653,"duration_ms":103130,"significance":"If correct, the paper would be a meaningful contribution: it would provide the first approximate Stackelberg equilibrium for a hard-cap, mean-field, common-noise emission game with a semi-explicit leader strategy and non-trivial policy trade-offs. The Malliavin-based treatment of the running supremum density for a reflected geometric Brownian motion is a genuine technical idea, and the numerical comparative statics are informative and falsifiable. However, the central leader-verification argument depends on a regularity property that is asserted but not proved, and the follower ANE proof outsources a key step to lecture notes. The contribution is therefore conditional on completing and verifying those components.","major_comments":[{"comment":"The regularity p(s,u;t,z) locally Lipschitz in t is load-bearing but unproved. Lemma 4.3 establishes only local boundedness of the densities p, p^{Q1}, and p^{Q2}; it contains no t-Lipschitz estimate. In the proof of Proposition 4.2, the claim that P(sup_{r in [t,s]} V^{t,z}(r) <= m) is locally Lipschitz in t is justified only by 'local Lipschitz continuity of E_t', but no Lipschitz estimate for E_t in t is supplied, and E_t depends on t through W(t) and the solution Z^{t,z}. The subsequent Girsanov/time-change paragraph also states an identity between the laws of V^{t,z} and U^{t,z} under Q and P that is not established as written, since V^{t,z} is defined using W rather than the Q-Brownian motion. Consequently, the bound on the term L4/h in (4.24) does not close, Psi may fail to belong to W^{1,2,1;infinity}_{loc}([0,T] x A), and the verification in Corollary 4.4, and hence Theorem 4.5, is not proved as written.","section":"Section 4.3, proof of Proposition 4.2 and Eq. (4.24)"},{"comment":"Lemma 3.3 states the unique best response q^{*,u}(t) = (2c)^{-1}(a+u(t)X^{*,u}(t;m_q^u)) and the running-supremum representation of A^{*,u}, but its proof is omitted. This lemma is the foundation for the MFE in Lemma 3.5 and for the finite-n ANE in Theorem 3.6. The DPP/HJB verification for a reflected SDE whose drift depends on the control and whose regulator is the running supremum of a controlled diffusion is not a one-line argument. The authors need to provide the proof or a precise reference with conditions under which this verification is valid.","section":"Section 3.2, Lemma 3.3"},{"comment":"The epsilon-Nash property is not actually verified in the proof. The argument establishes the state-process convergences (3.14) and (3.16) and then says that the conclusion follows 'by adopting a similar argument in the proof of Theorem 8.3 of Lacker (2018).' It does not demonstrate, for each i, the payoff comparison J_i^{(n)}(q^{*, (n)}(u)) >= sup_{q_i in U_i} J_i^{(n)}(q_i, q^{*, (n)}_{-i}(u)) - epsilon(n) with a quantified epsilon(n) -> 0, nor does it control the deviating player's payoff through the reflected dynamics (3.15), where A^{*,u,(n)}_{-i} itself depends on the deviating control q_i. Without this, the ANE half of the Stackelberg equilibrium is asserted rather than proved.","section":"Section 3.3, proof of Theorem 3.6"}],"minor_comments":[{"comment":"The symbol b in the displayed formula q^{*,u}(t) = (a - b m_q^{*,u}(t))/(2c) is undefined and is inconsistent with the subsequent formula (3.5). Please correct the expression or remove the undefined symbol.","section":"Section 3.2, Lemma 3.5"},{"comment":"The auxiliary follower problem (3.12) contains an exponential discount factor e^{-rho t}, but the parameter rho is never defined and the discount factor is absent from the original follower objective (2.7). Please clarify the role of rho and explain why the discounted auxiliary payoff comparison is valid for the undiscounted finite-n game.","section":"Section 3.3, Eq. (3.12)"},{"comment":"The in-text citation 'Nualart and Nualart 2018' does not match the reference list, which lists Nualart and Nualart (2006). Please harmonize the citation and the bibliographic entry.","section":"References and text"},{"comment":"The Monte Carlo figures would be more informative with the number of paths, time-discretization step, and confidence bands or standard errors reported; this would help assess whether the reported monotonicity patterns are numerically stable.","section":"Section 5, numerical analysis"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unproved t-Lipschitz regularity of the running-supremum density used in Proposition 4.2; this is the key load-bearing point for the leader's optimality and hence for Theorem 4.5. I recommend asking the authors for a complete proof of that regularity or a modified argument that avoids it. The paper also relies heavily on the authors' own prior work and on Lacker's lecture notes for the ANE step; the final version should make those dependencies precise and ideally self-contained for the advertised claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper does something genuinely new: it combines a leader–follower Stackelberg hierarchy with mean-field interaction and a hard emission cap enforced by state reflection, and it gives an approximate equilibrium theorem with a semi-explicit leader strategy. That combination is not in earlier Bo et al. work, which had reflection or MFG but not the two-level price mechanism. The modeling of transboundary pollution via a Skorokhod-reflected state is clean, and reducing the leader's HJB equation through Cole–Hopf to a linear Robin PDE with a probabilistic representation is a sensible route. The numerical section illustrates a plausible, policy-relevant insight: a looser cap can contract output through the leader's pre-emptive price adjustment, while faster tightening can support output with more abatement.\n\nThe soft spots are real but uneven. Lemma 3.3's proof is omitted, and that lemma carries the follower's best response. Lemma 3.5 uses an undefined b in q* = (a - b m)/(2c), which makes the MFE formula unreadable as written. Theorem 3.6 outsources the ANE argument to Lacker's lecture notes and gives only convergence sketches in the appendix; for a paper whose main theorem is approximate equilibrium, this is weak. The numerics are illustrative only, with no code, data, or statistical detail.\n\nThe load-bearing gap is in Proposition 4.2. To show t -> Psi is locally Lipschitz, the proof needs the density p(s,u;t,z) of the running supremum to be locally Lipschitz in t. Lemma 4.3 only proves local boundedness. The assertion that local Lipschitz continuity of the Girsanov density E_t gives the needed t-regularity is not demonstrated, and no estimate of E_t's Lipschitz constant in t is supplied. The L4/h bound does not close without that regularity. Since Corollary 4.4 and Theorem 4.5 depend on this verification, the leader's optimality and the approximate Stackelberg equilibrium are not fully established as written.\n\nIf the density regularity can be supplied—and it may well be true, given known results for reflected diffusions—the framework and the main theorem hold up. As written, the proof is incomplete at a specific, identifiable point. This paper is for researchers in mean-field games with reflection, hierarchical stochastic control, and climate policy modeling. It deserves a serious referee: the problem is worth solving, and the gap is likely fixable. I would not cite the main theorem in its current form, but I would follow a revision closely. Send it to peer review with a request to fill Lemma 3.3 and Proposition 4.2 and fix the b typo.","headline":"Genuinely new hierarchical mean-field Stackelberg model with reflection, but the leader's verification rests on an unproved t-Lipschitz density regularity and several deferred proofs.","tokens_in":30111,"tokens_out":3316,"would_cite":false,"duration_ms":35542,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A16","91A23","91A65","93E20","60H30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A hard-cap carbon game has an asymptotically exact Stackelberg equilibrium, with a semi-explicit optimal price rule.","keywords":["carbon emissions","transboundary pollution","mean field Stackelberg game","state reflection","local time","approximate Stackelberg equilibrium","reflected geometric Brownian motion","stochastic optimal control"],"falsifier":"Choose a simple cap schedule (for instance a constant or linear $\\Phi$), fix the parameters, and estimate the density of $\\sup_{r\\in[t,s]}U^{t,z}(r)$ near the boundary where the process touches the cap using Monte Carlo or an analytic calculation; if the density's derivative in $t$ or in the level is unbounded at any parameter point, Lemma 4.3 fails and the leader's optimality proof collapses.","tokens_in":29009,"feed_emoji":"🌍","tokens_out":10289,"duration_ms":99433,"temperature":0.7,"pith_summary":"This paper tries to establish that a two-level carbon governance game remains solvable when cumulative emissions are forced below a hard cap. One central regulator sets a price-adjustment policy first; $n$ heterogeneous regions then choose production to maximize profits, and their emissions aggregate into a common reflected state that cannot exceed a deterministic ceiling. The authors prove that, for large $n$, these strategies form an approximate Stackelberg equilibrium: the error vanishes as $n$ grows, and the leader's optimal price adjustment has a semi-explicit formula. The formula is an expectation involving a reflected geometric Brownian motion and its local time, so it can be evaluated by Monte Carlo simulation. If the proof is right, hard emission caps can be embedded in hierarchical mean-field policy games rather than approximated only through soft penalties.","feed_headline":"Hard emission caps have a solvable Stackelberg equilibrium","feed_subtitle":"As regions grow numerous, price-setting and regional output form an asymptotically exact carbon equilibrium.","key_machinery":"The load-bearing object is the dynamic Skorokhod map: the reflected state is the unreflected process minus the running supremum of its excess over the cap, so the hard cap is enforced by a local-time process that increases only when emissions touch the ceiling. For the leader's problem, a Cole-Hopf transformation turns the nonlinear HJB equation into a linear Robin PDE whose probabilistic representation is built from a reflected geometric Brownian motion $Y^{t,x,z}$ together with its local time $R^{t,x,z}$. The regularity engine is Lemma 4.3: the running supremum of the auxiliary process $U^{t,z}$ has locally bounded densities, locally Lipschitz in time, under the original measure and under two Girsanov tilts. That regularity makes the gradient of the representation well defined and justifies the feedback formula for $u^{*,(n)}$; the final verification uses a generalized Itô formula for Sobolev-space solutions.","core_discovery":"The central claim is Theorem 4.5: the leader's strategy $u^{*,n}$ in (4.14) and the followers' strategies $q_i^{*,n}(u)$ in (3.8) constitute an $(\\epsilon,0)$-Stackelberg equilibrium in the sense of Definition 2.2, with $\\epsilon(n)\\to 0$. The leader's optimal price adjustment has the probabilistic representation $$$u^{{*,(n)}}$(t,x,z)=-\\frac{\\mu x}{2c_P}\\$nu^{{(n)}}$\\Big(\\frac1{2c}\\Big)\\, \\frac{ \\mathbb{E}\\,[$e^{{-\\eta(c_A\\nabla_x R^{t,x,z}}$(T)+c_E\\nabla_x $Y^{{t,x,z}}$(T))}]}{ \\mathbb{E}\\,[$e^{{-\\eta(c_A R^{t,x,z}}$(T)+c_E $Y^{{t,x,z}}$(T))}]},$$ with the gradients expressed explicitly through the hitting time $\\tau^t_{x,z}$ and the geometric factor $K^t(s)$. The proof follows the standard backward-induction route: solve the followers' mean-field game to obtain an $\\epsilon$-Nash equilibrium, then verify the leader's control against that response, using a Malliavin-calculus regularity result for the running maximum of a transformed process.","pith_inferences":["The same reflected-state Stackelberg machinery should transfer to other common-pool resource problems with hard ceilings and common noise, such as water quotas, fishery catch limits, or congestion caps, wherever the state is a positive multiplicative diffusion.","If a fully explicit density for the running supremum of the reflected geometric Brownian motion were found, the approximate equilibrium could be upgraded to an explicit one and the Malliavin regularity step would become unnecessary.","The 'looser cap contracts output' outcome is demonstrated for one cost structure; testing it across abatement cost functions and price-sensitivity parameters would reveal whether it is a general mechanism or a parameter-dependent artifact.","The leader's exact optimality assumes followers coordinate on the approximate-Nash strategies; in a decentralized rollout, followers' miscoordination would feed back into the leader's cost and reintroduce error at both levels."],"forward_implications":["For large region populations the equilibrium is asymptotically exact, so the leader can compute her price policy from one reflected-geometric-Brownian-motion expectation while regions follow the linear rule $q_i=(a_i+uX)/(2c_i)$.","The leader's layer is exactly optimal against the followers' approximate best responses: the only vanishing error is the followers' $\\epsilon$-Nash error, so $\\epsilon_2=0$ in the approximate Stackelberg equilibrium.","The mean-field consistency condition reduces to $m_q^*(t)=\\nu(a/2c)+\\nu(1/2c)u(t)X^*(t)$, which gives a closed feedback structure linking pricing, regional production, and the emission state.","Simulations show a policy trade-off: a looser long-run cap can trigger pre-emptive price cuts that contract output, while a faster-tightening cap can be paired with price support and higher abatement to keep emissions under the lower ceiling."],"supporting_citations":[{"why":"Supplies the approximate $(\\epsilon_1,\\epsilon_2)$-Stackelberg equilibrium notion that lets the authors replace exact equilibrium with vanishing-error strategies.","marker":"Moon and Başar 2018"},{"why":"Provides the mean-field approximation argument used to turn the $n$-follower Nash game into an $\\epsilon$-Nash equilibrium.","marker":"Lacker 2018"},{"why":"Gives the mean-field game of controls with state reflections and the dynamic Skorokhod formulation on which the reflected emission dynamics rest.","marker":"Bo et al. 2025"},{"why":"Gives the explicit running-supremum form of the solution to the dynamic Skorokhod problem used throughout the paper.","marker":"PiliPenko 2014"},{"why":"Supplies the locally bounded density result for the running maximum that Lemma 4.3 adapts to the transformed process.","marker":"Coutin and Pontier 2019"},{"why":"Provides the Malliavin calculus tools used to prove existence and regularity of the density of the supremum.","marker":"Nualart and Nualart 2018"},{"why":"Supplies the generalized Itô formula for Sobolev-space functions used in the verification of the leader's optimality.","marker":"Krylov 2009"}],"fun_headline_variants":["One price rule to cap them all: mean-field Stackelberg equilibrium","Hard carbon caps, infinite regions: a solvable pricing game","Leader sets prices, regions cut emissions: exact equilibrium emerges","Mean-field Stackelberg game: reflection enforces the cap","As regions scale up, carbon pricing hits equilibrium exactly"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof of the leader's optimality rests on a regularity estimate: the probability density of the running maximum of a certain transformed emission process must be locally Lipschitz in time and in the starting value, under both the original measure and two change-of-measure tilts; if that estimate fails, the leader's optimal-control characterization and the equilibrium theorem lose their proof.","fun_headline_variants_meta":{"raw":{"variants":["One price rule to cap them all: mean-field Stackelberg equilibrium","Hard carbon caps, infinite regions: a solvable pricing game","Leader sets prices, regions cut emissions: exact equilibrium emerges","Mean-field Stackelberg game: reflection enforces the cap","As regions scale up, carbon pricing hits equilibrium exactly"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001373,"raw_usage":{"total_tokens":5545,"prompt_tokens":906,"completion_tokens":4639,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":4554}},"tokens_in":522,"tokens_out":4639,"duration_ms":37542,"temperature":1.0,"reasoning_tokens":4554,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:34:19.478229+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Choose a simple cap schedule (for instance a constant or linear $\\Phi$), fix the parameters, and estimate the density of $\\sup_{r\\in[t,s]}U^{t,z}(r)$ near the boundary where the process touches the cap using Monte Carlo or an analytic calculation; if the density's derivative in $t$ or in the level is unbounded at any parameter point, Lemma 4.3 fails and the leader's optimality proof collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the approximate $(\\epsilon_1,\\epsilon_2)$-Stackelberg equilibrium notion that lets the authors replace exact equilibrium with vanishing-error strategies."},{"cited_title":"(2018): Mean field game and interacting particle systems","cited_arxiv_id":null,"evidence_quote":"Provides the mean-field approximation argument used to turn the $n$-follower Nash game into an $\\epsilon$-Nash equilibrium."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the explicit running-supremum form of the solution to the dynamic Skorokhod problem used throughout the paper."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the locally bounded density result for the running maximum that Lemma 4.3 adapts to the transformed process."},{"cited_title":"(2009): Controlled Diffusion Processes","cited_arxiv_id":null,"evidence_quote":"Supplies the generalized Itô formula for Sobolev-space functions used in the verification of the leader's optimality."}],"review_version":1}