REVIEW 3 major objections 6 minor 54 references
A new probabilistic approach for mean field games of optimal stopping
T0 review · 3 major / 6 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Mean-field optimal stopping equilibria in randomized strategies coincide exactly with solutions of a new coupled reflected forward-backward McKean–Vlasov system.
desk verdict A substantial, mostly rigorous new FBSDE characterization of randomized-optimal-stopping MFGs; the KFG-based existence is solid, but the Tarski extremal/learning track leans on an unverified structural monotonicity assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the coupled MKV-RFBSDE system: a forward-backward system in which the randomized stopping strategy L is solved for as part of the unknown, rather than recovered from an external flow of measures. The two integral conditions on L—that the measure −dL be supported on the contact set of the value with the obstacle and on the flat set of the reflection process—are the mechanism that makes a candidate survival process an optimal response.
What would settle it
In a one-dimensional Markovian example with Brownian state, f=0, and h(t,x)=x, compute the reflected BSDE candidate and the survival process L; the equivalence predicts that the support of −dL is contained in {Y=ξ}∩{A=0}, so any positive stopping mass outside that set would refute the characterization. Alternatively, produce L≤L′ satisfying the standing assumptions for which the gap Y−h is not ordered, which would falsify the monotonicity assumption behind the Tarski route.
Extended reading notes
Core claim
The central discovery is that randomized-strategy equilibria of optimal-stopping mean field games are not a separate fixed-point object: they are exactly the L-component of a solution to a coupled reflected forward-backward McKean–Vlasov system. In the system, L is an adapted, non-increasing survival process taking values in [0,1], the state X evolves with coefficients averaged against the surviving population, and the reflected backward component (Y,Z,A) solves an RBSDE with obstacle built from L. Optimality is encoded by two contact conditions: the measure −dL can charge only times where Y equals the obstacle, and only times where the reflection process A has not yet increased. The paper p
Load-bearing premise
The most fragile premise is the order-preservation assumption that whenever one survival process stays below another, the gap between the continuation value and the stopping payoff is pointwise ordered in the same direction; the Tarski-based existence of extremal equilibria and the learning algorithms collapse if this monotonicity fails.
Editorial extensions
If this is right
- Randomized mean-field equilibria exist under the paper's assumptions, so pure-strategy non-existence is circumvented by allowing survival processes.
- Any solution of the coupled system is automatically an equilibrium, and any equilibrium produces a solution of the system; the game and the system are the same problem.
- Under monotonicity assumptions there are minimal and maximal equilibria, ordered by survival probability, with iterative learning schemes that converge to them.
- A mean-field equilibrium induces an ε-Nash equilibrium for the N-player stopping game, with the approximation error vanishing as N grows.
- The probabilistic system is equivalent to a constrained obstacle-problem PDE system, providing an analytic route to the same equilibria.
Reading between the lines
- Because the optimality conditions are stated purely through contact sets of Y and A, the same two-condition test could serve as a Snell-envelope-style criterion for randomized stopping in single-agent problems.
- The fixed point lives directly on survival processes, which suggests a natural numerical loop—solve the RBSDE for a given L, update L by the contact sets, iterate—that the order-theoretic route shows converges to extremal equilibria in monotone settings.
- The non-Markovian extension noted in the paper indicates the two contact conditions are filtration-relative, so the characterization may persist with common noise or partial information.
- The finite-player approximation is stated for i.i.d. copies of the equilibrium strategy; a natural continuation is to quantify deviations under dependent initial data or common noise.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a probabilistic formulation of mean-field games of optimal stopping (OS-MFGs) with randomized stopping strategies, based on a coupled reflected forward-backward McKean–Vlasov system (2.2)-(2.5). The equilibrium object is a quintuple (X, Y, Z, A, L), where L is a [0,1]-valued, non-increasing càdlàg process representing the survival/randomized stopping strategy. The two Skorokhod-type conditions involving (Y−ξ) and A are proposed as optimality conditions for randomized stopping, and are claimed to be new even in the classical single-agent setting. The main results are: (i) existence of solutions to the MKV-RFBSDE system by a Kakutani–Fan–Glicksberg fixed-point argument (Theorem 2.14), under Assumptions 1–3; (ii) uniqueness under a Lasry–Lions monotonicity assumption (Theorem 2.16); (iii) an equivalence theorem between solutions of the system and OS-MFG equilibria in randomized strategies (Theorem 5.2); (iv) an alternative existence proof of extremal solutions and learning algorithms via Tarski’s fixed-point theorem under Assumptions 5–6 (Theorems 4.9–4.10); (v) an approximate Nash equilibrium result for N-player games (Theorem 6.6); and (vi) a connection with the PDE/obstacle-problem approach of Bertucci (Theorem 7.1). The paper is carefully written and contains many nontrivial auxiliary results on continuity, compactness, and monotonicity of the maps involved.
Significance. If the results hold, this is a substantial contribution to the theory of OS-MFGs with randomized strategies. The KFG-based existence theorem and the equilibrium equivalence are coherent and are established without fitted parameters or self-citation loops; the two Skorokhod conditions are a genuine novelty. The approximate Nash theorem and the PDE bridge further increase the paper’s utility. However, the order-theoretic branch is conditional on Assumption 6(iii), a non-primitive monotonicity condition on the solution map Γ1, and the PDE section contains a subtle inconsistency at time zero. These issues do not invalidate the core KFG existence argument, but they do require substantial additional work or careful repositioning before publication.
major comments (3)
- [Section 4, Assumption 6(iii) and Lemma 4.7(ii)] The Tarski-based results (Theorems 4.9 and 4.10) depend on monotonicity of the best-response selection R. Lemma 4.7(ii) proves monotonicity of R(S) using Assumption 6(iii), which postulates that L ≤_V L' implies pointwise ordering of the gaps Y−ξ and Y'−ξ'. This is not a condition on the primitives b, σ, f, h, φ; it is a structural assumption on the solution map Γ1, and it is essentially the kind of final monotonicity that the Tarski construction is meant to produce. Remark 4.1 only gives a scalar example, and Assumption 8 provides sufficient conditions for Assumption 7 (τmin = τmax), not for Assumption 6(iii). If Assumption 6(iii) fails, R need not be monotone, Tarski’s theorem cannot be applied, and the extremal-solution and learning-algorithm results collapse. The paper should either prove Assumption 6(iii) from primitive conditions in a meaningful class, or explicitly reposition the
- [Section 7, definition of m_t and Theorem 7.1] The measure flow is defined as m_t(B) = E[1_B(X_t)L_t] for t ∈ (0,T], but m_0 is set to μ0 = Law(X0). Since V only requires L_{0−}=1 and does not require L_0=1, a solution may have L_0 < 1, corresponding to stopping mass at time 0. In that case the flow m_t just after zero has total mass E[L_0] < 1, while the PDE initial condition m_0 = μ0 has total mass 1. The Fokker–Planck inequality in Theorem 7.1(ii) is derived by integrating from 0− and using L_{0−}=1, so it does not correspond to the stated initial condition. A rigorous connection with [5] needs either m_0 = E[δ_{X_0}L_0] or an explicit jump/source term at t=0. As written, the claimed PDE bridge is not fully justified.
- [Theorem 4.10, Step 3] The passage to the limit in the Skorokhod integrals is not fully justified. In the last displayed estimate of Step 3, the term ∫(Ỹ_t − ξ̃_t)d(L^n_t − L̃_t) is said to be handled by the same technique as in Proposition 2.13. But Lemma 2.12 is only proved for Itô-type integrands with bounded coefficients, and Ỹ − ξ̃ is not shown to be such a process: ξ̃_t = h(t, X̃_t, E∫φ(t−s)dL̃_s) with h merely continuous. Additional regularity or a direct argument is needed to conclude that this term vanishes. Since Theorem 4.10 is the main convergence result for the learning schemes, this gap should be closed.
minor comments (6)
- [General] There are several typographical glitches in section headings, e.g., 'W ell-posedness' and 'T echnical results'.
- [Lemma 6.5] Lemma 6.5 is proved by saying it follows along the same lines as Lemma 6.3. While plausible, the deviation introduces an asymmetric term for player 1; a few more details would improve verifiability.
- [Theorem 2.16] The proof references Definition 5.1 and Theorem 5.2, which appear later in the paper. The reader is forced to jump ahead; consider stating the needed inequality from the equilibrium definition or moving the uniqueness result after Section 5.
- [Remark 4.3 / Assumption 8] Assumption 8 is introduced inside Remark 4.3, but Theorem 7.1 refers to 'Assumptions 8.a, 8.b and 8.c'. It would be cleaner to state Assumption 8 as a formal assumption outside a remark.
- [Section 7] The assumptions on h are inconsistent: Theorem 7.1 assumes h ∈ W^{1,2}([0,T]×R^d), while Assumption 8.b assumes h ∈ C^{1,2}. The Itô-formula arguments require C^{1,2}-type regularity; the Sobolev regularity should be either reconciled or justified.
- [Theorem 4.10, Step 1] The pointwise supremum L = sup_n L^n of càdlàg non-increasing processes is not automatically càdlàg. The right-continuous modification should be taken explicitly; the H^2 convergence statement needs a short justification.
Circularity Check
No significant circularity: central KFG existence and equilibrium equivalence are self-contained; Assumption 6(iii) is an explicit structural hypothesis, not a circular derivation.
full rationale
The paper's main derivation chain is not circular. Theorem 2.14 constructs the best-response correspondence Γ from the RBSDE solution map and proves non-emptiness, convexity and graph closedness using standard RBSDE stability results (El Karoui et al., external) and self-contained weak-compactness arguments; the Kakutani-Fan-Glicksberg fixed point then yields a solution of (2.2)-(2.5). No fitted parameter is renamed as a prediction, and no load-bearing self-citation is used: the many self-citations occur in the literature review or as technical pointers, not as the justification of the existence or equivalence theorems. Theorem 5.2 reduces to Theorem 3.3, which proves that the two Skorokhod conditions are equivalent to optimality in the randomized-stopping problem; this is a genuine equivalence proved from the RBSDE representation, not an identity by construction. The uniqueness result (Theorem 2.16) invokes Theorem 5.2 as a forward reference, but Theorem 5.2 itself does not depend on Theorem 2.16, so there is no circular dependency. The Tarski branch relies on Assumption 6(iii), which postulates the monotonicity of the gap Y−ξ in L; this is an explicit, non-primitive structural hypothesis rather than a derived conclusion, and the paper does not present it as a prediction from the model. That is a fragility or correctness risk, not circularity under the standards here. Overall the central existence and equivalence results stand on independent arguments, and the order-theoretic results are conditional on a clearly stated assumption.
Assumptions & free parameters
assumptions (9)
- domain assumption Assumption 1: global Lipschitz/linear growth of b, σ, ¯b, ¯σ
- domain assumption Assumption 2: Lipschitz/polynomial growth of f, h, ¯f; h continuous; ϕ ∈ C^1
- domain assumption Assumption 3: C^{1,2} regularity and polynomial growth of interaction functions ¯b, ¯σ, ¯f
- ad hoc to paper Assumption 4(ii): Lasry-Lions-type monotonicity with equality-iff-L=L' condition
- domain assumption Assumption 6(i)-(ii): non-negativity and componentwise monotonicity of ¯b, ¯f, b, f, h, and ϕ'
- ad hoc to paper Assumption 6(iii): monotonicity of the gap Y-ξ with respect to L
- ad hoc to paper Assumption 7: τmin = τmax for each L
- domain assumption Assumption 9: local Lipschitz growth of f and h for the N-player approximation
- standard math Standard results: RBSDE existence/uniqueness, KFG fixed-point theorem, Tarski fixed-point theorem, Helly/Arzelà-Ascoli
Cite this review
Pith. "Pith review of A new probabilistic approach for mean field games of optimal stopping." pith.science (2026). https://pith.science/paper/KH5V7SUA
@misc{pith2026260721062,
author = {Pith},
title = {Pith review of: A new probabilistic approach for mean field games of optimal stopping},
year = {2026},
howpublished = {\url{https://pith.science/paper/KH5V7SUA}},
note = {Machine review of arXiv:2607.21062}
}
abstract
We propose a novel probabilistic formulation for optimal stopping mean field games (OS-MFGs) with randomized strategies. We characterize mean field equilibria through a new class of coupled forward-backward systems, termed coupled reflected forward-backward McKean--Vlasov stochastic differential equations (MKV-RFBSDEs). An equilibrium is represented by a quintuple $(X,Y,Z,A,L)$, where $L$ is an adapted, $[0,1]$-valued, non-increasing c\`adl\`ag process representing the randomized stopping strategy. The optimality of randomized stopping strategies is characterized through two novel Skorokhod-type conditions involving $L$. This characterization is new even for classical optimal stopping problems without mean field interactions. We rigorously prove an equivalence between solutions of the MKV-RFBSDE system and OS-MFG equilibria in randomized strategies. We establish the existence of equilibria by applying the Kakutani--Fan--Glicksberg fixed-point theorem to a set-valued best-response correspondence, relying on new stability, compactness, and continuity results for the coupled MKV-RFBSDE system. We also prove uniqueness under suitable conditions. Under alternative monotonicity assumptions, we develop a new order-theoretic approach based on Tarski's fixed-point theorem, yielding the existence of extremal equilibria and constructive schemes for the minimal and maximal solutions. We further show that a mean field equilibrium induces an approximate Nash equilibrium for the associated $N$-player stopping game. Finally, we connect our probabilistic formulation with the analytical approach characterized by a coupled system of constrained partial differential equations.
Reference graph
Works this paper leans on
-
[5]
Bertucci
C. Bertucci. Optimal stopping in mean field games, an obstacle problem approach.Journal de Math´ ematiques Pures et Appliqu´ ees, 120:165–194, 2018
2018
-
[1]
Acciaio, J
B. Acciaio, J. Backhoff-Veraguas, and R. Carmona. Extended mean field control problems: stochastic maximum principle and transport perspective.SIAM journal on Control and Opti- mization, 57(6):3666–3693, 2019
2019
-
[2]
C. D. Aliprantis and K. C. Border.Infinite Dimensional Analysis: A Hitchhiker’s Guide. Springer, Berlin, Heidelberg, 3 edition, 2006
2006
-
[3]
M. Basei and H. Pham. Linear-quadratic McKean-Vlasov stochastic control problems with random coefficients on finite and infinite horizon, and applications.Preprint arXiv:1711.09390, 2017
arXiv 2017
-
[4]
Bensoussan, J
A. Bensoussan, J. Frehse, and P. Yam.Mean field games and mean field type control theory, volume 101. Springer, 2013
2013
-
[6]
J.-M. Bismut. Temps d’arrˆ et optimal, quasi-temps d’arrˆ et et retournement du temps.The Annals of Probability, pages 933–964, 1979
1979
-
[7]
Bouveret, R
G. Bouveret, R. Dumitrescu, and P. Tankov. Mean-field games of optimal stopping: a relaxed solution approach.SIAM Journal on Control and Optimization, 58(4):1795–1821, 2020
2020
-
[8]
Br´ ezis.Functional analysis, Sobolev spaces and partial differential equations, volume 2
H. Br´ ezis.Functional analysis, Sobolev spaces and partial differential equations, volume 2. Springer, 2011
2011
Show all 54 references
-
[9]
Cardaliaguet
P. Cardaliaguet. Notes on mean field games. Technical report, 2010
2010
-
[10]
Cardaliaguet, J
P. Cardaliaguet, J. Jackson, and P. E. Souganidis. Mean field control with stopping.Preprint arXiv:2603.21204, 2026
2026
-
[11]
Carmona and F
R. Carmona and F. Delarue. Probabilistic analysis of mean-field games.SIAM Journal on Control and Optimization, 51(4):2705–2734, 2013
2013
-
[12]
Carmona and F
R. Carmona and F. Delarue.Probabilistic theory of mean field games with applications. I, volume 83. Springer, Cham, 2018
2018
-
[13]
Carmona and F
R. Carmona and F. Delarue.Probabilistic theory of mean field games with applications. II, volume 84. Springer, Cham, 2018
2018
-
[14]
Carmona, F
R. Carmona, F. Delarue, and A. Lachapelle. Control of mckean–vlasov dynamics versus mean field games.Mathematics and Financial Economics, 7(2):131–166, 2013
2013
-
[15]
Carmona, F
R. Carmona, F. Delarue, and D. Lacker. Mean field games with common noise. 2016
2016
-
[16]
Carmona, F
R. Carmona, F. Delarue, and D. Lacker. Mean field games of timing and models for bank runs. Applied Mathematics & Optimization, 76(1):217–260, 2017
2017
-
[17]
Cosso and L
A. Cosso and L. Perelli. Mean field optimal stopping with uncontrolled state.Preprint arXiv:2503.04269, 2025
2025 arXiv
-
[18]
B. A. Davey and H. A. Priestley.Introduction to lattices and order. Cambridge university press, 2002. 39
2002
-
[19]
Dianetti, R
J. Dianetti, R. Dumitrescu, G. Ferrari, and R. Xu. Entropy regularization in mean-field games of optimal stopping.Preprint arXiv:2509.18821, 2025
2025
-
[20]
Dianetti, G
J. Dianetti, G. Ferrari, M. Fischer, and M. Nendel. A unifying framework for submodular mean field games.Mathematics of Operations Research, 48(3):1679–1710, 2023
2023
-
[21]
Djehiche and R
B. Djehiche and R. Dumitrescu. Zero-sum mean-field dynkin games: characterization and convergence.Mathematics of Operations Research, 51(2):1385–1412, 2026
2026
-
[22]
Djehiche, R
B. Djehiche, R. Elie, and S. Hamad` ene. Mean-field reflected backward stochastic differential equations.Preprint arXiv:1911.06079, 2019
1911 arXiv
-
[23]
A propagation of chaos result for weakly interacting nonlinear snell envelopes.Stochastic Processes and their Applications, 188:104669, 2025
Boualem Djehiche, Roxana Dumitrescu, and Jia Zeng. A propagation of chaos result for weakly interacting nonlinear snell envelopes.Stochastic Processes and their Applications, 188:104669, 2025
2025
-
[24]
Dumitrescu, M
R. Dumitrescu, M. Leutscher, and P. Tankov. Control and optimal stopping mean field games: a linear programming approach.Electronic Journal of Probability, 26:1–49, 2021
2021
-
[25]
Dumitrescu, M
R. Dumitrescu, M. Leutscher, and P. Tankov. Linear programming fictitious play algorithm for mean field games with optimal stopping and absorption.ESAIM: Mathematical Modelling and Numerical Analysis, 57(2):953–990, 2023
2023
-
[26]
B. Dupire. Functional Itˆ o calculus.Quant. Finance, 19(5):721–729, 2019
2019
-
[27]
El Karoui
N. El Karoui. Les aspects probabilistes du contrˆ ole stochastique. In ´Ecole d’´ et´ e de Probabilit´ es de Saint-Flour IX-1979, pages 73–238. Springer, 2006
1979
-
[28]
El Karoui, C
N. El Karoui, C. Kapoudjian, E. Pardoux, S. Peng, and M. C. Quenez. Reflected solutions of backward SDE’s, and related obstacle problems for PDE’s.Ann. Probab., 25(2):702–737, 1997
1997
-
[29]
El Karoui, J.-P
N. El Karoui, J.-P. Lepeltier, and A. Millet. A probabilistic approach to the reduite in optimal stopping.Probab. Math. Statist, 13(1):97–121, 1992
1992
-
[30]
S. N. Ethier and T. G. Kurtz.Markov processes. John Wiley & Sons, Inc., New York, 1986
1986
-
[31]
Ferrari and A
G. Ferrari and A. Pajola. Existence of strong randomized equilibria in mean-field games of optimal stopping with common noise.Preprint arXiv:2507.19123, 2025
2025 arXiv
-
[32]
P. J. Graber. Linear quadratic mean field type control and mean field games with common noise, with application to production of an exhaustible resource.Applied Mathematics & Optimization, 74(3):459–486, 2016
2016
-
[33]
Continuous-time mean field games: a primal-dual characterization.arXiv preprint arXiv:2503.01042, 2025
Xin Guo, Anran Hu, Jiacheng Zhang, and Yufei Zhang. Continuous-time mean field games: a primal-dual characterization.arXiv preprint arXiv:2503.01042, 2025
2025 arXiv
-
[34]
Mf-omo: An optimization formulation of mean-field games.SIAM Journal on Control and Optimization, 62(1):243–270, 2024
Xin Guo, Anran Hu, and Junzi Zhang. Mf-omo: An optimization formulation of mean-field games.SIAM Journal on Control and Optimization, 62(1):243–270, 2024
2024
-
[35]
X. He, X. Tan, and J. Zou. A mean-field version of bank–el karoui’s representation of stochastic processes.The Annals of Applied Probability, 35(5):3334–3377, 2025
2025
-
[36]
Huang, R
M. Huang, R. P. Malham´ e, and P. E. Caines. Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle.Communica- tions in Information & Systems, 6(3):221–252, 2006
2006
-
[37]
S. D. Jacka. Local times, optimal stopping and semimartingales.Ann. Probab., 21(1):329–339, 1993
1993
-
[38]
Y. M. Kabanov. Hedging and liquidation under transaction costs in currency markets.Finance and Stochastics, 3(2):237–248, 1999
1999
-
[39]
Kallenberg.Foundations of modern probability
O. Kallenberg.Foundations of modern probability. Springer, 3 edition, 2021
2021
-
[40]
Karatzas and S
I. Karatzas and S. Shreve.Brownian motion and stochastic calculus. springer, 2014
2014
-
[41]
D. Lacker. A general characterization of the mean field limit for stochastic differential games. Probability Theory and Related Fields, 165(3):581–648, 2016
2016
-
[42]
Mean field games via controlled martingale problems: existence of markovian equilibria.Stochastic Processes and their Applications, 125(7):2856–2894, 2015
Daniel Lacker. Mean field games via controlled martingale problems: existence of markovian equilibria.Stochastic Processes and their Applications, 125(7):2856–2894, 2015
2015
-
[43]
Lasry and P.-L
J.-M. Lasry and P.-L. Lions. Mean field games.Japanese Journal of Mathematics, 2(1):229–260, 2007
2007
-
[44]
P.-L. Lions. Th´ eorie des jeux de champ moyen et applications.Cours du College de France. http://www. college-de-france. fr/default/EN/all/equ der/audio video. jsp, 2007
2007
-
[45]
H. P. McKean. A class of markov processes associated with nonlinear parabolic equations. Proceedings of the National Academy of Sciences, 56(6):1907–1911, 1966. 40
1907
-
[46]
P.-A. Meyer. Convergence faible et compacit´ e des temps d’arrˆ et d’apres Baxter et Chacon. In S´ eminaire de Probabilit´ es XII: Universit´ e de Strasbourg 1976/77, pages 411–423. Springer, 2006
1976
-
[47]
Nualart.The Malliavin calculus and related topics
D. Nualart.The Malliavin calculus and related topics. Springer, 2006
2006
-
[48]
M. Nutz. A mean field game of optimal stopping.SIAM Journal on Control and Optimization, 56(2):1206–1221, 2018
2018
-
[49]
Peng and M
S. Peng and M. Xu. The smallestg-supermartingale and reflected bsde with single and double l2 obstacles. InAnnales de l’IHP Probabilit´ es et statistiques, volume 41, pages 605–630, 2005
2005
-
[50]
Possama ¨ ı and M
D. Possama ¨ ı and M. Talbi. Mean-field games of optimal stopping: master equation and weak equilibria.Applied Mathematics & Optimization, 92(3):1–32, 2025
2025
-
[51]
Revuz and M
D. Revuz and M. Yor.Continuous martingales and Brownian motion. Springer Science & Business Media, 2013
2013
-
[52]
Talbi, N
M. Talbi, N. Touzi, and J. Zhang. Dynamic programming equation for the mean field optimal stopping problem.SIAM Journal on Control and Optimization, 61(4):2140–2164, 2023
2023
-
[53]
A. Tarski. A lattice-theoretical fixpoint theorem and its applications.Pacific Journal of Math- ematics, 5(2):285–309, 1955
1955
-
[54]
Touzi and N
N. Touzi and N. Vieille. Continuous-time dynkin games with mixed strategies.SIAM Journal on Control and Optimization, 41(4):1073–1088, 2002. 41
2002
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.