{"id":"88b3e952-6a81-4820-8973-01e706356db5","arxiv_id":"2607.21062","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Randomized mean-field equilibria of optimal-stopping games are characterized by a coupled reflected McKean–Vlasov forward-backward SDE system whose survival process L is an endogenous part of the solution.","lead":"This paper introduces a new probabilistic formulation for mean-field games of optimal stopping, modeling randomized stopping strategies through a coupled system of reflected McKean–Vlasov forward-backward stochastic differential equations. It proves existence, uniqueness, N-player approximation, and a bridge to the PDE/obstacle approach.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4's extremal/learning results hinge on Assumption 6(iii), a non-primitive monotonicity condition on the solution map Γ1; no derivation from primitives is supplied.","rationale":"The reader's weakest assumption—Assumption 6(iii)—is indeed the point where the paper's order-theoretic results rest on a hypothesis that is close to the desired conclusion. I independently checked the structure of the Tarski argument: without Assumption 6(iii), Lemma 4.7(ii) cannot establish the monotonicity of R, and therefore Theorem 4.9, the learning algorithms, and Proposition 5.6 lose their foundation. This does not undermine Theorem 2.14 or Theorem 5.2, which rely on the KFG fixed-point argument and appear internally sound. The paper deserves credit for the genuinely new probabilistic formulation, the Skorokhod-type optimality conditions, and the rigorous KFG existence proof, but the advertised extremal/learning results should either be proved under primitive sufficient conditions or explicitly demarcated as conditional on a structural monotonicity hypothesis. The reader's CONDITIONAL verdict is appropriate, and my stress-test does not alter it.","tokens_in":52899,"tokens_out":35844,"duration_ms":378461,"concrete_test":"In the one-dimensional Markovian setting of Remark 4.1, take h(t,x,w)=x+ψ(w) for a strictly increasing, nonlinear ψ (e.g., ψ(w)=sqrt(w+1)), f(t,x,m)=λφ(m) with increasing φ and λ>0, and choose two ordered strategies L≤L′ (for instance, pure stopping at τ≤τ′). Compute the RBSDE gap Y−ξ explicitly from the associated obstacle problem/PDE for both strategies and check whether the inequality in Assumption 6(iii) holds for all such L≤L′. A single failure would show Assumption 6(iii) is not implied by Assumptions 1,2,5,6(i,ii), and Section 4 needs either a proof from primitives or an explicit additional standing assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The KFG-based existence (Theorem 2.14) and the equilibrium equivalence (Theorem 5.2) are, as far as I can see, internally coherent and do not require Assumption 6(iii). The load-bearing weakness is in the order-theoretic branch. The Tarski argument in Theorems 4.9/4.10 requires the selection R(L)=ess-inf Γ(L) to be monotone in L. Lemma 4.8(ii) obtains this from Lemma 4.7(ii), and the final step of Lemma 4.7(ii) uses Assumption 6(iii) to infer Y_t=ξ_t from Y'_t=ξ'_t, Y_t≤Y'_t, and the assumed ordering of the gaps. Assumption 6(iii) is not a condition on the primitives b,σ,f,h,ϕ: it asserts that the solution map L ↦ (Y−ξ) is pointwise order-preserving, which is essentially the kind of monotonicity the Tarski construction is intended to produce. The only supporting evidence is the scalar example in Remark 4.1 plus separate sufficient conditions for Assumption 7 (τ_min=τ_max), not for Assumption 6(iii). If Assumption 6(iii) fails, R need not be monotone, Tarski's theorem cannot be applied, and the minimal/maximal equilibrium and learning-algorithm results collapse. This supports the CONDITIONAL verdict: the central existence results stand, but Section 4 is conditional on an unverified structural assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a probabilistic formulation of mean-field games of optimal stopping (OS-MFGs) with randomized stopping strategies, based on a coupled reflected forward-backward McKean–Vlasov system (2.2)-(2.5). The equilibrium object is a quintuple (X, Y, Z, A, L), where L is a [0,1]-valued, non-increasing càdlàg process representing the survival/randomized stopping strategy. The two Skorokhod-type conditions involving (Y−ξ) and A are proposed as optimality conditions for randomized stopping, and are claimed to be new even in the classical single-agent setting. The main results are: (i) existence of solutions to the MKV-RFBSDE system by a Kakutani–Fan–Glicksberg fixed-point argument (Theorem 2.14), under Assumptions 1–3; (ii) uniqueness under a Lasry–Lions monotonicity assumption (Theorem 2.16); (iii) an equivalence theorem between solutions of the system and OS-MFG equilibria in randomized strategies (Theorem 5.2); (iv) an alternative existence proof of extremal solutions and learning algorithms via Tarski’s fixed-point theorem under Assumptions 5–6 (Theorems 4.9–4.10); (v) an approximate Nash equilibrium result for N-player games (Theorem 6.6); and (vi) a connection with the PDE/obstacle-problem approach of Bertucci (Theorem 7.1). The paper is carefully written and contains many nontrivial auxiliary results on continuity, compactness, and monotonicity of the maps involved.","tokens_in":53254,"tokens_out":22962,"duration_ms":223226,"significance":"If the results hold, this is a substantial contribution to the theory of OS-MFGs with randomized strategies. The KFG-based existence theorem and the equilibrium equivalence are coherent and are established without fitted parameters or self-citation loops; the two Skorokhod conditions are a genuine novelty. The approximate Nash theorem and the PDE bridge further increase the paper’s utility. However, the order-theoretic branch is conditional on Assumption 6(iii), a non-primitive monotonicity condition on the solution map Γ1, and the PDE section contains a subtle inconsistency at time zero. These issues do not invalidate the core KFG existence argument, but they do require substantial additional work or careful repositioning before publication.","major_comments":[{"comment":"The Tarski-based results (Theorems 4.9 and 4.10) depend on monotonicity of the best-response selection R. Lemma 4.7(ii) proves monotonicity of R(S) using Assumption 6(iii), which postulates that L ≤_V L' implies pointwise ordering of the gaps Y−ξ and Y'−ξ'. This is not a condition on the primitives b, σ, f, h, φ; it is a structural assumption on the solution map Γ1, and it is essentially the kind of final monotonicity that the Tarski construction is meant to produce. Remark 4.1 only gives a scalar example, and Assumption 8 provides sufficient conditions for Assumption 7 (τmin = τmax), not for Assumption 6(iii). If Assumption 6(iii) fails, R need not be monotone, Tarski’s theorem cannot be applied, and the extremal-solution and learning-algorithm results collapse. The paper should either prove Assumption 6(iii) from primitive conditions in a meaningful class, or explicitly reposition the","section":"Section 4, Assumption 6(iii) and Lemma 4.7(ii)"},{"comment":"The measure flow is defined as m_t(B) = E[1_B(X_t)L_t] for t ∈ (0,T], but m_0 is set to μ0 = Law(X0). Since V only requires L_{0−}=1 and does not require L_0=1, a solution may have L_0 < 1, corresponding to stopping mass at time 0. In that case the flow m_t just after zero has total mass E[L_0] < 1, while the PDE initial condition m_0 = μ0 has total mass 1. The Fokker–Planck inequality in Theorem 7.1(ii) is derived by integrating from 0− and using L_{0−}=1, so it does not correspond to the stated initial condition. A rigorous connection with [5] needs either m_0 = E[δ_{X_0}L_0] or an explicit jump/source term at t=0. As written, the claimed PDE bridge is not fully justified.","section":"Section 7, definition of m_t and Theorem 7.1"},{"comment":"The passage to the limit in the Skorokhod integrals is not fully justified. In the last displayed estimate of Step 3, the term ∫(Ỹ_t − ξ̃_t)d(L^n_t − L̃_t) is said to be handled by the same technique as in Proposition 2.13. But Lemma 2.12 is only proved for Itô-type integrands with bounded coefficients, and Ỹ − ξ̃ is not shown to be such a process: ξ̃_t = h(t, X̃_t, E∫φ(t−s)dL̃_s) with h merely continuous. Additional regularity or a direct argument is needed to conclude that this term vanishes. Since Theorem 4.10 is the main convergence result for the learning schemes, this gap should be closed.","section":"Theorem 4.10, Step 3"}],"minor_comments":[{"comment":"There are several typographical glitches in section headings, e.g., 'W ell-posedness' and 'T echnical results'.","section":"General"},{"comment":"Lemma 6.5 is proved by saying it follows along the same lines as Lemma 6.3. While plausible, the deviation introduces an asymmetric term for player 1; a few more details would improve verifiability.","section":"Lemma 6.5"},{"comment":"The proof references Definition 5.1 and Theorem 5.2, which appear later in the paper. The reader is forced to jump ahead; consider stating the needed inequality from the equilibrium definition or moving the uniqueness result after Section 5.","section":"Theorem 2.16"},{"comment":"Assumption 8 is introduced inside Remark 4.3, but Theorem 7.1 refers to 'Assumptions 8.a, 8.b and 8.c'. It would be cleaner to state Assumption 8 as a formal assumption outside a remark.","section":"Remark 4.3 / Assumption 8"},{"comment":"The assumptions on h are inconsistent: Theorem 7.1 assumes h ∈ W^{1,2}([0,T]×R^d), while Assumption 8.b assumes h ∈ C^{1,2}. The Itô-formula arguments require C^{1,2}-type regularity; the Sobolev regularity should be either reconciled or justified.","section":"Section 7"},{"comment":"The pointwise supremum L = sup_n L^n of càdlàg non-increasing processes is not automatically càdlàg. The right-continuous modification should be taken explicitly; the H^2 convergence statement needs a short justification.","section":"Theorem 4.10, Step 1"}],"recommendation":"major_revision","confidential_remarks":"The core KFG existence result and the equivalence theorem appear sound and could be publishable. My main concern is Section 4: the Tarski branch is conditional on Assumption 6(iii), which is essentially the desired monotonicity of the best-response map, and no primitive sufficient conditions are supplied. The PDE section also has a time-zero mass inconsistency. I recommend major revision, not rejection, because the central probabilistic existence/equivalence claims are defensible and the issues, while load-bearing for the advertised Section 4 and Section 7 results, are local enough to be addressed in a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers a genuinely new probabilistic characterization of randomized OS-MFG equilibria as fixed points of a coupled reflected MKV-FBSDE system. The two Skorokhod-type optimality conditions are new even in the classical one-player setting, and the equivalence theorem (5.2) is the right completeness statement. The existence proof via Kakutani–Fan–Glicksberg (Theorem 2.14) is coherent; the stability and compactness machinery in Section 2 is the heart of the paper and it appears to work. The approximate-Nash result in Section 6 is also a credible bonus.\n\nThe soft spot is the order-theoretic branch in Section 4. Assumption 6(iii) is not a condition on primitives: it postulates that the solution map L ↦ (Y−ξ) is pointwise order-preserving. That is very close to the monotonicity that Tarski is meant to deliver. The paper gives only a scalar example (Remark 4.1) and a list of sufficient conditions for Assumption 7 (τmin=τmax), not for 6(iii). If 6(iii) fails, the minimal/maximal and learning-algorithm results collapse. This is a real caveat, not a manufactured one; the stress-test note is on target. Also, Assumption 7 is restrictive, and the PDE connection in Section 7 is formal and needs extra smoothness/structural assumptions. None of this threatens the central existence/equivalence claims, which do not use 6(iii), but the advertised \"extremal and learning\" results are conditional.\n\nWho benefits: anyone working on probabilistic methods for MFGs with stopping, and people using RBSDEs for randomized stopping. The paper deserves a serious referee if it hasn't been sent out already; it is a clear advance over the existing obstacle-PDE and LP approaches, and the KFG track is likely correct. The authors should either prove 6(iii) from primitives, replace it, or clearly mark Section 4 as conditional.\n\nRecommendation: send it to a good referee, and require the authors to address the status of 6(iii) before publication.","headline":"A substantial, mostly rigorous new FBSDE characterization of randomized-optimal-stopping MFGs; the KFG-based existence is solid, but the Tarski extremal/learning track leans on an unverified structural monotonicity assumption.","tokens_in":53755,"tokens_out":1517,"would_cite":true,"duration_ms":18501,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A16","60G40","60H10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Mean-field optimal stopping equilibria in randomized strategies coincide exactly with solutions of a new coupled reflected forward-backward McKean–Vlasov system.","keywords":["mean field games","optimal stopping","randomized stopping strategies","McKean–Vlasov equations","reflected backward SDEs","survival processes","Nash equilibrium","obstacle problems"],"falsifier":"In a one-dimensional Markovian example with Brownian state, f=0, and h(t,x)=x, compute the reflected BSDE candidate and the survival process L; the equivalence predicts that the support of −dL is contained in {Y=ξ}∩{A=0}, so any positive stopping mass outside that set would refute the characterization. Alternatively, produce L≤L′ satisfying the standing assumptions for which the gap Y−h is not ordered, which would falsify the monotonicity assumption behind the Tarski route.","tokens_in":52743,"feed_emoji":"🎯","tokens_out":8606,"duration_ms":85584,"temperature":0.7,"pith_summary":"This paper establishes a complete probabilistic characterization of mean field games of optimal stopping when players may randomize their stopping times. It proves that an equilibrium is the same thing as a quintuple (X,Y,Z,A,L) solving a coupled system in which the state process, a reflected backward value process, and a survival process L are determined together. Two new optimality conditions—the value must meet the stopping payoff only where L can place mass, and no stopping mass may appear once reflection has begun—turn strategic optimality into equations. Existence is obtained under two complementary sets of assumptions, and the system yields approximate Nash equilibria for large finite-player games and connects to an obstacle-problem PDE formulation.","feed_headline":"One system solves mean-field stopping games with random strategies","feed_subtitle":"Solving for the survival process directly; optimality is two contact conditions; existence and N-player results.","key_machinery":"The load-bearing object is the coupled MKV-RFBSDE system: a forward-backward system in which the randomized stopping strategy L is solved for as part of the unknown, rather than recovered from an external flow of measures. The two integral conditions on L—that the measure −dL be supported on the contact set of the value with the obstacle and on the flat set of the reflection process—are the mechanism that makes a candidate survival process an optimal response.","core_discovery":"The central discovery is that randomized-strategy equilibria of optimal-stopping mean field games are not a separate fixed-point object: they are exactly the L-component of a solution to a coupled reflected forward-backward McKean–Vlasov system. In the system, L is an adapted, non-increasing survival process taking values in [0,1], the state X evolves with coefficients averaged against the surviving population, and the reflected backward component (Y,Z,A) solves an RBSDE with obstacle built from L. Optimality is encoded by two contact conditions: the measure −dL can charge only times where Y equals the obstacle, and only times where the reflection process A has not yet increased. The paper p","pith_inferences":["Because the optimality conditions are stated purely through contact sets of Y and A, the same two-condition test could serve as a Snell-envelope-style criterion for randomized stopping in single-agent problems.","The fixed point lives directly on survival processes, which suggests a natural numerical loop—solve the RBSDE for a given L, update L by the contact sets, iterate—that the order-theoretic route shows converges to extremal equilibria in monotone settings.","The non-Markovian extension noted in the paper indicates the two contact conditions are filtration-relative, so the characterization may persist with common noise or partial information.","The finite-player approximation is stated for i.i.d. copies of the equilibrium strategy; a natural continuation is to quantify deviations under dependent initial data or common noise."],"forward_implications":["Randomized mean-field equilibria exist under the paper's assumptions, so pure-strategy non-existence is circumvented by allowing survival processes.","Any solution of the coupled system is automatically an equilibrium, and any equilibrium produces a solution of the system; the game and the system are the same problem.","Under monotonicity assumptions there are minimal and maximal equilibria, ordered by survival probability, with iterative learning schemes that converge to them.","A mean-field equilibrium induces an ε-Nash equilibrium for the N-player stopping game, with the approximation error vanishing as N grows.","The probabilistic system is equivalent to a constrained obstacle-problem PDE system, providing an analytic route to the same equilibria."],"fun_headline_variants":["Random stopping in mean-field games equals reflected McKean-Vlasov system","Randomized-stopping equilibria are solutions of coupled reflected FBSDEs","Mean-field stopping games solved via one coupled MKV-RFBSDE system","Survival process L: the key to optimal stopping mean-field games","From N-player games to MKV-RFBSDE: a unified solution for stopping"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The most fragile premise is the order-preservation assumption that whenever one survival process stays below another, the gap between the continuation value and the stopping payoff is pointwise ordered in the same direction; the Tarski-based existence of extremal equilibria and the learning algorithms collapse if this monotonicity fails.","fun_headline_variants_meta":{"raw":{"variants":["Random stopping in mean-field games equals reflected McKean-Vlasov system","Randomized-stopping equilibria are solutions of coupled reflected FBSDEs","Mean-field stopping games solved via one coupled MKV-RFBSDE system","Survival process L: the key to optimal stopping mean-field games","From N-player games to MKV-RFBSDE: a unified solution for stopping"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1240,"prompt_tokens":820,"completion_tokens":420,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":319}},"tokens_in":564,"tokens_out":420,"duration_ms":4847,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T08:35:04.870351+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In a one-dimensional Markovian example with Brownian state, f=0, and h(t,x)=x, compute the reflected BSDE candidate and the survival process L; the equivalence predicts that the support of −dL is contained in {Y=ξ}∩{A=0}, so any positive stopping mass outside that set would refute the characterization. Alternatively, produce L≤L′ satisfying the standing assumptions for which the gap Y−h is not ordered, which would falsify the monotonicity assumption behind the Tarski route.","supporting_citations":[],"review_version":1}