{"id":"0faa402d-8bbe-496b-b6dd-7a6b294034a9","arxiv_id":"2507.00853","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For target-based quantile-competition mean-field games, the equilibrium is characterized by decoupled ordinary differential equations with an explicit epsilon-Nash error of order 1/sqrt(N).","lead":"A new class of mean-field games where agents are ranked by whether their final value clears the population's alpha-quantile is introduced, and an exact solution is found for the 'target' version. This makes rank-based competition in large groups, such as venture capital funding rounds, mathematically tractable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.3's constructed N-player best response is anticipating and not admissible, so the epsilon-Nash proof is incomplete as written; an adaptedness fix is needed.","rationale":"The reader's weakest-assumption analysis isolates exactly the load-bearing gap: the N-player best response used in Step 1 of Theorem 3.3 is not adapted to the agent's filtration, so the epsilon-Nash conclusion is not established by the given proof. I agree with this assessment and find no independent, more serious flaw in the paper's central argument. The decoupled FBODE characterization in Theorem 3.1, the uniqueness result in Proposition 3.2, and the numerical exploration of the threshold-based formulation are all consistent with the claims as stated; the threshold-based consistency and convergence are explicitly left open in Section 4, so they are not a hidden defect. The concern is localized to Theorem 3.3, but it is load-bearing because the epsilon-Nash property is a headline contribution. The fix is plausibly straightforward: replace the anticipating backward-ODE control with the genuinely adapted conditional-expectation control and rework the bounds in Step 2, or impose an additional argument showing the anticipating control is epsilon-optimal. Because the paper otherwise appears sound and the gap is repairable, the appropriate disposition is the same conditional acceptance the reader recommended.","tokens_in":21219,"tokens_out":6436,"duration_ms":82736,"concrete_test":"Re-derive the finite-N best response from the stochastic maximum principle for a fixed admissible deviation. For the linear dynamics and terminal cost \\lambda/2(x_T^i - q_T^{[N]})^2, the adjoint BSDE is dy_t = z_t dw^i_t with y_T = \\lambda(x_T^i - q_T^{[N]}), yielding y_t = E[\\lambda(x_T^i - q_T^{[N]}) | \\mathcal{F}_t^{[N]}] and \\hat u_t^i = -(b/r)y_t. Then compare this adapted formula with (3.49)-(3.50) in a small numerical example, e.g. N=3 with the parameters in Table 1: if the two controls differ at any t<T, the ODE-based control is not the SMP best response and Step 1 of Theorem 3.3 fails. A sharper analytical check is to verify whether the process M_t = \\hat\\theta_t^{[N]} + \\lambda E[q_T^{[N]} | \\mathcal{F}_t^{[N]}] is a martingale; the backward ODE (3.50) generically gives it nonzero drift, confirming the anticipating flaw.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central theorem 3.3 rests on the claim that the deviating agent's best response in the N-player game is \\hat u^i_t = -(b/r)(\\eta_t \\hat x^i_t + \\hat\\theta^{\\alpha,[N]}_t), where \\hat\\theta solves the backward ODE (3.50) with random terminal condition \\hat\\theta_T = -\\lambda q^{\\alpha,[N]}_T. Because q^{\\alpha,[N]}_T is \\mathcal{F}^{[N]}_T-measurable but not \\mathcal{F}^{[N]}_t-measurable for t<T, the solution \\hat\\theta_t is an anticipating functional of the whole sample, and \\hat u^i is not \\mathcal{F}^{[N]}-adapted, hence not admissible under (2.5). The stochastic maximum principle invoked at (3.48) applies to admissible controls only, so identifying \\hat u^i as the infimum is unjustified. This is not cosmetic: the backward ODE has no martingale term, so it cannot represent the conditional expectation E[\\lambda(\\hat x_T^i - q_T^{[N]})|\\mathcal{F}_t^{[N]}] that the correct adjoint BSDE would produce. Since the epsilon-Nash bound (3.43) is obtained by comparing finite-N deviating costs through this \\hat u^i, the proof of Theorem 3.3 is incomplete as written. The result is plausibly repairable, e.g. by using the genuinely adapted adjoint y_t = E[\\lambda(\\hat x_T^i - q_T^{[N]})|\\mathcal{F}_t^{[N]}] and controlling the resulting martingale term, but the current argument does not establish the advertised epsilon-Nash property.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies a class of mean-field games in which agents' terminal costs depend on the sample α-quantile of the population's terminal states, motivated by ranking competitions. Two formulations are introduced: a target-based one, where agents aim to land exactly on the quantile, and a threshold-based one, where agents are penalized only for falling below the quantile. For the target-based formulation, the paper derives a fully decoupled forward-backward ODE system characterizing the equilibrium quantile process, proves existence and uniqueness of solutions, and claims an ε-Nash property for the finite-N game with an error of order sqrt(1/N) * sqrt(α(1−α)) divided by the terminal-state density at the quantile. For the threshold-based formulation, the paper gives a semi-explicit optimal control representation and solves the mean-field consistency condition numerically via a fixed-point iteration. The model is applied to early-stage venture capital, where a fund selects top start-ups, and a set of numerical experiments with sensitivity analysis is presented.","tokens_in":21505,"tokens_out":8582,"duration_ms":99184,"significance":"If the ε-Nash claim is correct, the paper makes a useful contribution: it provides one of the first analytically tractable mean-field game models with quantile-type interactions, and it does so without ad hoc fitted constants. The core equilibrium characterization in Theorem 3.1 and Proposition 3.2 is clean and self-contained: the Gaussian quantile identity (3.21) and the decoupling through η+π=0 are correct, and the resulting FBODE system is explicit. The threshold-based numerics and the venture-capital application are novel and clearly presented. However, the significance of the paper hinges on the ε-Nash property, and the proof of Theorem 3.3 has a serious adaptedness gap that is not cosmetic. The contribution remains potentially valuable, but the advertised finite-player approximation result is not established as written.","major_comments":[{"comment":"The control hat u^i_t = -(b/r)(eta_t hat x^i_t + hattheta^{alpha,[N]}_t) asserted to achieve the infimum in (3.48) is not admissible. By (3.50), hattheta^{alpha,[N]}_t solves a backward ODE with terminal condition hattheta^{alpha,[N]}_T = -lambda q^{alpha,[N]}_T, and for t<T the sample quantile q^{alpha,[N]}_T is mathcal{F}^{[N]}_T-measurable but generally not mathcal{F}^{[N]}_t-measurable, since it depends on future increments of all agents' Brownian motions, including agent i's own w^i. Consequently hattheta^{alpha,[N]}_t is an anticipating functional of the full sample, and hat u^i is not mathcal{F}^{[N]}-adapted; it therefore violates the admissibility condition U^{[N]} in (2.5). The stochastic maximum principle invoked at (3.48) applies only to admissible controls, so the identification of hat u^i with the infimum is unjustified. The correct adjoint for a deviating agent would be a BSDE with a martingale term representing the conditional expectation of lambda(hat x^i_T - q^{alpha,[N]}_T) given mathcal{F}^{[N]}_t; because q^{alpha,[N]}_T is a nonlinear functional of the whole population, that martingale term need not vanish and the linear feedback structure (3.49) is not the true best response. Since the chain of inequalities (3.52)–(3.68) compares arbitrary finite-N deviations through this hat u^i, the proof of the epsilon-Nash bound (3.43) is incomplete. The claim may be repairable with a genuinely adapted finite-N best-response construction, but that requires new arguments rather than a minor correction.","section":"Section 3.3, proof of Theorem 3.3, Eqs. (3.48)–(3.51)"},{"comment":"The bound in (3.69) uses E[(q^{alpha,[N]}_T - bar q^alpha_T)^2]^{1/2} and concludes in (3.72) a rate O(1/sqrt{N}) from the quantile central limit theorem (3.70). The CLT is a statement of convergence in distribution; it does not by itself give convergence of second moments, which is what (3.69) requires. Uniform integrability of N(q^{alpha,[N]}_T - bar q^alpha_T)^2 is plausible (e.g., from Gaussian tail bounds for order statistics), but it must be proved before the epsilon-Nash rate is justified. In addition, (3.72) writes mu(T,bar q^alpha_T) where (3.70) and (3.43) use p(T,bar q^alpha_T); the notation should be made consistent.","section":"Section 3.3, Step 3, Eqs. (3.69)–(3.72)"}],"minor_comments":[{"comment":"The sensitivity analysis is useful, but the paper does not report any error bars or confidence intervals for the finite-N simulations (e.g., in Figures 4, 5, and 7); a brief discussion of Monte Carlo variability would strengthen the comparison between the target-based and threshold-based formulations.","section":"Section 5, general"},{"comment":"References [22] and [43] appear to describe the same article (same title, same authors) with different publication statuses; one should be removed or cross-referenced to avoid duplication.","section":"Introduction, references"},{"comment":"The paper claims in the introduction that the epsilon-Nash proof 'is also applicable to the model in [43]' and that it relaxes a uniform bound assumption, but this claim is not revisited or substantiated anywhere in the text; either provide the argument or soften the claim.","section":"Introduction, contribution statement"},{"comment":"The proof of Proposition 3.2 asserts uniqueness of the solution to (3.3)-(3.7) but only spells out the argument for (3.3), relying on [45] for a Riccati-type equation that is, in fact, explicitly solvable; adding the closed form eta_t = 1/(1/lambda + (b^2/r)(T-t)) would make the existence and uniqueness claim immediately transparent.","section":"Proposition 3.2, proof"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of math.OC and the MFG literature, and the target-based equilibrium derivation is sound. The main obstacle is the proof of Theorem 3.3: the proposed finite-N best response is not adapted, so the advertised epsilon-Nash property is not established. I would encourage the editor to seek a revision where the authors either supply a genuinely adapted best-response argument or revise the claim to reflect what is actually proved. The issue is serious but appears local to Section 3.3, and I do not see grounds for rejection if the proof can be repaired."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nQuick read of arXiv:2507.00853. My take: it is a genuinely useful paper with a real gap in the headline epsilon-Nash proof, but the gap is localized and likely repairable.\n\nWhat is new: the target-based ranking quantilized MFG, where each agent pays a quadratic penalty for terminal state differing from the population alpha-quantile. The limiting equilibrium is characterized by a fully decoupled FBODE system (3.3)-(3.7), solved explicitly. The decoupling trick—using theta = pi qbar + phi and showing eta+pi=0—is clean and correct. The numerical experiments comparing target vs threshold formulations in a venture-capital staging setting are thoughtful, and the sensitivity analysis is a nice bonus. I agree the epsilon-Nash result is absent from prior quantilized MFG work.\n\nWhere it is soft: Theorem 3.3. The deviating agent's best response is written as (3.49) with hat-theta solving the backward ODE (3.50) whose terminal condition is -lambda q^{alpha,[N]}_T, the sample quantile. That terminal condition is not known at time t<T, so hat-theta_t is anticipating and the control is not F^{[N]}-adapted. It is not admissible under (2.5), so the stochastic maximum principle argument at (3.48) does not apply. The epsilon-Nash bound (3.43) depends on that identification. This matches the stress-test concern. The fix is plausible: replace the anticipating backward ODE with the genuinely adapted adjoint y_t = E[lambda(hat x_T^i - q_T^{[N]})|F_t], allow a martingale term, and bound it. But as written the proof does not go through. The threshold-based formulation is explicitly left open, which is fine and honestly stated.\n\nOther soft spots are minor: the CLT rate (3.43) requires the density at the quantile to be nonzero, which holds for the Gaussian limit, but the finite-N empirical quantile convergence is cited rather than proved from their setup. The application is a stylized model, not calibrated to real VC data, but it is clearly presented as an illustration.\n\nBottom line: the analytic target-based solution is new and correct, and the epsilon-Nash claim is plausible but needs a repaired adaptedness argument. A serious referee should engage with it; I would send it out with a request for a corrected proof of Theorem 3.3. I would cite the decoupling result and the application if I worked in this area.","headline":"A useful new quantilized MFG model with a clean analytic solution, but the epsilon-Nash proof has an adaptedness gap that needs repair.","tokens_in":22103,"tokens_out":2769,"would_cite":true,"duration_ms":32220,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A16","49N80"],"pacs":[],"model":"deepseek-v4-flash","headline":"An explicit equilibrium for target-based ranking quantilized mean-field games, with an epsilon-Nash guarantee for the finite-player game.","keywords":["mean-field games","quantilized mean-field games","alpha-quantiles","ranking games","epsilon-Nash equilibrium","forward-backward ordinary differential equations","early-stage venture investments","start-up competition"],"falsifier":"Check whether $\\hat u^i_t = -\\frac{b}{r}(\\eta_t \\hat x^i_t+\\hat\\theta^{\\alpha,[N]}_t)$ with $\\hat\\theta^{\\alpha,[N]}_T=-\\lambda q_T^{\\alpha,[N]}$ is adapted: for $N$ finite and $\\alpha\\in(0,1)$, $q_T^{\\alpha,[N]}$ is not $\\sigma(x^i_0,w^i_s: s\\le t)$-measurable for $t<T$, so the backward equation (3.50) is anticipating. A direct calculation of $J_i^{[N]}(\\hat u^i,\\cdot)$ compared with the true dynamic-programming optimum for a deviator would show whether the non-adapted control can be replaced by an admissible one that achieves the same cost; if not, Step 1 of the proof (relation (3.48)) has no justification and the claimed bound is not established by the given argument.","tokens_in":20950,"feed_emoji":"🏆","tokens_out":9401,"duration_ms":97963,"temperature":0.7,"pith_summary":"The paper studies mean-field games in which agents are ranked by their terminal state against the population's $\\alpha$-quantile, so the top $(1-\\alpha)\\%$ qualify, and it proposes two competing formulations. In the target-based version, each agent tries to land exactly on the quantile, and the paper proves that the equilibrium quantile path and the optimal effort strategy are characterized by a fully decoupled system of five forward-backward ordinary differential equations. It then shows that using these limiting strategies in the $N$-player game forms an $\\epsilon$-Nash equilibrium, with error of order $\\sqrt{1/N}\\,\\sqrt{\\alpha(1-\\alpha)}/p(T,\\bar q^\\alpha_T)$, where $p$ is the terminal-state density at the quantile. The threshold-based version, where agents only need to reach or exceed the quantile, is solved semi-explicitly and handled numerically. The motivating application is early-stage venture capital: a VC firm funds competing start-ups and selects those whose terminal market values reach the endogenously determined quantile threshold.","feed_headline":"Quantile ranking games get a near-Nash equilibrium","feed_subtitle":"Target-based mean-field solution is provably near-optimal for N players and maps to VC start-up contests.","key_machinery":"The load-bearing object is the sample $\\alpha$-quantile $q_T^{\\alpha,[N]}$ of the terminal states, defined by (2.3), and its deterministic limit $\\bar q^\\alpha_t$. In the limiting game the quantile of the Gaussian state law has the explicit form $\\bar q^\\alpha_t = \\mathbb{E}[x^\\star_t] + X^\\alpha\\sqrt{\\mathbb{V}[x^\\star_t]}$, with $X^\\alpha=Q(\\alpha,\\mathcal N(0,1))$. The machinery is a three-step loop: (i) the stochastic maximum principle for the quadratic terminal cost $(\\lambda/2)(x_T-q_T^\\alpha)^2$, producing Riccati equations for the adjoint coefficients; (ii) differentiation of the Gaussian quantile identity to obtain the quantile-path ODE; (iii) the identity $\\pi_t=-\\eta_t$, which follows from the Riccati terminal conditions and decouples the system into (3.3)--(3.7). The finite-$N$ error estimate is carried by a central-limit theorem for sample quantiles, which turns the terminal mismatch between $q_T^{\\alpha,[N]}$ and $\\bar q_T^\\alpha$ into the rate displayed in (3.43).","core_discovery":"The central discovery is that the target-based ranking model has a closed-form mean-field equilibrium, and that this equilibrium is nearly optimal for each agent in the finite population. Fixing the limiting quantile process, the representative agent's optimal control is $u^\\star_t = -\\frac{b}{r}(\\eta_t x^\\star_t + \\pi_t \\bar q^\\alpha_t + \\phi^\\alpha_t)$, where the coefficient processes solve the decoupled FBODEs (3.3)--(3.7). The argument first uses the stochastic maximum principle with the adjoint ansatz $y_t=\\eta_t x_t+\\theta^\\alpha_t$, then imposes the quantilized consistency condition $q^\\alpha_t = \\mathbb{E}[x^\\star_t]+X^\\alpha\\sqrt{\\mathbb{V}[x^\\star_t]}$, where $X^\\alpha$ is the standard-normal quantile. A key simplification is the identity $\\eta_t+\\pi_t=0$, which turns coupled equations into the explicit system. Theorem 3.3 extends these limiting strategies to the $N$-player game: for any unilateral deviation, the cost gain is at most $\\epsilon^\\alpha_N=O\\!\\left(\\sqrt{1/N}\\,\\sqrt{\\alpha(1-\\alpha)}\\big/p(T,\\bar q^\\alpha_T)\\right)$, obtained from the central limit theorem for sample quantiles.","pith_inferences":["Editorial inference: The identity $\\pi_t=-\\eta_t$ is likely a general phenomenon for symmetric quadratic terminal costs whose coefficients cancel; if so, the decoupling ansatz (3.34) would transfer to other target-type MFGs, including heterogeneous-agent versions, as long as the representative-agent Gaussianity is preserved.","Editorial inference: The error rate (3.43) contains the factor $1/p(T,\\bar q^\\alpha_T)$, so the approximation should worsen in low-density regions of the terminal distribution, for instance when $\\alpha$ is very close to 0 or 1, or when volatility is small and the density at the quantile is thin; the paper's numerics do not probe this boundary regime.","Editorial inference: Since the VC coordinator's funding allocation is treated as an exogenous deterministic support $\\gamma_t$, a natural extension is a principal-agent layer where $\\gamma_t$ is chosen optimally; the target-based FBODEs could serve as the reduced-form constraint in such a design problem."],"forward_implications":["The target-based equilibrium can be computed by solving five decoupled scalar ODEs, so the quantile path and effort policy are available without fixed-point iteration.","In the finite-$N$ game, any agent who deviates from the proposed strategy gains at most $\\epsilon^\\alpha_N$, which vanishes as $N\\to\\infty$; hence the limiting strategy is asymptotically a Nash equilibrium.","The same $\\epsilon$-Nash argument is claimed to cover quantilized games with quantile-of-control interactions and to remove the earlier uniform boundedness assumption on square-integrable deviating strategies.","The threshold-based equilibrium has no closed form, but in the venture-capital experiments its quantile stays systematically slightly above the target-based one, and the two distributions of terminal values are close."],"supporting_citations":[{"why":"Supplies the existence and uniqueness result for the Riccati ODE in Proposition 3.2.","marker":"[45]"},{"why":"Provides the central limit theorem for sample quantiles that yields the epsilon-Nash rate (3.43).","marker":"[46]"},{"why":"Introduces quantilized mean-field games and the quantile consistency condition that the present ranking model extends.","marker":"[21]"},{"why":"Establishes the LQG quantilized MFG setting underlying the target-based derivation.","marker":"[22]"},{"why":"Recent quantile-dependent-cost MFG work whose lack of an epsilon-Nash property motivates the new argument.","marker":"[23]"},{"why":"Quantile-of-control MFG whose epsilon-Nash proof the paper says its own argument relaxes.","marker":"[43]"},{"why":"Fixed-point numerical scheme adapted to solve the threshold-based consistency condition.","marker":"[16]"}],"fun_headline_variants":["Quantile ranking games reach near-Nash equilibrium","Closed-form equilibrium for quantile mean-field games","VC start-up contests get near-optimal strategies","Target-based quantile games yield analytic equilibrium","Provably near-optimal play in quantile ranking games"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof of the $\\epsilon$-Nash property assumes that the strategy it identifies as the best deviating response is actually available to the agent; but that strategy's terminal condition involves the final sample quantile, which is not known until time $T$, so the strategy may not be adapted to the agent's information flow and may fall outside the admissible set.","fun_headline_variants_meta":{"raw":{"variants":["Quantile ranking games reach near-Nash equilibrium","Closed-form equilibrium for quantile mean-field games","VC start-up contests get near-optimal strategies","Target-based quantile games yield analytic equilibrium","Provably near-optimal play in quantile ranking games"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000325,"raw_usage":{"total_tokens":1922,"prompt_tokens":1145,"completion_tokens":777,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":761,"completion_tokens_details":{"reasoning_tokens":705}},"tokens_in":761,"tokens_out":777,"duration_ms":9003,"temperature":1.0,"reasoning_tokens":705,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:06:58.343981+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether $\\hat u^i_t = -\\frac{b}{r}(\\eta_t \\hat x^i_t+\\hat\\theta^{\\alpha,[N]}_t)$ with $\\hat\\theta^{\\alpha,[N]}_T=-\\lambda q_T^{\\alpha,[N]}$ is adapted: for $N$ finite and $\\alpha\\in(0,1)$, $q_T^{\\alpha,[N]}$ is not $\\sigma(x^i_0,w^i_s: s\\le t)$-measurable for $t<T$, so the backward equation (3.50) is anticipating. A direct calculation of $J_i^{[N]}(\\hat u^i,\\cdot)$ compared with the true dynamic-programming optimum for a deviator would show whether the non-adapted control can be replaced by an admissible one that achieves the same cost; if not, Step 1 of the proof (relation (3.48)) has no justification and the claimed bound is not established by the given argument.","supporting_citations":[{"cited_title":"Freiling, G","cited_arxiv_id":null,"evidence_quote":"Supplies the existence and uniqueness result for the Riccati ODE in Proposition 3.2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the central limit theorem for sample quantiles that yields the epsilon-Nash rate (3.43)."},{"cited_title":"Foguen-Tchuendom, R","cited_arxiv_id":null,"evidence_quote":"Introduces quantilized mean-field games and the quantile consistency condition that the present ranking model extends."},{"cited_title":"Foguen-Tchuendom, R","cited_arxiv_id":null,"evidence_quote":"Establishes the LQG quantilized MFG setting underlying the target-based derivation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Recent quantile-dependent-cost MFG work whose lack of an epsilon-Nash property motivates the new argument."},{"cited_title":"Foguen-Tchuendom, R","cited_arxiv_id":null,"evidence_quote":"Quantile-of-control MFG whose epsilon-Nash proof the paper says its own argument relaxes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Fixed-point numerical scheme adapted to solve the threshold-based consistency condition."}],"review_version":1}