Pith. sign in

REVIEW 3 major objections 3 minor 2 cited by

Convergence Rates of Time Discretization in Extended Mean Field Control

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read For linear-convex extended mean field control problems, this paper proves piecewise constant policies approximate the optimal cost with order 1/2 and the optimal control with order 1/4, and proves first-order value convergence under smoothn

desk verdict Genuinely new convergence rates for time-discretized extended MFC, but the half-order results are conditional on an assumption verified only in special cases. read the letter →

arxiv 2509.00904 v1 pith:F5WQ4WXD submitted 2025-08-31 math.OC cs.NAmath.NA

classification math.OCcs.NAmath.NA MSC 49N8049N6060H3565L70
keywords extendedmeanfieldcontrolcontrolledMcKean–Vlasovdiffusionpiecewiseconstanttimediscretizationerrorestimatepathregularity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper works out how well a continuous-time extended mean field control problem can be solved by discretizing time and using controls that are constant on each interval—the piecewise constant approximation at the heart of most numerical schemes. For linear-convex problems, it proves that the true optimal control is 1/2-Hölder continuous in time, and from that regularity derives two rates: the optimal cost is matched to within order 1/2 of the step size, and any near-optimal discrete control is within order 1/4 of the true optimal control in L2. For general extended MFC problems with sufficiently smooth value functions, it proves the value functions converge at first order, matching the best rate known for classical stochastic control without mean-field interaction. These are the first such rates for extended MFC, and the control approximation result is new even for classical control problems.

What carries the argument

Two mechanisms carry the argument. In the linear-convex part, a feedback map α̂ (Assumption H.2) realizes the optimality condition pointwise: given the current state, adjoint process, and their joint law, it returns the optimal control value, with a prescribed 1/2-Hölder time regularity. This map reduces the generally non-Markovian Pontryagin system to a coupled McKean–Vlasov FBSDE; Lipschitz and monotonicity properties of the reduced coefficients give well-posedness by a continuation method, and BSDE regularity bounds yield the 1/2-Hölder path regularity of X and Y, hence of the optimal control. In the smooth part, iterated Itô expansions along measure flows—through the operators Lx and L(x

What would settle it

Take a linear-convex extended MFC problem satisfying the structural assumptions but whose optimality condition cannot be written through a pointwise feedback map, such as the non-Markovian quadratic example identified in the paper's reference [1]. Compute the optimal control's time increments and the discrete-time value errors for increasing numbers of steps: if the control modulus is worse than 1/2-Hölder, or if the value error decays slower than N^{-1/2}, then the feedback-map assumption is essential to the claimed rates.

Watch

Extended reading notes

Core claim

The central claim is that piecewise constant controls are not only computationally convenient; their approximation error can be quantified. Under linear-convex structure, the unique optimal control is 1/2-Hölder in time, the discrete-time value error is O(|π|^{1/2}), and the H2 error of the optimal control is O(|π|^{1/4}). Under sufficient regularity of the value function and its measure derivatives, the value error improves to first order. The half-order results follow by writing the optimality condition through the stochastic maximum principle, representing it with a feedback map, and reducing the non-Markovian Pontryagin system to a McKean–Vlasov forward-backward SDE whose path regularity

Load-bearing premise

The linear-convex rates are conditional on the existence of a feedback map that satisfies the pointwise optimality condition and is 1/2-Hölder in time; the paper verifies this map only for special cost structures, so the rates are not proven for every linear-convex problem.

Editorial extensions

If this is right

  • For linear-convex extended MFC, an N-step piecewise constant scheme approximates the optimal cost with error O(N^{-1/2}), so halving the error requires quadrupling the number of time steps.
  • Near-optimal discrete controls are within O(N^{-1/4}+√ε) of the true optimal control in L2, a strong convergence guarantee for discretized control processes that was not previously available even in classical control.
  • Under enough value-function smoothness, the mean-field dependence does not slow the discretization: value functions converge at first order, matching the best-known rate for classical control.
  • The 1/2-Hölder regularity of the optimal control is the input that carries both linear-convex rates; sharper regularity in special cases would automatically imply faster discretization rates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The value-rate and control-rate differ by a factor of two through the strong convexity gap of the cost, suggesting a general conversion principle: in convex MFC settings, cost accuracy can be converted into control accuracy with a square-root loss of rate.
  • The paper leaves open whether the half-order rates hold for linear-convex problems where the pointwise optimality condition cannot be realized by a feedback map; testing the non-Markovian quadratic example flagged in the paper would show what slower rate emerges there.
  • The first-order result requires bounded derivatives of the value function in both space and measure; relaxing to semi-concavity or semi-convexity, as in the classical 1/3-rate theory, is a natural next step for non-smooth mean-field models.
  • Because the value rate is twice the control rate, numerical schemes that stop when value estimates stabilize may overstate the accuracy of the implied policy; directly monitoring policy differences would be a safer stopping rule.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper analyzes the error of piecewise-constant time discretization for extended mean field control problems. In the first part, for linear-convex problems with affine drift and uncontrolled diffusion, the authors use the stochastic maximum principle and a coupled MV-FBSDE to prove that the optimal control is 1/2-Hölder continuous in time (Theorem 2.5), and then derive order-1/2 convergence of the value function and order-1/4 strong convergence of the optimal control (Theorems 2.6 and 2.7). These results require an additional feedback-map assumption H.2, which is verified only in special cases (Propositions 2.2--2.4). In the second part, under the regularity assumption H.5 on the data and the piecewise-constant value function, the authors prove first-order convergence of the value functions (Theorem 2.8) via an iterated Itô expansion and an approximate dynamic programming principle, and illustrate the rate numerically on controlled Cucker--Smale models.

Significance. The conditional theorems are technically substantial. The 1/2-Hölder regularity result and, in particular, the strong $L^2$ convergence of discrete optimal controls (Theorem 2.7) appear new even for classical stochastic control without mean-field interaction. The first-order value convergence in Theorem 2.8 extends the best-known rate from [16] to extended MFC with control-law dependence, and the Cucker--Smale experiments provide useful numerical confirmation. The paper is also honest about its limitations: Remark 2.2 flags the difficulty of verifying H.2, and Remark 2.5 explicitly acknowledges that H.5 is not established even in the classical setting. The main weakness is that the advertised headline claims in the abstract are broader than the proved theorems: the half-order results require H.2, and the first-order result requires H.5. If the authors revise the presentation to state the precise assumptions and provide the missing proof in Theorem 3.4, the paper would be a solid contribution.

major comments (3)
  1. [Abstract and §2.1.2 (H.2)] The abstract and Introduction state the half-order results for 'linear-convex extended MFC problems' without flagging H.2. However, Theorems 2.5–2.7 require H.2, a Lipschitz feedback map α̂ satisfying the pointwise optimality condition (2.9), in addition to the linear-convex structure H.1. H.2 is not implied by H.1; Remark 2.2 itself says that constructing such a map for general coefficients 'appears challenging.' Propositions 2.2–2.4 verify H.2 only for special classes: f independent of the control-law marginal, separable f with b2 = 0, and quadratic-in-control f in one dimension. Thus the class of problems for which the half-order rates are proved is strictly smaller than advertised. Please either prove H.2 under H.1 (or a clearly stated sub-class), or revise the abstract and introduction to state that the half-order results are conditional on H.2 and its verified special cases.
  2. [§3.2, proof of Theorem 3.4] The proof of Theorem 3.4 applies [30, Theorem 5.2.2(i)] to the decoupled FBSDE (3.11) and then states that the proof 'can be extended' to F0-measurable initial data and time-measurable coefficients. This extension is not proved. The step is load-bearing: it yields the bound |Z_t| ≤ C|σ(t, X_t, P_{X_t})| and consequently the Hölder regularity of Y, which is used in Theorem 2.5 and then in Theorems 2.6–2.7. Without a rigorous justification of the extension, the 1/2-Hölder regularity theorem is incomplete as stated. Please provide a proof of the extension or replace the reference by a theorem that covers the present generality.
  3. [§2.2, H.5 and Theorem 2.8] Theorem 2.8 is stated under H.5, but H.5(2) requires Vπ^c ∈ C^{1,2}_2 and boundedness of L^{(x,a)}L^x Vπ^c. Remark 2.5 explicitly acknowledges that the existence of these derivatives for the piecewise-constant-control value function 'has not established' even without mean-field interaction. Hence the first-order convergence result is conditional on an unverified regularity property. While the abstract says 'under sufficient regularity of the value functions,' the phrasing 'we further show' gives the impression of a proved general theorem. Please reframe Theorem 2.8 as a conditional result and make the strength of H.5 more prominent in the abstract and introduction.
minor comments (3)
  1. [Abstract and Theorem 2.6] The order-1/2 result for |Vπ(ξ0) − V(ξ0)| requires A compact; without compactness only the one-sided bound Vπ − V ≤ O(|π|^{1/2}) is proved. Please mention this restriction when summarizing the cost approximation result.
  2. [Lemma 2.1, Eq. (2.7)] In the terminal condition for the adjoint equation (2.7), the expression uses P_{X_t^α} in the second term; this appears to be a typo and should read P_{X_T^α}. Please correct.
  3. [Proof of Proposition 3.1] The derivation of the Lipschitz bound for ϕ via Lemma A.1 is very compressed, especially the step 'another application of Lemma A.1 gives...'. A short explanation of how the supremum with the composed functions is converted to an infimum over couplings would improve readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: main results are conditional theorems proved from explicit assumptions; the key limitations are scope (H.2/H.5) rather than circular reasoning.

full rationale

The derivation chain is self-contained in the sense that every main theorem is a conditional statement proved from stated assumptions using standard external tools (stochastic maximum principle, continuation methods for MV-FBSDEs, Itô calculus, Malliavin regularity). No parameter is fitted to a target quantity, and no conclusion is used as an input to its own proof. Assumption H.2 postulates a feedback map satisfying the pointwise optimality condition (2.9) with 1/2-Hölder time regularity; the paper explicitly verifies H.2 in special cases (Propositions 2.2–2.4) and states in Remark 2.2 that constructing such a map for general coefficients is challenging. Thus the abstract's phrasing that the 1/2-Hölder regularity holds for 'linear-convex extended MFC problems' is broader than the proven conditional scope, but this is a scope/limitation issue, not circularity. Similarly, Theorem 2.8 assumes H.5, including regularity of the discrete value function V^c_π, and then proves first-order convergence; this is a standard conditional regularity result, and the paper itself flags that existence of such derivatives has not been established (Remark 2.5). Self-citations [16] and [26] are used for comparison, technique extension, and a numerical benchmark, not as load-bearing justification of the new convergence rates. No equation is defined in terms of a target result, and no 'prediction' reduces by construction to a fitted input or to an assumed version of itself.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper's theorems are conditional on assumptions H.2 and H.5 which are not proven for the general case. H.2 is verified in three special classes, so the abstract's claim for all linear-convex problems overstates the result. H.5 is a high-regularity assumption whose validity for piecewise constant controls is acknowledged as unproven.

assumptions (6)
  • domain assumption Pontryagin maximum principle for extended MFC (Lemma 2.1, from [1])
    Used to characterize optimal controls via adjoint process.
  • domain assumption Well-posedness of McKean-Vlasov FBSDE via continuation method (Proposition 3.3, based on [3,5,14])
    Ensures existence/uniqueness of the coupled forward-backward system.
  • standard math Ito formula for flows of measures (Theorem B.1 from [6])
    Used for local expansions in the dynamic programming proof.
  • ad hoc to paper Existence of feedback map in H.2 (assumed, verified in special cases)
    The half-order rates depend on this assumption; Propositions 2.2-2.4 only cover three special classes.
  • ad hoc to paper Regularity of value function in H.5 (assumed, not established even in classical case)
    First-order rate requires high-order smoothness of V_c^pi, acknowledged as unproven in Remark 2.5.
  • ad hoc to paper Extension of Zhang's path regularity theorem to random initial data and time-measurable coefficients (Theorem 3.4)
    An unproven extension of [30, Theorem 5.2.2] is asserted to justify bounded Z.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Convergence Rates of Time Discretization in Extended Mean Field Control." pith.science (2026). https://pith.science/paper/F5WQ4WXD

@misc{pith2026250900904,
  author       = {Pith},
  title        = {Pith review of: Convergence Rates of Time Discretization in Extended Mean Field Control},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F5WQ4WXD}},
  note         = {Machine review of arXiv:2509.00904}
}
abstract

Piecewise constant control approximation provides a practical framework for designing numerical schemes of continuous-time control problems. We analyze the accuracy of such approximations for extended mean field control (MFC) problems, where the dynamics and costs depend on the joint distribution of states and controls. For linear-convex extended MFC problems, we show that the optimal control is $1/2$-H\"older continuous in time. Using this regularity, we prove that the optimal cost of the continuous-time problem can be approximated by piecewise constant controls with order $1/2$, while the optimal control itself can be approximated with order $1/4$. For general extended MFC problems, we further show that, under sufficient regularity of the value functions, the value functions converge with an improved first-order rate, matching the best-known rate for classical control problems without mean field interaction, and consistent with the numerical observations for MFC of Cucker-Smale models.

Figures

Figures reproduced from arXiv: 2509.00904 by the authors.

Figure 1
Figure 1. Value function convergence (β = 0, γ1 = 0.1, d = 1) [PITH_FULL_IMAGE:figures/full_fig_p029_1.png] view at source ↗
Figure 4
Figure 4. Convergence in time (β = 0, γ1 = 0.1, d = 3) Finally, we consider β = 1, which gives a non-trivial coupling between x and v and a fully 2d–dimensional models. We keep the rest of the parameters the same as before, while K = 400. Recall that γ1 = 0.1, T = 1, N = O(103 ), d = 1, M = 128, and σ = 0.1. This is not a LQ model and therefore we do not have an explicit solution to the optimal feedback control. Nonetheless, … view at source ↗
Figure 6
Figure 6. Convergence in time (β = 1, γ1 = 0.1, d = 1) A Proofs of Propositions 3.1 and 3.2 and 3.3 To prove Proposition 3.1, we recall the following Kantorovich duality theorem, which follows as a special case of [28, Theorem 5.10]. Lemma A.1. Let (X , µ) and (Y, ν) be two Polish probability spaces and let ω : X × Y → [0, ∞) be a continuous function. Then we have that inf κ∈Π(µ,ν) Z X ×Y ω(x, y) dκ(x, y) = sup (ψ,φ)∈Cb(X)×Cb… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

    math.OC 2026-07 accept novelty 6.0 of 10

    Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.

  2. NeuralChaos: Optimal Adapted Approximation of Square Integrable Predictable Processes

    math.PR 2026-07 conditional novelty 5.0 of 10

    A finite-sampling neural architecture is dense in the Hilbert space of square-integrable predictable processes and attains best-N-term chaoslet rates for compressible or Malliavin-regular processes.

Reference graph

Works this paper leans on

30 extracted references · 26 canonical work pages · cited by 2 Pith papers

  1. [16]

    Improved order 1 /4 convergence of Krylov’s piecewise constant policy approximation

    E. R. Jakobsen, A. Picarelli, and C. Reisinger. “Improved order 1 /4 convergence of Krylov’s piecewise constant policy approximation”. Electron. Comm. Probab. 24 (2019), pp. 1–10

  2. [30]

    Backward Stochastic Differential Equations: From Linear to Fully Nonlinear Theory

    J. Zhang. “Backward Stochastic Differential Equations: From Linear to Fully Nonlinear Theory”. Springer, New York 86 (2017). 38

  3. [1]

    Extended mean field control prob- lems: stochastic maximum principle and transport perspective

    B. Acciaio, J. Backhoff-Veraguas, and R. Carmona. “Extended mean field control prob- lems: stochastic maximum principle and transport perspective”. SIAM J. Control Optim. 57 (2019), pp. 3666–3693

  4. [2]

    Linear-quadratic McKean-Vlasov stochastic control problems with random coefficients on finite and infinite horizon, and applications

    M. Basei and H. Pham. “Linear-quadratic McKean-Vlasov stochastic control problems with random coefficients on finite and infinite horizon, and applications”. arXiv preprint arXiv:1711.09390 (2017)

  5. [3]

    Well-posedness of mean-field type forward-backward stochastic differential equations

    A. Bensoussan, S. Yam, and Z. Zhang. “Well-posedness of mean-field type forward-backward stochastic differential equations”. Stochastic Process. Appl. 125.9 (2015), pp. 3327–3354

  6. [4]

    Viscosity solutions for controlled McKean-Vlasov jump-diffusions

    M. Burzoni, V. Ignazio, H. Soner, and A. M. Reppen. “Viscosity solutions for controlled McKean-Vlasov jump-diffusions”. SIAM J. Control Optim. 58.3 (2020), pp. 1676–1699

  7. [5]

    Forward-backward stochastic differential equations and con- trolled McKean-Vlasov dynamics

    R. Carmona and F. Delarue. “Forward-backward stochastic differential equations and con- trolled McKean-Vlasov dynamics”. Ann. Probab. 43.5 (2015), pp. 2647–2700

  8. [6]

    Probabilistic theory of mean field games with applications I: Mean-field FBSDEs, control, and games

    R. Carmona and F. Delarue. “Probabilistic theory of mean field games with applications I: Mean-field FBSDEs, control, and games”. Probability Theory and Stochastic Modelling, Springer 83 (2018)

Show all 30 references
  1. [7]

    Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: II-The finite horizon case

    R. Carmona and M. Lauri` ere. “Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: II-The finite horizon case”. Ann. Appl. Probab. 32.6 (2022), pp. 4065–4105

  2. [8]

    A probabilistic approach to classical solutions of the master equation for large population equilibria

    J. F. Chassagneux, D. Crisan, and F. Delarue. “A probabilistic approach to classical solutions of the master equation for large population equilibria”. Mem. Amer. Math. Soc. 280.1379 (2022)

  3. [9]

    Master Bellman equation in the Wasserstein space: Uniqueness of viscosity solutions

    A. Cosso, F. Gozzi, I. Kharroubi, H. Pham, and M. Rosestolato. “Master Bellman equation in the Wasserstein space: Uniqueness of viscosity solutions”. Trans. Amer. Math. Soc. 377.1 (2024), pp. 31–83

  4. [10]

    Emergent behavior in flocks

    F. Cucker and S. Smale. “Emergent behavior in flocks”. IEEE Trans. Automat. Control 52.5 (2007), pp. 852–862

  5. [11]

    Extended mean field control problem: a propagation of chaos result

    M. F. Djete. “Extended mean field control problem: a propagation of chaos result”. Electron. J. Probab. 27 (2022), pp. 1–53

  6. [12]

    McKean-Vlasov optimal control: the dynamic programming principle

    M. F. Djete, D. Possamai, and X. Tan. “McKean-Vlasov optimal control: the dynamic programming principle”. Ann. Probab. 50.2 (2022), pp. 791–833

  7. [13]

    Extended McKean-Vlasov optimal stochastic control applied to smart grid management

    E. Gobet and M. Grangereau. “Extended McKean-Vlasov optimal stochastic control applied to smart grid management”. ESAIM Contr. Op. Ca. Va. 28 (2022), p. 40

  8. [14]

    Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls

    X. Guo, A. Hu, and Y. Zhang. “Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls”. SIAM J. Control and Optim. 61.2 (2023), pp. 755–787

  9. [15]

    Itˆ o’s formula for flows of measures on semimartingales

    X. Guo, H. Pham, and X. Wei. “Itˆ o’s formula for flows of measures on semimartingales”. Stochastic Process. Appl. 159 (2023), pp. 350–390

  10. [17]

    Adam: A Method for Stochastic Optimization

    D. P. Kingma and J. Ba. “Adam: A Method for Stochastic Optimization”. arXiv preprint arXiv:1412.6980 (2014). 37

  11. [18]

    Approximating value functions for controlled degenerate diffusion processes by using piece-wise constant policies

    N. V. Krylov. “Approximating value functions for controlled degenerate diffusion processes by using piece-wise constant policies”. Electron. J. Probab. 4 (1999), pp. 1–19

  12. [19]

    Convergence of large population games to mean field games with interaction through the controls

    M. Lauri` ere and L. Tangpi. “Convergence of large population games to mean field games with interaction through the controls”. SIAM J. Math. Anal. 54.3 (2022), pp. 3535–3574

  13. [20]

    On the convergence rate for piece-wise constant policy approximation of stochastic optimal control problem

    T. Legrand. “On the convergence rate for piece-wise constant policy approximation of stochastic optimal control problem”. MA thesis. NTNU, 2023

  14. [21]

    Mean field analysis of controlled Cucker- Smale type flocking: Linear analysis and perturbation equations

    M. Nourian, P. E. Caines, and R. P. Malham´ e. “Mean field analysis of controlled Cucker- Smale type flocking: Linear analysis and perturbation equations”. IF AC Proc. Vol. 44.1 (2011), pp. 4471–4476

  15. [22]

    Mean-field neural networks-based algorithms for McKean-Vlasov control problems

    H. Pham and X. Warin. “Mean-field neural networks-based algorithms for McKean-Vlasov control problems”. arXiv preprint arXiv:2212.11518 (2022)

  16. [23]

    Bellman equation and viscosity solutions for mean-field stochastic control problem

    H. Pham and X. Wei. “Bellman equation and viscosity solutions for mean-field stochastic control problem”. ESAIM Contr. Op. Ca. Va. 24 (2018), pp. 437–461

  17. [24]

    Extended mean field control: a finite-dimensional numerical approximation

    A. Picarelli, M. Scaratti, and J. Tam. “Extended mean field control: a finite-dimensional numerical approximation”. arXiv preprint arXiv:2503.20510 (2025)

  18. [25]

    Freidlin-Wentzell LDP in path space for McKean- Vlasov equations and the functional iterated logarithm law

    G. dos Reis, W. Salkeld, and J. Tugaut. “Freidlin-Wentzell LDP in path space for McKean- Vlasov equations and the functional iterated logarithm law”.Ann. Appl. Probab. 29.3 (2019), pp. 1487–1540

  19. [26]

    A fast iterative PDE-based algorithm for feedback controls of nonsmooth mean-field control problems

    C. Reisinger, W. Stockinger, and Y. Zhang. “A fast iterative PDE-based algorithm for feedback controls of nonsmooth mean-field control problems”. SIAM J. Sci. Comput. 46.4 (2024), A2737–A2773

  20. [27]

    Exploration-exploitation trade-off for continuous- time episodic reinforcement learning with linear-convex models

    L. Szpruch, T. Treetanthiploet, and Y. Zhang. “Exploration-exploitation trade-off for continuous- time episodic reinforcement learning with linear-convex models”.arXiv preprint arXiv:2112.10264 (2021)

  21. [28]

    Optimal Transport: Old and New

    C. Villani. “Optimal Transport: Old and New”. Springer-Verlag, Berlin (2009)

  22. [29]

    A linear-quadratic optimal control problem for mean-field stochastic differential equations

    J. Yong. “A linear-quadratic optimal control problem for mean-field stochastic differential equations”. SIAM J. Control Optim. 51.4 (2013), pp. 2809–2838

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.