REVIEW 3 major objections 3 minor 2 cited by
Convergence Rates of Time Discretization in Extended Mean Field Control
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read For linear-convex extended mean field control problems, this paper proves piecewise constant policies approximate the optimal cost with order 1/2 and the optimal control with order 1/4, and proves first-order value convergence under smoothn
desk verdict Genuinely new convergence rates for time-discretized extended MFC, but the half-order results are conditional on an assumption verified only in special cases. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two mechanisms carry the argument. In the linear-convex part, a feedback map α̂ (Assumption H.2) realizes the optimality condition pointwise: given the current state, adjoint process, and their joint law, it returns the optimal control value, with a prescribed 1/2-Hölder time regularity. This map reduces the generally non-Markovian Pontryagin system to a coupled McKean–Vlasov FBSDE; Lipschitz and monotonicity properties of the reduced coefficients give well-posedness by a continuation method, and BSDE regularity bounds yield the 1/2-Hölder path regularity of X and Y, hence of the optimal control. In the smooth part, iterated Itô expansions along measure flows—through the operators Lx and L(x
What would settle it
Take a linear-convex extended MFC problem satisfying the structural assumptions but whose optimality condition cannot be written through a pointwise feedback map, such as the non-Markovian quadratic example identified in the paper's reference [1]. Compute the optimal control's time increments and the discrete-time value errors for increasing numbers of steps: if the control modulus is worse than 1/2-Hölder, or if the value error decays slower than N^{-1/2}, then the feedback-map assumption is essential to the claimed rates.
Extended reading notes
Core claim
The central claim is that piecewise constant controls are not only computationally convenient; their approximation error can be quantified. Under linear-convex structure, the unique optimal control is 1/2-Hölder in time, the discrete-time value error is O(|π|^{1/2}), and the H2 error of the optimal control is O(|π|^{1/4}). Under sufficient regularity of the value function and its measure derivatives, the value error improves to first order. The half-order results follow by writing the optimality condition through the stochastic maximum principle, representing it with a feedback map, and reducing the non-Markovian Pontryagin system to a McKean–Vlasov forward-backward SDE whose path regularity
Load-bearing premise
The linear-convex rates are conditional on the existence of a feedback map that satisfies the pointwise optimality condition and is 1/2-Hölder in time; the paper verifies this map only for special cost structures, so the rates are not proven for every linear-convex problem.
Editorial extensions
If this is right
- For linear-convex extended MFC, an N-step piecewise constant scheme approximates the optimal cost with error O(N^{-1/2}), so halving the error requires quadrupling the number of time steps.
- Near-optimal discrete controls are within O(N^{-1/4}+√ε) of the true optimal control in L2, a strong convergence guarantee for discretized control processes that was not previously available even in classical control.
- Under enough value-function smoothness, the mean-field dependence does not slow the discretization: value functions converge at first order, matching the best-known rate for classical control.
- The 1/2-Hölder regularity of the optimal control is the input that carries both linear-convex rates; sharper regularity in special cases would automatically imply faster discretization rates.
Reading between the lines
- The value-rate and control-rate differ by a factor of two through the strong convexity gap of the cost, suggesting a general conversion principle: in convex MFC settings, cost accuracy can be converted into control accuracy with a square-root loss of rate.
- The paper leaves open whether the half-order rates hold for linear-convex problems where the pointwise optimality condition cannot be realized by a feedback map; testing the non-Markovian quadratic example flagged in the paper would show what slower rate emerges there.
- The first-order result requires bounded derivatives of the value function in both space and measure; relaxing to semi-concavity or semi-convexity, as in the classical 1/3-rate theory, is a natural next step for non-smooth mean-field models.
- Because the value rate is twice the control rate, numerical schemes that stop when value estimates stabilize may overstate the accuracy of the implied policy; directly monitoring policy differences would be a safer stopping rule.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes the error of piecewise-constant time discretization for extended mean field control problems. In the first part, for linear-convex problems with affine drift and uncontrolled diffusion, the authors use the stochastic maximum principle and a coupled MV-FBSDE to prove that the optimal control is 1/2-Hölder continuous in time (Theorem 2.5), and then derive order-1/2 convergence of the value function and order-1/4 strong convergence of the optimal control (Theorems 2.6 and 2.7). These results require an additional feedback-map assumption H.2, which is verified only in special cases (Propositions 2.2--2.4). In the second part, under the regularity assumption H.5 on the data and the piecewise-constant value function, the authors prove first-order convergence of the value functions (Theorem 2.8) via an iterated Itô expansion and an approximate dynamic programming principle, and illustrate the rate numerically on controlled Cucker--Smale models.
Significance. The conditional theorems are technically substantial. The 1/2-Hölder regularity result and, in particular, the strong $L^2$ convergence of discrete optimal controls (Theorem 2.7) appear new even for classical stochastic control without mean-field interaction. The first-order value convergence in Theorem 2.8 extends the best-known rate from [16] to extended MFC with control-law dependence, and the Cucker--Smale experiments provide useful numerical confirmation. The paper is also honest about its limitations: Remark 2.2 flags the difficulty of verifying H.2, and Remark 2.5 explicitly acknowledges that H.5 is not established even in the classical setting. The main weakness is that the advertised headline claims in the abstract are broader than the proved theorems: the half-order results require H.2, and the first-order result requires H.5. If the authors revise the presentation to state the precise assumptions and provide the missing proof in Theorem 3.4, the paper would be a solid contribution.
major comments (3)
- [Abstract and §2.1.2 (H.2)] The abstract and Introduction state the half-order results for 'linear-convex extended MFC problems' without flagging H.2. However, Theorems 2.5–2.7 require H.2, a Lipschitz feedback map α̂ satisfying the pointwise optimality condition (2.9), in addition to the linear-convex structure H.1. H.2 is not implied by H.1; Remark 2.2 itself says that constructing such a map for general coefficients 'appears challenging.' Propositions 2.2–2.4 verify H.2 only for special classes: f independent of the control-law marginal, separable f with b2 = 0, and quadratic-in-control f in one dimension. Thus the class of problems for which the half-order rates are proved is strictly smaller than advertised. Please either prove H.2 under H.1 (or a clearly stated sub-class), or revise the abstract and introduction to state that the half-order results are conditional on H.2 and its verified special cases.
- [§3.2, proof of Theorem 3.4] The proof of Theorem 3.4 applies [30, Theorem 5.2.2(i)] to the decoupled FBSDE (3.11) and then states that the proof 'can be extended' to F0-measurable initial data and time-measurable coefficients. This extension is not proved. The step is load-bearing: it yields the bound |Z_t| ≤ C|σ(t, X_t, P_{X_t})| and consequently the Hölder regularity of Y, which is used in Theorem 2.5 and then in Theorems 2.6–2.7. Without a rigorous justification of the extension, the 1/2-Hölder regularity theorem is incomplete as stated. Please provide a proof of the extension or replace the reference by a theorem that covers the present generality.
- [§2.2, H.5 and Theorem 2.8] Theorem 2.8 is stated under H.5, but H.5(2) requires Vπ^c ∈ C^{1,2}_2 and boundedness of L^{(x,a)}L^x Vπ^c. Remark 2.5 explicitly acknowledges that the existence of these derivatives for the piecewise-constant-control value function 'has not established' even without mean-field interaction. Hence the first-order convergence result is conditional on an unverified regularity property. While the abstract says 'under sufficient regularity of the value functions,' the phrasing 'we further show' gives the impression of a proved general theorem. Please reframe Theorem 2.8 as a conditional result and make the strength of H.5 more prominent in the abstract and introduction.
minor comments (3)
- [Abstract and Theorem 2.6] The order-1/2 result for |Vπ(ξ0) − V(ξ0)| requires A compact; without compactness only the one-sided bound Vπ − V ≤ O(|π|^{1/2}) is proved. Please mention this restriction when summarizing the cost approximation result.
- [Lemma 2.1, Eq. (2.7)] In the terminal condition for the adjoint equation (2.7), the expression uses P_{X_t^α} in the second term; this appears to be a typo and should read P_{X_T^α}. Please correct.
- [Proof of Proposition 3.1] The derivation of the Lipschitz bound for ϕ via Lemma A.1 is very compressed, especially the step 'another application of Lemma A.1 gives...'. A short explanation of how the supremum with the composed functions is converted to an infimum over couplings would improve readability.
Circularity Check
No significant circularity: main results are conditional theorems proved from explicit assumptions; the key limitations are scope (H.2/H.5) rather than circular reasoning.
full rationale
The derivation chain is self-contained in the sense that every main theorem is a conditional statement proved from stated assumptions using standard external tools (stochastic maximum principle, continuation methods for MV-FBSDEs, Itô calculus, Malliavin regularity). No parameter is fitted to a target quantity, and no conclusion is used as an input to its own proof. Assumption H.2 postulates a feedback map satisfying the pointwise optimality condition (2.9) with 1/2-Hölder time regularity; the paper explicitly verifies H.2 in special cases (Propositions 2.2–2.4) and states in Remark 2.2 that constructing such a map for general coefficients is challenging. Thus the abstract's phrasing that the 1/2-Hölder regularity holds for 'linear-convex extended MFC problems' is broader than the proven conditional scope, but this is a scope/limitation issue, not circularity. Similarly, Theorem 2.8 assumes H.5, including regularity of the discrete value function V^c_π, and then proves first-order convergence; this is a standard conditional regularity result, and the paper itself flags that existence of such derivatives has not been established (Remark 2.5). Self-citations [16] and [26] are used for comparison, technique extension, and a numerical benchmark, not as load-bearing justification of the new convergence rates. No equation is defined in terms of a target result, and no 'prediction' reduces by construction to a fitted input or to an assumed version of itself.
Assumptions & free parameters
assumptions (6)
- domain assumption Pontryagin maximum principle for extended MFC (Lemma 2.1, from [1])
- domain assumption Well-posedness of McKean-Vlasov FBSDE via continuation method (Proposition 3.3, based on [3,5,14])
- standard math Ito formula for flows of measures (Theorem B.1 from [6])
- ad hoc to paper Existence of feedback map in H.2 (assumed, verified in special cases)
- ad hoc to paper Regularity of value function in H.5 (assumed, not established even in classical case)
- ad hoc to paper Extension of Zhang's path regularity theorem to random initial data and time-measurable coefficients (Theorem 3.4)
Cite this review
Pith. "Pith review of Convergence Rates of Time Discretization in Extended Mean Field Control." pith.science (2026). https://pith.science/paper/F5WQ4WXD
@misc{pith2026250900904,
author = {Pith},
title = {Pith review of: Convergence Rates of Time Discretization in Extended Mean Field Control},
year = {2026},
howpublished = {\url{https://pith.science/paper/F5WQ4WXD}},
note = {Machine review of arXiv:2509.00904}
}
abstract
Piecewise constant control approximation provides a practical framework for designing numerical schemes of continuous-time control problems. We analyze the accuracy of such approximations for extended mean field control (MFC) problems, where the dynamics and costs depend on the joint distribution of states and controls. For linear-convex extended MFC problems, we show that the optimal control is $1/2$-H\"older continuous in time. Using this regularity, we prove that the optimal cost of the continuous-time problem can be approximated by piecewise constant controls with order $1/2$, while the optimal control itself can be approximated with order $1/4$. For general extended MFC problems, we further show that, under sufficient regularity of the value functions, the value functions converge with an improved first-order rate, matching the best-known rate for classical control problems without mean field interaction, and consistent with the numerical observations for MFC of Cucker-Smale models.
Figures
Forward citations
Cited by 2 Pith papers
-
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies
Model-free deterministic policy gradients and a continuous-time deep actor-critic algorithm solve extended mean-field control problems whose dynamics and rewards depend on the joint state-control law.
-
NeuralChaos: Optimal Adapted Approximation of Square Integrable Predictable Processes
A finite-sampling neural architecture is dense in the Hilbert space of square-integrable predictable processes and attains best-N-term chaoslet rates for compressible or Malliavin-regular processes.
Reference graph
Works this paper leans on
-
[16]
Improved order 1 /4 convergence of Krylov’s piecewise constant policy approximation
E. R. Jakobsen, A. Picarelli, and C. Reisinger. “Improved order 1 /4 convergence of Krylov’s piecewise constant policy approximation”. Electron. Comm. Probab. 24 (2019), pp. 1–10
work page 2019
-
[30]
Backward Stochastic Differential Equations: From Linear to Fully Nonlinear Theory
J. Zhang. “Backward Stochastic Differential Equations: From Linear to Fully Nonlinear Theory”. Springer, New York 86 (2017). 38
work page 2017
-
[1]
Extended mean field control prob- lems: stochastic maximum principle and transport perspective
B. Acciaio, J. Backhoff-Veraguas, and R. Carmona. “Extended mean field control prob- lems: stochastic maximum principle and transport perspective”. SIAM J. Control Optim. 57 (2019), pp. 3666–3693
work page 2019
-
[2]
M. Basei and H. Pham. “Linear-quadratic McKean-Vlasov stochastic control problems with random coefficients on finite and infinite horizon, and applications”. arXiv preprint arXiv:1711.09390 (2017)
work page Pith review arXiv 2017
-
[3]
Well-posedness of mean-field type forward-backward stochastic differential equations
A. Bensoussan, S. Yam, and Z. Zhang. “Well-posedness of mean-field type forward-backward stochastic differential equations”. Stochastic Process. Appl. 125.9 (2015), pp. 3327–3354
work page 2015
-
[4]
Viscosity solutions for controlled McKean-Vlasov jump-diffusions
M. Burzoni, V. Ignazio, H. Soner, and A. M. Reppen. “Viscosity solutions for controlled McKean-Vlasov jump-diffusions”. SIAM J. Control Optim. 58.3 (2020), pp. 1676–1699
work page 2020
-
[5]
Forward-backward stochastic differential equations and con- trolled McKean-Vlasov dynamics
R. Carmona and F. Delarue. “Forward-backward stochastic differential equations and con- trolled McKean-Vlasov dynamics”. Ann. Probab. 43.5 (2015), pp. 2647–2700
work page 2015
-
[6]
Probabilistic theory of mean field games with applications I: Mean-field FBSDEs, control, and games
R. Carmona and F. Delarue. “Probabilistic theory of mean field games with applications I: Mean-field FBSDEs, control, and games”. Probability Theory and Stochastic Modelling, Springer 83 (2018)
work page 2018
Show all 30 references
-
[7]
Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: II-The finite horizon case
R. Carmona and M. Lauri` ere. “Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: II-The finite horizon case”. Ann. Appl. Probab. 32.6 (2022), pp. 4065–4105
2022
-
[8]
A probabilistic approach to classical solutions of the master equation for large population equilibria
J. F. Chassagneux, D. Crisan, and F. Delarue. “A probabilistic approach to classical solutions of the master equation for large population equilibria”. Mem. Amer. Math. Soc. 280.1379 (2022)
2022
-
[9]
Master Bellman equation in the Wasserstein space: Uniqueness of viscosity solutions
A. Cosso, F. Gozzi, I. Kharroubi, H. Pham, and M. Rosestolato. “Master Bellman equation in the Wasserstein space: Uniqueness of viscosity solutions”. Trans. Amer. Math. Soc. 377.1 (2024), pp. 31–83
2024
-
[10]
Emergent behavior in flocks
F. Cucker and S. Smale. “Emergent behavior in flocks”. IEEE Trans. Automat. Control 52.5 (2007), pp. 852–862
2007
-
[11]
Extended mean field control problem: a propagation of chaos result
M. F. Djete. “Extended mean field control problem: a propagation of chaos result”. Electron. J. Probab. 27 (2022), pp. 1–53
2022
-
[12]
McKean-Vlasov optimal control: the dynamic programming principle
M. F. Djete, D. Possamai, and X. Tan. “McKean-Vlasov optimal control: the dynamic programming principle”. Ann. Probab. 50.2 (2022), pp. 791–833
2022
-
[13]
Extended McKean-Vlasov optimal stochastic control applied to smart grid management
E. Gobet and M. Grangereau. “Extended McKean-Vlasov optimal stochastic control applied to smart grid management”. ESAIM Contr. Op. Ca. Va. 28 (2022), p. 40
2022
-
[14]
Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls
X. Guo, A. Hu, and Y. Zhang. “Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls”. SIAM J. Control and Optim. 61.2 (2023), pp. 755–787
2023
-
[15]
Itˆ o’s formula for flows of measures on semimartingales
X. Guo, H. Pham, and X. Wei. “Itˆ o’s formula for flows of measures on semimartingales”. Stochastic Process. Appl. 159 (2023), pp. 350–390
2023
-
[17]
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba. “Adam: A Method for Stochastic Optimization”. arXiv preprint arXiv:1412.6980 (2014). 37
2014 arXiv
-
[18]
Approximating value functions for controlled degenerate diffusion processes by using piece-wise constant policies
N. V. Krylov. “Approximating value functions for controlled degenerate diffusion processes by using piece-wise constant policies”. Electron. J. Probab. 4 (1999), pp. 1–19
1999
-
[19]
Convergence of large population games to mean field games with interaction through the controls
M. Lauri` ere and L. Tangpi. “Convergence of large population games to mean field games with interaction through the controls”. SIAM J. Math. Anal. 54.3 (2022), pp. 3535–3574
2022
-
[20]
On the convergence rate for piece-wise constant policy approximation of stochastic optimal control problem
T. Legrand. “On the convergence rate for piece-wise constant policy approximation of stochastic optimal control problem”. MA thesis. NTNU, 2023
2023
-
[21]
Mean field analysis of controlled Cucker- Smale type flocking: Linear analysis and perturbation equations
M. Nourian, P. E. Caines, and R. P. Malham´ e. “Mean field analysis of controlled Cucker- Smale type flocking: Linear analysis and perturbation equations”. IF AC Proc. Vol. 44.1 (2011), pp. 4471–4476
2011
-
[22]
Mean-field neural networks-based algorithms for McKean-Vlasov control problems
H. Pham and X. Warin. “Mean-field neural networks-based algorithms for McKean-Vlasov control problems”. arXiv preprint arXiv:2212.11518 (2022)
2022 arXiv
-
[23]
Bellman equation and viscosity solutions for mean-field stochastic control problem
H. Pham and X. Wei. “Bellman equation and viscosity solutions for mean-field stochastic control problem”. ESAIM Contr. Op. Ca. Va. 24 (2018), pp. 437–461
2018
-
[24]
Extended mean field control: a finite-dimensional numerical approximation
A. Picarelli, M. Scaratti, and J. Tam. “Extended mean field control: a finite-dimensional numerical approximation”. arXiv preprint arXiv:2503.20510 (2025)
2025
-
[25]
Freidlin-Wentzell LDP in path space for McKean- Vlasov equations and the functional iterated logarithm law
G. dos Reis, W. Salkeld, and J. Tugaut. “Freidlin-Wentzell LDP in path space for McKean- Vlasov equations and the functional iterated logarithm law”.Ann. Appl. Probab. 29.3 (2019), pp. 1487–1540
2019
-
[26]
A fast iterative PDE-based algorithm for feedback controls of nonsmooth mean-field control problems
C. Reisinger, W. Stockinger, and Y. Zhang. “A fast iterative PDE-based algorithm for feedback controls of nonsmooth mean-field control problems”. SIAM J. Sci. Comput. 46.4 (2024), A2737–A2773
2024
-
[27]
Exploration-exploitation trade-off for continuous- time episodic reinforcement learning with linear-convex models
L. Szpruch, T. Treetanthiploet, and Y. Zhang. “Exploration-exploitation trade-off for continuous- time episodic reinforcement learning with linear-convex models”.arXiv preprint arXiv:2112.10264 (2021)
2021 arXiv
-
[28]
Optimal Transport: Old and New
C. Villani. “Optimal Transport: Old and New”. Springer-Verlag, Berlin (2009)
2009
-
[29]
A linear-quadratic optimal control problem for mean-field stochastic differential equations
J. Yong. “A linear-quadratic optimal control problem for mean-field stochastic differential equations”. SIAM J. Control Optim. 51.4 (2013), pp. 2809–2838
2013
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.