REVIEW 5 minor 31 references
Donsker-Type Theorem for BSDEs: Rate of Convergence
T0 review · 0 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Random-walk driven BSDEs approximate Brownian BSDEs in Wasserstein distance at rate $n^{-(\alpha\wedge\varepsilon/2)}$.
desk verdict The paper delivers the conjectured optimal Donsker rate n^{-(α∧ε/2)} for BSDEs in Wasserstein distance, and the proof survives a close read; the remaining issues are typos and density, not substance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery consists of the value function $u$ of the semilinear heat equation, its discrete finite-difference analogue $U_n$, the discrete gradient $\Delta_n(t,x)=h^{-1/2}(U_n(t+h,x+\sqrt{h})-U_n(t+h,x-\sqrt{h}))/2$, and the representation formulas that write $u$ and $\nabla u$ as expectations against Brownian increments and $U_n$ and $\Delta_n$ as expectations against random-walk increments. The rate transfer relies on Lemma 6, which states that $\nabla u(t,\cdot)$ is $\varepsilon$-Hölder with constant $C(T-t)^{-1/2}$ even though $f$ is only Hölder in space, and on the Wasserstein central-limit bound $W_r(B^{n,t,x}_s,B^{t,x}_s)\leq C n^{-1/2}$ for the scaled random walk. A Gronwall/Volterra argument then closes the estimate for $\Delta_n$.
What would settle it
Set $T=1$, take $f=0$ and $g(x)=|x|^\varepsilon$ with $\varepsilon\in(0,1)$, and compute the random-walk scheme exactly for $n=2^m$. At a fixed time $s<1$, measure the Wasserstein-2 distance between $Z^{n,0,0}_s$ and $\nabla u(s,B_s)$ for the heat-equation value function $u$; the theorem predicts decay like $n^{-\varepsilon/2}$, so a slower observed exponent disproves it. One can also check Lemma 6 directly by testing whether $|\nabla u(t,x)-\nabla u(t,y)|\le C(T-t)^{-1/2}|x-y|^\varepsilon$ holds for a sequence $t\uparrow T$.
Extended reading notes
Core claim
The central claim, on the paper's own terms, is that the Wasserstein distance between the random-walk-driven BSDE solution and the Brownian BSDE solution decays at the rate $n^{-(\alpha\wedge\varepsilon/2)}$, with the $Z$ component carrying the natural factor $(T-s)^{-1/2}$ that diverges near the terminal time. The proof works by reducing the stochastic approximation to a deterministic comparison: the random-walk BSDE is represented through a finite-difference value function $U_n$ and discrete gradient $\Delta_n$, while the Brownian BSDE is represented through the solution $u$ of the associated semilinear heat equation. The paper proves the pointwise estimates $|u-U_n|\leq C(1+|x|)^\varepsilon n^{-(\alpha\wedge\varepsilon/2)}$ and $|\nabla u-\Delta_n|\leq C(1+|x|)^\varepsilon (T-t)^{-1/2}n^{-(\alpha\wedge\varepsilon/2)}$, and then transfers these bounds to Wasserstein distance using the Hölder regularity of $u$ and $\nabla u$.
Load-bearing premise
The argument stands on a precise regularity estimate: the spatial derivative of the PDE solution must be Hölder continuous with a constant proportional to $(T-t)^{-1/2}$, and if that constant blew up any faster the integral estimates for the $Z$ error would diverge and the rate would fail.
Editorial extensions
If this is right
- For terminal data with $\varepsilon=1$ and time regularity $\alpha\ge 1/2$, the approximation attains the $n^{-1/2}$ Wasserstein rate, matching the random-walk CLT; the discretization does not slow down the weak error.
- The $Z$ error must be expected to grow like $(T-s)^{-1/2}$ near $T$, so any practical simulation based on this scheme needs to handle that singularity.
- The pointwise finite-difference rates for $U_n$ and $\Delta_n$ provide a standalone statement about semilinear heat equations with Hölder data.
- The Wasserstein formulation is essential: the same quantities measured in strong $L^p$ norms would not show this clean rate, because the $Z$ processes are not close pathwise.
Reading between the lines
- The rate $n^{-\varepsilon/2}$ suggests that the dominant error is the weak approximation of Brownian motion by the random walk, with the nonlinearity contributing only through the Hölder modulus; one would expect the same rate for any CLT-scaled Markov chain satisfying a Wasserstein CLT bound with rate $n^{-1/2}$.
- The paper leaves implicit that the same proof should extend to multidimensional Brownian motion and random walks with i.i.d. increments with exponential moments, provided the Wasserstein CLT estimate and the gradient regularity lemma have multidimensional analogues.
- When the data are smoother than Hölder, the bound saturates at $n^{-1/2}$, so further improvement would require a higher-order weak scheme rather than the plain random walk; this is consistent with the known behaviour of binomial tree approximations.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proves a quantitative Donsker-type theorem for the Briand-Delyon-Mémin random-walk approximation of Markovian BSDEs. Under Assumption (A1), where the terminal condition g is ε-Hölder and the generator f is α-Hölder in time, ε-Hölder in space, and Lipschitz in (y,z), the authors show that the random-walk-driven solution (Y^{n,t,x}, Z^{n,t,x}) converges in L^r-Wasserstein distance to the Brownian BSDE solution (Y^{t,x}, Z^{t,x}) at rate n^{-(α∧ε/2)} for Y, and with the additional factor (T-s)^{-1/2} for Z (Theorem 10). The proof uses Rio's Wasserstein central limit theorem, the Feynman-Kac/PDE representation, new regularity estimates for u and ∇u (Lemma 6, proved via inf-convolution in Appendix A.3), discrete BSDE a priori estimates (Lemma 12), and a two-stage Gronwall argument in Proposition 11.
Significance. If the result holds, it is a substantial improvement over the previous rate n^{-ε/4} obtained by Geiss-Labart-Luoto, and it matches the expected optimal rate n^{-ε/2}; for ε=1 and α≥1/2 it recovers the n^{-1/2} rate of Rio's Wasserstein CLT. The Wasserstein approach is well chosen for this problem and gives clean statements with explicit singular behavior of the Z-error near the terminal time. The proof is structurally self-contained: the essential inputs are Rio's theorem, Zhang's representation, and classical BSDE a priori estimates, and no constants are fitted or ad hoc. The main regularity lemma is delicate, but the inf-convolution argument in the appendix is a reasonable strategy and the overall chain of estimates is internally consistent.
minor comments (5)
- [Section 4, Eq. (30)] The exponent in the display should be −(1−ε)/2 rather than (1−ε)/2. The subsequent applications in Lemma 6(bii) and Proposition 9 consistently use the negative-exponent form, so I regard this as a typographical slip, but the displayed inequality as written is false and should be corrected.
- [Appendix A.3, Step 2] The claim that the inf-convolution fη satisfies (3) uniformly in η is not immediate for the ε-Hölder constant in x. The uniformity follows by combining the η-Lipschitz bound with the L∞-approximation error in (46); I recommend adding a short justification or a reference, since this uniformity is used to make the constants in (48) independent of η.
- [Section 5, proof of Theorem 10, grid-point case] The identity Z^{n,t,x}_s = Δ_n(s−h/2, B^{n,t,x}_{s−h/2}) relies on the convention that U_n is evaluated at the lower grid point; a brief reminder of this convention in the display would prevent confusion, since without it the time shift looks off by h/2.
- [Appendix A.3, use of Lemma 14] The inequality in (50) has the weakly singular kernel (s−r)^{−ε/2}, whereas Lemma 14 is stated for the kernel (s−r)^{−1/2}. The conclusion used in the text follows from a standard generalized Volterra Gronwall inequality, but the paper should either state that version or explain why Lemma 14 applies; as written this step is not immediate.
- [Section 3, Proposition 3] The triangle-inequality step in the proof of Proposition 3 is hard to parse because the overline/underline notation for grid endpoints is not visible in the display; please ensure that the Brownian increments B_{\underline{t}} and B_{\overline{s}} are clearly distinguished from B_t and B_s.
Circularity Check
No significant circularity: the rate n^{-(α∧ε/2)} is derived from independent external inputs (Rio's Wasserstein bound, Zhang's representation, El Karoui et al. a priori estimates) plus a self-contained PDE regularity argument.
full rationale
The paper's central result, Theorem 10, is not equivalent to any input by construction. The Wasserstein rate for the scaled random walk is taken from Rio [15], which is independent of the authors, and the representation formulas for u and ∇u are taken from Zhang [19], also independent. The only regularity input beyond the Lipschitz case, Lemma 6(b), is proved in Appendix A.3 by inf-convolution: the authors regularize f to a Lipschitz generator f_η, apply the already-proved Step 1 (based on [19]) to f_η, and then pass to the limit using (47) and dominated convergence. No fitted constant appears, and no target rate is assumed. The authors' own earlier papers [9,10] are cited only to state the previously known slower rate n^{-ε/4}, not as a component of the proof; [5] and [6] are used for existence, uniqueness, and a discrete a priori estimate that is generalized and proved in Lemma 12; [11] is a standard BSDE estimate. One displayed exponent in (30) has the wrong sign (it should be −(1−ε)/2), but every subsequent application uses the negative exponent, so this is a typo rather than a circular dependence. Thus there is no circular step to report.
Assumptions & free parameters
assumptions (7)
- standard math Existence and uniqueness of L^p solutions to the BSDE (2) under Assumption (A1), cited from [4, Theorem 4.2].
- standard math Rio's theorem: W_ψ(n^{-1/2} S_n, G) ≤ C n^{-1/2} for i.i.d. increments with exponential moments ([15, Theorem 2.1]).
- standard math Ma-Zhang representation theorem and Zhang's formulas (6) and (7), expressing u and ∇u as expectations of g and F evaluated along Brownian bridges ([13, 19]).
- standard math Zhang's regularity: u ∈ C^{0,1}([0,T[×R) and Z = ∇u for Lipschitz generators ([19, Theorem 3.2]).
- standard math Classical a priori estimates for BSDEs ([8, Proposition 2.1]) and their discrete analogues ([6, Proposition 7]).
- standard math Volterra Gronwall lemma for singular kernels ([12, Exercise 4, page 190]).
- standard math The inf-convolution f_η is η-Lipschitz in x and satisfies |f_η - f| ≤ c η^{-ε/(1-ε)} under the Hölder condition (3).
Cite this review
Pith. "Pith review of Donsker-Type Theorem for BSDEs: Rate of Convergence." pith.science (2026). https://pith.science/paper/RHUJLO67
@misc{pith2026190801188,
author = {Pith},
title = {Pith review of: Donsker-Type Theorem for BSDEs: Rate of Convergence},
year = {2026},
howpublished = {\url{https://pith.science/paper/RHUJLO67}},
note = {Machine review of arXiv:1908.01188}
}
read the original abstract
In this paper, we study in the Markovian case the rate of convergence in the Wasserstein distance of an approximation of the solution to a BSDE given by a BSDE which is driven by a scaled random walk as introduced in Briand, Delyon and M{\'e}min (Electron. Comm. Probab. 6(2001),1-14).
Reference graph
Works this paper leans on
-
[1]
Introduction. In this paper, we are concerned with the discretization of so lutions to BSDEs of the form Yt =G(B) + ∫ T t f (Bs,Y s,Z s)ds − ∫ T t ZsdBs, 0 ≤t ≤T, where B is a standard Brownian motion. These equations have been int roduced by Jean- Michel Bismut for linear generators in [ 2] and by Étienne Pardoux and Shige Peng for Lipschitz generators i...
-
[2]
In all the sequel, T > 0 is a fixed positive real number
Notation. In all the sequel, T > 0 is a fixed positive real number. We work on a complete probabi lity space (Ω , F, P) carrying a standard real Brownian motion {Bt}0≤t≤T, and {Ft}0≤t≤T stands for the augmented filtration of B which is right continuous and complete. 2 We consider the following BSDE Yt =g(BT ) + ∫ T t f (s,B s,Y s,Z s)ds − ∫ T t ZsdBs, 0 ≤t ...
-
[3]
Scaled random walk and Wasserstein distance. One starting point of our paper is the following result of Emm anuel Rio [ 15] (Theorem 2.1); see also [ 16]. This result covers, up to a generalization, the case where the generator vanishes, i.e. f ≡ 0. Letψ be the convex function defined by ψ (x) = e|x| − 1. The Orlicz norm associated to this function ψ of an...
-
[4]
Regularity results on u, Un, ∇u and ∆ n. Let us start by known regularity properties of the function u that follow from classical a priori estimates for BSDEs. Lemma 5. Under Assumption (A1) there exists a constant C >0 depending on (T,ε,f,g ) such that, for all (t,x ) ∈ [0,T ] × R, |u(t,x )| ≤ C (1 + |x|)ε, ‖u(t, ·)‖ε ≤C, ‖u(·,x )‖ε/ 2 ≤C (1 + |x|)ε. Pro...
-
[5]
Main results. In this section, we state the main result of this paper which g ives the rate of convergence in the Wasserstein distance between the solution to the BSDE (
-
[6]
, On the robustness of backward stochastic differential equat ions, Stochastic Process. Appl. 97 (2002), no. 2, 229–253
work page 2002
-
[7]
For the following we want to remind the reader of Remark 7
and the solution to the BSDE driven by the scaled random walk ( 12). For the following we want to remind the reader of Remark 7. Theorem 10. Under (A1), for any r ∈ [1, ∞[, there exists a constant Cr > 0 depending at most on (T,α,ε,f,g,r ) such that for all x ∈ R, (i) Wr ( Yn,t,x s ,Y t,x s ) ≤Cr (1 + |x|)ε n−(α ∧ ε 2 ) for all 0 ≤t ≤s ≤T, (ii) Wr ( Zn,t,...
-
[8]
N. El Karoui, S. Peng, and M.C. Quenez, Backward Stochastic Differential Equations in Finance, Math. Finance 7 (1997), 1–71
work page 1997
Show all 31 references
-
[9]
Geiss, C
C. Geiss, C. Labart, and A. Luoto, L2-Approximation rate of forward -backward SDEs using random walk, arXiv:1807.05889 (2018)
2018 arXiv
-
[10]
, Random walk approximation of BSDEs with Hölder continuous t erminal condition, arXiv:1806.07674, to appear in Bernoulli, (2019)
2019 arXiv
-
[11]
Geiss and J
S. Geiss and J. Ylinen, Decoupling on the Wiener Space, Related Besov Spaces, and Ap pli- cations to BSDEs. arXiv:1409.5322, to appear Memoirs AMS, (2019)
2019 arXiv
-
[12]
where (g,f ) is replaced by (¯g, ¯f ). Proof. Let n be such that T ‖f ‖ Lip/n< 1 and T ‖ ¯f ‖ Lip/n< 1. Since, ⟨Bn⟩t − ⟨Bn⟩s ≤ (t −s) + T/n , doing exactly the same computation as in the proof of Propos ition 7 in [ 6], we get, for a universal constant c ≥ 1, E [ sup σ ≤s≤τ |δ...
-
[13]
In view of the regularity of f in time, we have ⏐ ⏐ ⏐ ⏐ ⏐E [ ∫ T t ( f ( s, Θ n,t,x s ) −f ( s, Θ n,t,x s )) ds ] ⏐ ⏐ ⏐ ⏐ ⏐ ≤Cn −α
and ( 16), E [ ∫ T t f ( s, Θ n,t,x s ) ds ] = E [ ∫ T t ( f ( s, Θ n,t,x s ) −f ( s, Θ n,t,x s )) ] + E [ ∫ T t Fn ( s,B n,t,x s ) ds ] + E [ ∫ t t f ( s,B n,t,x s ,Y n,t,x s ,Z n,t,x s ) ds ] . In view of the regularity of f in time, we have ⏐ ⏐ ⏐ ⏐ ⏐E [ ∫ T t ( f ( s, Θ n,t...
-
[14]
Ambrosio, N
L. Ambrosio, N. Gigli and G. Savaré, Gradient Flows in Metric Spaces and in the Space of Probability Measures, Birkhäuser, Basel, Boston, Berlin (2005)
2005
-
[15]
Bismut, Théorie probabiliste du contrôle des diffusions , Mem
J.M. Bismut, Théorie probabiliste du contrôle des diffusions , Mem. AMS 176 (1973)
1973
-
[16]
Bouchard and N
B. Bouchard and N. Touzi, Discrete-time approximation and Monte-Carlo simulation o f backward stochastic differential equations , Stochastic Process. Appl. 111 (2004), no. 2, 175–206
2004
-
[17]
Briand, B
P. Briand, B. Delyon, Y. Hu, E. Pardoux and L. Stoica, Lp solutions of backward stochastic differential equations , Stochastic Process. Appl. 108 (2003) , 109–129
2003
-
[18]
With this notation in hand, we have, taking into account (
we also have that Fn(s,B n,t,x s ) = f (s, Θ n,t,x s ). With this notation in hand, we have, taking into account (
-
[19]
Briand, B
Ph. Briand, B. Delyon, and J. Mémin, Donsker–type theorem for BSDEs , Electron. Comm. Probab. 6 (2001), 1–14, (electronic)
2001
-
[20]
Jacod and A
J. Jacod and A. N. Shiryaev, Limit theorems for stochastic processes , Springer (2003)
2003
-
[22]
Choosing f (x) = x in ( 21), this implies the first result
for r = 1, W1 ( g ( Bn,t,x s ) ,g ( Bt,x s )) ≤ ‖g‖ε W1 ( Bn,t,x s ,B t,x s ) ε ≤ ‖g‖ε cε 1 ( T n ) ε/ 2 . Choosing f (x) = x in ( 21), this implies the first result. Let us prove the second assertion. We start by observing that , since Bs −Bt and Bn s −Bn t are centered random...
-
[23]
Henry, Geometric theory of semilinear parabolic equations, Lecture Notes in Mathematics 840, Springer, 1981
D. Henry, Geometric theory of semilinear parabolic equations, Lecture Notes in Mathematics 840, Springer, 1981
1981
-
[24]
Ma and J
J. Ma and J. Zhang, Representation theorems for backward stochastic different ial equations, Ann. Appl. Probab. 12 (2002), no. 4, 1390–1418
2002
-
[25]
Pardoux and S
É. Pardoux and S. Peng, Adapted solution of a backward stochastic differential equa tion, Systems Control Lett. 14 (1990), no. 1, 55–61
1990
-
[26]
Rio, Upper bounds for minimal distances in the central limit theo rem, Ann
E. Rio, Upper bounds for minimal distances in the central limit theo rem, Ann. Inst. Henri Poincaré Probab. Stat. 45 (2009), no. 3, 802–817
2009
-
[27]
, Asymptotic constants for minimal distance in the central li mit theorem, Electron. Commun. Probab. 16 (2011), 96–103
2011
-
[28]
Walsh, The rate of convergence of the binomial tree scheme , Finance Stoch
John B. Walsh, The rate of convergence of the binomial tree scheme , Finance Stoch. 7 (2003), no. 3, 337–361
2003
-
[29]
Zhang, A numerical scheme for BSDEs , Ann
J. Zhang, A numerical scheme for BSDEs , Ann. Appl. Probab. 14 (2004), no. 1, 459–488
2004
-
[30]
, Representation of solutions to BSDEs associated with a dege nerate FSDE , Ann. Appl. Probab. 15 (2005), no. 3, 1798–1831. 28
2005
-
[31]
19 Since (t −t) = (t −t) ε 2 (t −t)1− ε 2 ≤h ε 2 (s −t)1− ε 2 and the same upper bound holds for s −s (since s −s ≤h ≤s −t), we get H1(s) ≤ C√ T −s h ε 2 (s −t)(1−ε )/ 2(s −t) ε 2
of F gives H1(s) ≤C (s −t)(1+ε )/ 2 √ T −s (t −t) + (s −s) (s −t)(s −t) = C√ T −s (t −t) + (s −s) (s −t)(1−ε )/ 2(s −t). 19 Since (t −t) = (t −t) ε 2 (t −t)1− ε 2 ≤h ε 2 (s −t)1− ε 2 and the same upper bound holds for s −s (since s −s ≤h ≤s −t), we get H1(s) ≤ C√ T −s h ε 2 (s...
-
[32]
, Backward Stochastic Differential Equations , Springer, New York (2017). 29
2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.