Pith. sign in

REVIEW 5 minor 31 references

Donsker-Type Theorem for BSDEs: Rate of Convergence

T0 review · 0 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Random-walk driven BSDEs approximate Brownian BSDEs in Wasserstein distance at rate $n^{-(\alpha\wedge\varepsilon/2)}$.

desk verdict The paper delivers the conjectured optimal Donsker rate n^{-(α∧ε/2)} for BSDEs in Wasserstein distance, and the proof survives a close read; the remaining issues are typos and density, not substance. read the letter →

arxiv 1908.01188 v1 pith:RHUJLO67 submitted 2019-08-03 math.PR

classification math.PR MSC 60H1060H3560F0565C30
keywords WassersteindistancebackwardstochasticdifferentialequationsrandomwalkapproximationDonskertheoremrateofconvergencesemilinearheatequationHölderregularityfinite-differencescheme
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves a rate of convergence in Wasserstein distance for approximating a Markovian backward stochastic differential equation (BSDE) by a BSDE driven by a scaled random walk. Under Assumption (A1), where the terminal condition $g$ is $\varepsilon$-Hölder and the generator $f$ is $\alpha$-Hölder in time and $\varepsilon$-Hölder in space, the paper establishes $W_r(Y^{n,t,x}_s,Y^{t,x}_s)\leq C_r(1+|x|)^\varepsilon n^{-(\alpha\wedge\varepsilon/2)}$ and $W_r(Z^{n,t,x}_s,Z^{t,x}_s)\leq C_r(1+|x|)^\varepsilon (T-s)^{-1/2}n^{-(\alpha\wedge\varepsilon/2)}$. This confirms the expected improvement over the previous rate $n^{-\varepsilon/4}$ and, for smooth data, reaches the $n^{-1/2}$ rate of the classical random-walk central limit theorem. The result matters because it identifies the exact weak error of a simple Donsker-type discretization of BSDEs with irregular data, which is the relevant regime for Monte Carlo simulation.

What carries the argument

The load-bearing machinery consists of the value function $u$ of the semilinear heat equation, its discrete finite-difference analogue $U_n$, the discrete gradient $\Delta_n(t,x)=h^{-1/2}(U_n(t+h,x+\sqrt{h})-U_n(t+h,x-\sqrt{h}))/2$, and the representation formulas that write $u$ and $\nabla u$ as expectations against Brownian increments and $U_n$ and $\Delta_n$ as expectations against random-walk increments. The rate transfer relies on Lemma 6, which states that $\nabla u(t,\cdot)$ is $\varepsilon$-Hölder with constant $C(T-t)^{-1/2}$ even though $f$ is only Hölder in space, and on the Wasserstein central-limit bound $W_r(B^{n,t,x}_s,B^{t,x}_s)\leq C n^{-1/2}$ for the scaled random walk. A Gronwall/Volterra argument then closes the estimate for $\Delta_n$.

What would settle it

Set $T=1$, take $f=0$ and $g(x)=|x|^\varepsilon$ with $\varepsilon\in(0,1)$, and compute the random-walk scheme exactly for $n=2^m$. At a fixed time $s<1$, measure the Wasserstein-2 distance between $Z^{n,0,0}_s$ and $\nabla u(s,B_s)$ for the heat-equation value function $u$; the theorem predicts decay like $n^{-\varepsilon/2}$, so a slower observed exponent disproves it. One can also check Lemma 6 directly by testing whether $|\nabla u(t,x)-\nabla u(t,y)|\le C(T-t)^{-1/2}|x-y|^\varepsilon$ holds for a sequence $t\uparrow T$.

Watch

Extended reading notes

Core claim

The central claim, on the paper's own terms, is that the Wasserstein distance between the random-walk-driven BSDE solution and the Brownian BSDE solution decays at the rate $n^{-(\alpha\wedge\varepsilon/2)}$, with the $Z$ component carrying the natural factor $(T-s)^{-1/2}$ that diverges near the terminal time. The proof works by reducing the stochastic approximation to a deterministic comparison: the random-walk BSDE is represented through a finite-difference value function $U_n$ and discrete gradient $\Delta_n$, while the Brownian BSDE is represented through the solution $u$ of the associated semilinear heat equation. The paper proves the pointwise estimates $|u-U_n|\leq C(1+|x|)^\varepsilon n^{-(\alpha\wedge\varepsilon/2)}$ and $|\nabla u-\Delta_n|\leq C(1+|x|)^\varepsilon (T-t)^{-1/2}n^{-(\alpha\wedge\varepsilon/2)}$, and then transfers these bounds to Wasserstein distance using the Hölder regularity of $u$ and $\nabla u$.

Load-bearing premise

The argument stands on a precise regularity estimate: the spatial derivative of the PDE solution must be Hölder continuous with a constant proportional to $(T-t)^{-1/2}$, and if that constant blew up any faster the integral estimates for the $Z$ error would diverge and the rate would fail.

Editorial extensions

If this is right

  • For terminal data with $\varepsilon=1$ and time regularity $\alpha\ge 1/2$, the approximation attains the $n^{-1/2}$ Wasserstein rate, matching the random-walk CLT; the discretization does not slow down the weak error.
  • The $Z$ error must be expected to grow like $(T-s)^{-1/2}$ near $T$, so any practical simulation based on this scheme needs to handle that singularity.
  • The pointwise finite-difference rates for $U_n$ and $\Delta_n$ provide a standalone statement about semilinear heat equations with Hölder data.
  • The Wasserstein formulation is essential: the same quantities measured in strong $L^p$ norms would not show this clean rate, because the $Z$ processes are not close pathwise.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The rate $n^{-\varepsilon/2}$ suggests that the dominant error is the weak approximation of Brownian motion by the random walk, with the nonlinearity contributing only through the Hölder modulus; one would expect the same rate for any CLT-scaled Markov chain satisfying a Wasserstein CLT bound with rate $n^{-1/2}$.
  • The paper leaves implicit that the same proof should extend to multidimensional Brownian motion and random walks with i.i.d. increments with exponential moments, provided the Wasserstein CLT estimate and the gradient regularity lemma have multidimensional analogues.
  • When the data are smoother than Hölder, the bound saturates at $n^{-1/2}$, so further improvement would require a higher-order weak scheme rather than the plain random walk; this is consistent with the known behaviour of binomial tree approximations.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper proves a quantitative Donsker-type theorem for the Briand-Delyon-Mémin random-walk approximation of Markovian BSDEs. Under Assumption (A1), where the terminal condition g is ε-Hölder and the generator f is α-Hölder in time, ε-Hölder in space, and Lipschitz in (y,z), the authors show that the random-walk-driven solution (Y^{n,t,x}, Z^{n,t,x}) converges in L^r-Wasserstein distance to the Brownian BSDE solution (Y^{t,x}, Z^{t,x}) at rate n^{-(α∧ε/2)} for Y, and with the additional factor (T-s)^{-1/2} for Z (Theorem 10). The proof uses Rio's Wasserstein central limit theorem, the Feynman-Kac/PDE representation, new regularity estimates for u and ∇u (Lemma 6, proved via inf-convolution in Appendix A.3), discrete BSDE a priori estimates (Lemma 12), and a two-stage Gronwall argument in Proposition 11.

Significance. If the result holds, it is a substantial improvement over the previous rate n^{-ε/4} obtained by Geiss-Labart-Luoto, and it matches the expected optimal rate n^{-ε/2}; for ε=1 and α≥1/2 it recovers the n^{-1/2} rate of Rio's Wasserstein CLT. The Wasserstein approach is well chosen for this problem and gives clean statements with explicit singular behavior of the Z-error near the terminal time. The proof is structurally self-contained: the essential inputs are Rio's theorem, Zhang's representation, and classical BSDE a priori estimates, and no constants are fitted or ad hoc. The main regularity lemma is delicate, but the inf-convolution argument in the appendix is a reasonable strategy and the overall chain of estimates is internally consistent.

minor comments (5)
  1. [Section 4, Eq. (30)] The exponent in the display should be −(1−ε)/2 rather than (1−ε)/2. The subsequent applications in Lemma 6(bii) and Proposition 9 consistently use the negative-exponent form, so I regard this as a typographical slip, but the displayed inequality as written is false and should be corrected.
  2. [Appendix A.3, Step 2] The claim that the inf-convolution fη satisfies (3) uniformly in η is not immediate for the ε-Hölder constant in x. The uniformity follows by combining the η-Lipschitz bound with the L∞-approximation error in (46); I recommend adding a short justification or a reference, since this uniformity is used to make the constants in (48) independent of η.
  3. [Section 5, proof of Theorem 10, grid-point case] The identity Z^{n,t,x}_s = Δ_n(s−h/2, B^{n,t,x}_{s−h/2}) relies on the convention that U_n is evaluated at the lower grid point; a brief reminder of this convention in the display would prevent confusion, since without it the time shift looks off by h/2.
  4. [Appendix A.3, use of Lemma 14] The inequality in (50) has the weakly singular kernel (s−r)^{−ε/2}, whereas Lemma 14 is stated for the kernel (s−r)^{−1/2}. The conclusion used in the text follows from a standard generalized Volterra Gronwall inequality, but the paper should either state that version or explain why Lemma 14 applies; as written this step is not immediate.
  5. [Section 3, Proposition 3] The triangle-inequality step in the proof of Proposition 3 is hard to parse because the overline/underline notation for grid endpoints is not visible in the display; please ensure that the Brownian increments B_{\underline{t}} and B_{\overline{s}} are clearly distinguished from B_t and B_s.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the rate n^{-(α∧ε/2)} is derived from independent external inputs (Rio's Wasserstein bound, Zhang's representation, El Karoui et al. a priori estimates) plus a self-contained PDE regularity argument.

full rationale

The paper's central result, Theorem 10, is not equivalent to any input by construction. The Wasserstein rate for the scaled random walk is taken from Rio [15], which is independent of the authors, and the representation formulas for u and ∇u are taken from Zhang [19], also independent. The only regularity input beyond the Lipschitz case, Lemma 6(b), is proved in Appendix A.3 by inf-convolution: the authors regularize f to a Lipschitz generator f_η, apply the already-proved Step 1 (based on [19]) to f_η, and then pass to the limit using (47) and dominated convergence. No fitted constant appears, and no target rate is assumed. The authors' own earlier papers [9,10] are cited only to state the previously known slower rate n^{-ε/4}, not as a component of the proof; [5] and [6] are used for existence, uniqueness, and a discrete a priori estimate that is generalized and proved in Lemma 12; [11] is a standard BSDE estimate. One displayed exponent in (30) has the wrong sign (it should be −(1−ε)/2), but every subsequent application uses the negative exponent, so this is a typo rather than a circular dependence. Thus there is no circular step to report.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

No free parameters are fitted. The proof rests on established theorems from the BSDE and probability literature, all cited and used as external input. No new entities are introduced.

assumptions (7)
  • standard math Existence and uniqueness of L^p solutions to the BSDE (2) under Assumption (A1), cited from [4, Theorem 4.2].
    Invoked in Section 2 to define (Y,Z) and the Markovian value function u.
  • standard math Rio's theorem: W_ψ(n^{-1/2} S_n, G) ≤ C n^{-1/2} for i.i.d. increments with exponential moments ([15, Theorem 2.1]).
    Starting point in Section 3; Proposition 3 extends it to Brownian increments.
  • standard math Ma-Zhang representation theorem and Zhang's formulas (6) and (7), expressing u and ∇u as expectations of g and F evaluated along Brownian bridges ([13, 19]).
    Core representation used throughout Proposition 11.
  • standard math Zhang's regularity: u ∈ C^{0,1}([0,T[×R) and Z = ∇u for Lipschitz generators ([19, Theorem 3.2]).
    Baseline for Lemma 6, which extends to ε-Hölder f.
  • standard math Classical a priori estimates for BSDEs ([8, Proposition 2.1]) and their discrete analogues ([6, Proposition 7]).
    Used in Lemmas 5 and 8 and in Appendix A.1 to bound the approximating processes.
  • standard math Volterra Gronwall lemma for singular kernels ([12, Exercise 4, page 190]).
    Used in Lemma 14 to close the integral inequality for γ_n.
  • standard math The inf-convolution f_η is η-Lipschitz in x and satisfies |f_η - f| ≤ c η^{-ε/(1-ε)} under the Hölder condition (3).
    Used in Appendix A.3 to prove Lemma 6 for ε-Hölder f.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Donsker-Type Theorem for BSDEs: Rate of Convergence." pith.science (2026). https://pith.science/paper/RHUJLO67

@misc{pith2026190801188,
  author       = {Pith},
  title        = {Pith review of: Donsker-Type Theorem for BSDEs: Rate of Convergence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RHUJLO67}},
  note         = {Machine review of arXiv:1908.01188}
}
read the original abstract

In this paper, we study in the Markovian case the rate of convergence in the Wasserstein distance of an approximation of the solution to a BSDE given by a BSDE which is driven by a scaled random walk as introduced in Briand, Delyon and M{\'e}min (Electron. Comm. Probab. 6(2001),1-14).

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

31 extracted references · 31 canonical work pages

  1. [1]

    Introduction. In this paper, we are concerned with the discretization of so lutions to BSDEs of the form Yt =G(B) + ∫ T t f (Bs,Y s,Z s)ds − ∫ T t ZsdBs, 0 ≤t ≤T, where B is a standard Brownian motion. These equations have been int roduced by Jean- Michel Bismut for linear generators in [ 2] and by Étienne Pardoux and Shige Peng for Lipschitz generators i...

  2. [2]

    In all the sequel, T > 0 is a fixed positive real number

    Notation. In all the sequel, T > 0 is a fixed positive real number. We work on a complete probabi lity space (Ω , F, P) carrying a standard real Brownian motion {Bt}0≤t≤T, and {Ft}0≤t≤T stands for the augmented filtration of B which is right continuous and complete. 2 We consider the following BSDE Yt =g(BT ) + ∫ T t f (s,B s,Y s,Z s)ds − ∫ T t ZsdBs, 0 ≤t ...

  3. [3]

    One starting point of our paper is the following result of Emm anuel Rio [ 15] (Theorem 2.1); see also [ 16]

    Scaled random walk and Wasserstein distance. One starting point of our paper is the following result of Emm anuel Rio [ 15] (Theorem 2.1); see also [ 16]. This result covers, up to a generalization, the case where the generator vanishes, i.e. f ≡ 0. Letψ be the convex function defined by ψ (x) = e|x| − 1. The Orlicz norm associated to this function ψ of an...

  4. [4]

    Let us start by known regularity properties of the function u that follow from classical a priori estimates for BSDEs

    Regularity results on u, Un, ∇u and ∆ n. Let us start by known regularity properties of the function u that follow from classical a priori estimates for BSDEs. Lemma 5. Under Assumption (A1) there exists a constant C >0 depending on (T,ε,f,g ) such that, for all (t,x ) ∈ [0,T ] × R, |u(t,x )| ≤ C (1 + |x|)ε, ‖u(t, ·)‖ε ≤C, ‖u(·,x )‖ε/ 2 ≤C (1 + |x|)ε. Pro...

  5. [5]

    In this section, we state the main result of this paper which g ives the rate of convergence in the Wasserstein distance between the solution to the BSDE (

    Main results. In this section, we state the main result of this paper which g ives the rate of convergence in the Wasserstein distance between the solution to the BSDE (

  6. [6]

    , On the robustness of backward stochastic differential equat ions, Stochastic Process. Appl. 97 (2002), no. 2, 229–253

  7. [7]

    For the following we want to remind the reader of Remark 7

    and the solution to the BSDE driven by the scaled random walk ( 12). For the following we want to remind the reader of Remark 7. Theorem 10. Under (A1), for any r ∈ [1, ∞[, there exists a constant Cr > 0 depending at most on (T,α,ε,f,g,r ) such that for all x ∈ R, (i) Wr ( Yn,t,x s ,Y t,x s ) ≤Cr (1 + |x|)ε n−(α ∧ ε 2 ) for all 0 ≤t ≤s ≤T, (ii) Wr ( Zn,t,...

  8. [8]

    El Karoui, S

    N. El Karoui, S. Peng, and M.C. Quenez, Backward Stochastic Differential Equations in Finance, Math. Finance 7 (1997), 1–71

Show all 31 references
  1. [9]

    Geiss, C

    C. Geiss, C. Labart, and A. Luoto, L2-Approximation rate of forward -backward SDEs using random walk, arXiv:1807.05889 (2018)

  2. [10]

    , Random walk approximation of BSDEs with Hölder continuous t erminal condition, arXiv:1806.07674, to appear in Bernoulli, (2019)

  3. [11]

    Geiss and J

    S. Geiss and J. Ylinen, Decoupling on the Wiener Space, Related Besov Spaces, and Ap pli- cations to BSDEs. arXiv:1409.5322, to appear Memoirs AMS, (2019)

  4. [12]

    where (g,f ) is replaced by (¯g, ¯f ). Proof. Let n be such that T ‖f ‖ Lip/n< 1 and T ‖ ¯f ‖ Lip/n< 1. Since, ⟨Bn⟩t − ⟨Bn⟩s ≤ (t −s) + T/n , doing exactly the same computation as in the proof of Propos ition 7 in [ 6], we get, for a universal constant c ≥ 1, E [ sup σ ≤s≤τ |δ...

  5. [13]

    In view of the regularity of f in time, we have ⏐ ⏐ ⏐ ⏐ ⏐E [ ∫ T t ( f ( s, Θ n,t,x s ) −f ( s, Θ n,t,x s )) ds ] ⏐ ⏐ ⏐ ⏐ ⏐ ≤Cn −α

    and ( 16), E [ ∫ T t f ( s, Θ n,t,x s ) ds ] = E [ ∫ T t ( f ( s, Θ n,t,x s ) −f ( s, Θ n,t,x s )) ] + E [ ∫ T t Fn ( s,B n,t,x s ) ds ] + E [ ∫ t t f ( s,B n,t,x s ,Y n,t,x s ,Z n,t,x s ) ds ] . In view of the regularity of f in time, we have ⏐ ⏐ ⏐ ⏐ ⏐E [ ∫ T t ( f ( s, Θ n,t...

  6. [14]

    Ambrosio, N

    L. Ambrosio, N. Gigli and G. Savaré, Gradient Flows in Metric Spaces and in the Space of Probability Measures, Birkhäuser, Basel, Boston, Berlin (2005)

  7. [15]

    Bismut, Théorie probabiliste du contrôle des diffusions , Mem

    J.M. Bismut, Théorie probabiliste du contrôle des diffusions , Mem. AMS 176 (1973)

  8. [16]

    Bouchard and N

    B. Bouchard and N. Touzi, Discrete-time approximation and Monte-Carlo simulation o f backward stochastic differential equations , Stochastic Process. Appl. 111 (2004), no. 2, 175–206

  9. [17]

    Briand, B

    P. Briand, B. Delyon, Y. Hu, E. Pardoux and L. Stoica, Lp solutions of backward stochastic differential equations , Stochastic Process. Appl. 108 (2003) , 109–129

  10. [18]

    With this notation in hand, we have, taking into account (

    we also have that Fn(s,B n,t,x s ) = f (s, Θ n,t,x s ). With this notation in hand, we have, taking into account (

  11. [19]

    Briand, B

    Ph. Briand, B. Delyon, and J. Mémin, Donsker–type theorem for BSDEs , Electron. Comm. Probab. 6 (2001), 1–14, (electronic)

  12. [20]

    Jacod and A

    J. Jacod and A. N. Shiryaev, Limit theorems for stochastic processes , Springer (2003)

  13. [22]

    Choosing f (x) = x in ( 21), this implies the first result

    for r = 1, W1 ( g ( Bn,t,x s ) ,g ( Bt,x s )) ≤ ‖g‖ε W1 ( Bn,t,x s ,B t,x s ) ε ≤ ‖g‖ε cε 1 ( T n ) ε/ 2 . Choosing f (x) = x in ( 21), this implies the first result. Let us prove the second assertion. We start by observing that , since Bs −Bt and Bn s −Bn t are centered random...

  14. [23]

    Henry, Geometric theory of semilinear parabolic equations, Lecture Notes in Mathematics 840, Springer, 1981

    D. Henry, Geometric theory of semilinear parabolic equations, Lecture Notes in Mathematics 840, Springer, 1981

  15. [24]

    Ma and J

    J. Ma and J. Zhang, Representation theorems for backward stochastic different ial equations, Ann. Appl. Probab. 12 (2002), no. 4, 1390–1418

  16. [25]

    Pardoux and S

    É. Pardoux and S. Peng, Adapted solution of a backward stochastic differential equa tion, Systems Control Lett. 14 (1990), no. 1, 55–61

  17. [26]

    Rio, Upper bounds for minimal distances in the central limit theo rem, Ann

    E. Rio, Upper bounds for minimal distances in the central limit theo rem, Ann. Inst. Henri Poincaré Probab. Stat. 45 (2009), no. 3, 802–817

  18. [27]

    , Asymptotic constants for minimal distance in the central li mit theorem, Electron. Commun. Probab. 16 (2011), 96–103

  19. [28]

    Walsh, The rate of convergence of the binomial tree scheme , Finance Stoch

    John B. Walsh, The rate of convergence of the binomial tree scheme , Finance Stoch. 7 (2003), no. 3, 337–361

  20. [29]

    Zhang, A numerical scheme for BSDEs , Ann

    J. Zhang, A numerical scheme for BSDEs , Ann. Appl. Probab. 14 (2004), no. 1, 459–488

  21. [30]

    , Representation of solutions to BSDEs associated with a dege nerate FSDE , Ann. Appl. Probab. 15 (2005), no. 3, 1798–1831. 28

  22. [31]

    19 Since (t −t) = (t −t) ε 2 (t −t)1− ε 2 ≤h ε 2 (s −t)1− ε 2 and the same upper bound holds for s −s (since s −s ≤h ≤s −t), we get H1(s) ≤ C√ T −s h ε 2 (s −t)(1−ε )/ 2(s −t) ε 2

    of F gives H1(s) ≤C (s −t)(1+ε )/ 2 √ T −s (t −t) + (s −s) (s −t)(s −t) = C√ T −s (t −t) + (s −s) (s −t)(1−ε )/ 2(s −t). 19 Since (t −t) = (t −t) ε 2 (t −t)1− ε 2 ≤h ε 2 (s −t)1− ε 2 and the same upper bound holds for s −s (since s −s ≤h ≤s −t), we get H1(s) ≤ C√ T −s h ε 2 (s...

  23. [32]

    , Backward Stochastic Differential Equations , Springer, New York (2017). 29

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.