{"id":"acd9f816-feac-46da-84aa-17a548489417","arxiv_id":"2505.17032","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":0.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A review by the method's inventors that describes the Deep BSDE approach to high-dimensional PDEs and surveys follow-up work.","lead":"This paper is a review of the Deep BSDE method, a deep-learning approach for solving high-dimensional partial differential equations. It summarizes the method's formulation, subsequent algorithmic advances, and open theoretical problems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's load-bearing claim is the unsubstantiated priority assertion in §1/abstract that Deep BSDE was the first deep-learning method for general nonlinear PDEs in high dimensions; the review never systematically tests this against earlier work, including its own citation [35].","rationale":"The reader's weakest_assumption about the Pardoux–Peng equivalence is not the real risk: these are standard results and the review's citation is correct. The ψ/φ swap is a genuine typo but it is localized and does not threaten the method or the review's central message. The only claim whose failure would change the significance of the review is the priority/first assertion, which the reader also flagged as the strongest claim. However, the reader's conditional verdict does not rest on a systematic test of that priority claim. I therefore identify the unverified priority assertion as the load-bearing concern. The proposed literature check would settle it: if a clear predecessor is found, the priority sentences must be revised; if none is found, the claim stands. Since the mathematical description of Deep BSDE remains correct, the CONDITIONAL verdict is appropriate pending that check, so I recommend no change to the reader's verdict.","tokens_in":10223,"tokens_out":8498,"duration_ms":89444,"concrete_test":"Perform a targeted literature search (e.g., Google Scholar, zbMATH, MathSciNet) for neural-network PDE solvers published before 2017 that report numerical solutions of a semilinear parabolic PDE, HJB equation, or Schrödinger-type equation in dimension greater than or equal to 10, explicitly also checking Carleo–Troyer [13] and the workshop paper [35]. For each candidate, verify three conditions: (i) gradient-based training of neural networks, (ii) reported high-dimensional test cases, and (iii) publication before [24,40]. If any candidate meets all three, the §1 and abstract 'first' claim needs to be qualified; if none, the priority claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest load-bearing point is not the BSDE-to-PDE equivalence, which is standard and correctly cited, but the priority claim that anchors the abstract and §1: 'This method was the first numerical approach based on modern deep learning to effectively address general nonlinear PDEs in high dimensions.' The review asserts this without a systematic or even comparative literature account. The phrase 'general nonlinear PDEs' is exclusive enough to exclude the review's own citation [35] (a 2016 deep-learning method for high-dimensional stochastic control problems linked to HJB PDEs) only by an implicit boundary that is never stated. Earlier neural PDE work cited in §3 (e.g., [21,56,60]) is dismissed as low-dimensional rather than shown to be ineffective or non-scalable in high dimensions. Thus the central historical claim is credible but unestablished; if a pre-2017 neural method solved a high-dimensional nonlinear PDE effectively, the review's main contribution assertion would be false. This does not affect the mathematical correctness of the Deep BSDE formulation, but it is load-bearing for the review's purpose.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a short review of the Deep BSDE method for high-dimensional semilinear parabolic PDEs. It recalls the equivalence between such PDEs and backward stochastic differential equations, derives a variational/stochastic-control formulation, describes the neural-network discretization (forward Euler in time, terminal-condition loss), and surveys subsequent developments (BSDE-based schemes, PINNs/least-squares, Deep Ritz, weak adversarial networks) plus recent approximation and generalization theory. The abstract and Section 1 make the historical claim that Deep BSDE was the first numerical approach based on modern deep learning to effectively address general nonlinear PDEs in high dimensions; Section 4 outlines future directions in control, probabilistic modeling, quantum mechanics, kinetic equations, and theory.","tokens_in":10419,"tokens_out":14018,"duration_ms":125441,"significance":"If the historical priority claim is accepted, this is a useful concise reference: the mathematical summary is standard and essentially correct, the bibliography is comprehensive, and the paper explicitly acknowledges that full convergence analysis remains open. The review does not contain new proofs or numerical experiments, but it is a fair and compact exposition of the method. Its main weakness is that the 'first ... general nonlinear PDEs in high dimensions' claim is asserted rather than established; the mathematical content itself is not in question.","major_comments":[{"comment":"The sentence 'This method was the first numerical approach based on modern deep learning to effectively address general nonlinear PDEs in high dimensions' is a central historical claim for this review, but it is made without comparative evidence. The same paragraph credits [35] as the first deep-learning method for high-dimensional stochastic control problems, and HJB equations are nonlinear PDEs; Section 3 also acknowledges early neural PDE solvers [21,56,60]. To make the claim defensible, please either (i) state explicit exclusion criteria (why the HJB class is not 'general nonlinear PDEs', what 'modern deep learning' excludes, and what quantitative threshold 'effectively' implies), or (ii) qualify the claim, e.g. 'to our knowledge' or 'within the semilinear parabolic class considered here'. As written, the priority claim is credible but unestablished, and the unsupported 'hundreds or even thousands' dimension claim would benefit from a few representative numbers from [24,40].","section":"Section 1, first paragraph (and the abstract)"}],"minor_comments":[{"comment":"The roles of the two networks are reversed: with the notation of (2.5), (2.10), and (2.12), psi approximates u(0,·) and phi approximates sigma^* grad u, so the deterministic-case sentence should read 'psi_0 = u(0, xi_0) and phi_0 = sigma(0, xi_0)^* grad_x u(0, xi_0)', not the converse.","section":"Section 2, final paragraph"},{"comment":"The word 'euqations' in 'backward stochastic differential euqations' is a typo for 'equations'.","section":"Section 2, after (2.1)"},{"comment":"The text should read 'Feynman-Kac' rather than 'F eyman-Kac' and 'second-order BSDEs' rather than 'seconds order BSDEs'.","section":"Section 3, methods based on least-squares and BSDEs"},{"comment":"The sentence 'The minimizer of this variational problem is the solution to the PDE and vice versa' is slightly overbroad without standard regularity/well-posedness assumptions; please add a parenthetical reference to the assumptions under which (2.4) has a unique solution.","section":"Section 2, variational formulation"},{"comment":"The statement that the curse of dimensionality 'was originally coined' in this context needs a citation (standard attribution is to Bellman), or it should be softened.","section":"Section 4, optimal control paragraph"},{"comment":"Reference [39] ('Deep Picard iteration for high-dimensional nonlinear PDEs') is incomplete: it lacks a year and an arXiv or venue identifier.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The authors are the originators of the Deep BSDE method, so the 'first' claim in Section 1 is a claim about their own priority. A careful reader will want that claim either substantiated with a genuine comparison to [35] and the early neural-PDE literature, or explicitly scoped down. The paper also self-cites heavily; this is natural in a review of one's own method, but the historical discussion would be more convincing with a few independent citations for the priority question."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent, compact review of the Deep BSDE method by the people who introduced it. It does what a review should do: lays out the method, maps the follow-up literature, and flags open theory problems. No new results, which is fine.\n\nWhat it does well: the Section 2 derivation is clean. The BSDE-PDE equivalence via Itô and Pardoux-Peng is correctly stated and cited. The variational reformulation and the network discretization are easy to follow. Section 3 gives a useful taxonomy — BSDE-based methods, least-squares, Ritz, Galerkin, and theoretical analysis — across a large and messy literature. The reference list is broad.\n\nThe soft spots are real but limited. There is a notation swap in the final paragraph of Section 2: the text says φ0 = u(0,ξ0) and ψ0 = ∇_x u(0,ξ0), but based on the definitions in the same section it should be ψ0 = u(0,ξ0) and φ0 = ∇_x u(0,ξ0). Minor, but it could confuse an implementer.\n\nThe larger concern is the priority claim in the introduction: the Deep BSDE method was 'the first numerical approach based on modern deep learning to effectively address general nonlinear PDEs in high dimensions.' That is a load-bearing historical assertion, and the review does not support it with a systematic comparison. It acknowledges [35] as the first deep-learning method for high-dimensional stochastic control problems linked to HJB PDEs, and it mentions 1990s-era neural PDE solvers, but it never explains why those do not qualify as 'general nonlinear PDEs in high dimensions.' The claim may be true, but it is unestablished as written. For a review whose main contribution is contextual, that is a genuine weakness. It should be tempered or backed by a more careful literature account.\n\nThe mathematical core holds up. The BSDE-to-PDE equivalence is standard, and the cited uniqueness results are correct. The paper is appropriately cautious about theory, noting that optimization error remains poorly understood even in one dimension.\n\nWho this is for: a newcomer who wants a short, authoritative entry point into Deep BSDE and its offshoots. It will not change practice, but it will get someone oriented quickly.\n\nRecommendation: it deserves a serious referee. The notation error must be fixed, and the priority claim needs to be either supported or softened. After that, it is publishable as a review.","headline":"A clean, useful review of Deep BSDE by its inventors; the math is fine, but the 'first' claim in the introduction is stronger than the evidence supports.","tokens_in":10916,"tokens_out":4196,"would_cite":false,"duration_ms":38193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65M75","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The Deep BSDE method recasts high-dimensional nonlinear PDEs as stochastic control problems solved by neural networks; this review argues it was the first method of its kind and surveys the field it started.","keywords":["Deep BSDE method","high-dimensional PDEs","backward stochastic differential equations","curse of dimensionality","neural networks","semilinear parabolic PDEs","stochastic control","deep learning for PDEs"],"falsifier":"Choose a 100-dimensional semilinear parabolic PDE with a known solution and a nonlinear term $f$ that violates the Lipschitz-type conditions needed for the BSDE uniqueness results the review relies on. Run the Deep BSDE training, increasing network size and training steps, and compare the recovered $u(0,x_0)$ with the true value; if the error does not tend to zero, the method's validity outside the assumed equivalence regime is disproved.","tokens_in":10026,"feed_emoji":"🧠","tokens_out":11532,"duration_ms":98543,"temperature":0.7,"pith_summary":"This review argues that a particular reformulation—turning a semilinear parabolic PDE into a backward stochastic differential equation, then into a stochastic control problem—lets neural networks solve nonlinear PDEs in hundreds or even thousands of dimensions. The method, proposed in 2017, parameterizes the unknown solution and its gradient by feedforward networks, stacks one small subnetwork per time step, and trains by minimizing the squared mismatch between the network's terminal value and the PDE's terminal condition. The paper asserts that this was the first modern-deep-learning numerical method to effectively address general nonlinear high-dimensional PDEs, and it maps the subsequent wave of neural PDE solvers built on stochastic, least-squares, Ritz, and Galerkin formulations. A sympathetic reader would take away that the curse of dimensionality is not an absolute barrier once a stochastic representation and a trainable parameterization are both available.","feed_headline":"Neural networks beat the curse of dimensionality for nonlinear PDEs","feed_subtitle":"Rewriting the PDE as a backward stochastic equation bypasses the meshes that break down as dimension grows.","key_machinery":"The load-bearing machinery is the BSDE reformulation of the PDE, specifically the identity $Y_t=u(t,X_t)$ and $Z_t=[\\sigma(t,X_t)]^*\\nabla_x u(t,X_t)$ for a solution $u$. This identity, combined with existence-and-uniqueness results for BSDEs, converts the PDE into a stochastic control problem: choose a starting value $Y_0$ and a process $Z_t$ so that the forwardly defined $Y_t$ ends at $g(X_T)$. The Deep BSDE method discretizes time, assigns one small feedforward subnetwork to each time step to represent $Z_t$, and trains the whole stacked residual network by stochastic gradient descent on the terminal-matching loss $\\mathbb{E}|g(X_{t_N})-\\hat u|^2$. Because the Brownian paths and initial condition are sampled rather than stored, the loss is defined without a pre-existing training set.","core_discovery":"The paper's central claim is that the Deep BSDE method works. For a semilinear parabolic PDE of the form $\\partial_t u+\\mu\\cdot\\nabla_x u+\\frac12\\mathrm{Tr}(\\sigma\\sigma^*\\,\\mathrm{Hess}_x u)+f=0$, it uses the BSDE equivalence: with $X_t$ the forward diffusion, the pair $Y_t=u(t,X_t)$, $Z_t=[\\sigma(t,X_t)]^*\\nabla_x u(t,X_t)$ satisfies a backward equation, and uniqueness of the BSDE solution (from the cited existence-and-uniqueness theory) makes solving the PDE equivalent to minimizing $\\mathbb{E}|g(X_T)-Y_T|^2$ over the starting value $Y_0$ and the control process $Z_t$. The review then describes how neural subnetworks $\\psi_0,\\phi_n$ parameterize $Y_0$ and $Z_t$, how the stacked subnetworks form a deep residual network, and how the terminal-matching loss uses paths generated on the fly, giving effectively infinite training data. It further claims this was the first numerical method based on modern deep learning to effectively solve general nonlinear PDEs in high dimensions, a statement the review treats as historically load-bearing, and it places the method as the seed of several later families of deep PDE solvers.","pith_inferences":["The review does not quantify how large the constant in the polynomial parameter growth is; if the implied constants are enormous, the practical value of the no-curse-of-dimensionality results could be limited even though the method works on benchmarks.","The review's priority claim invites a concrete historical test: since it cites 1990s least-squares neural PDE methods, a reader can check whether those earlier methods were indeed confined to low dimensions, which would sharpen or weaken the 'first' statement.","The method's reliance on stochastic gradient descent suggests a robustness check the review leaves open: run the same high-dimensional benchmark with different optimizers and random seeds and compare terminal losses; if the outcome varies wildly, the bottleneck is optimization rather than representation."],"forward_implications":["For semilinear parabolic PDEs that admit the BSDE equivalence, the method returns numerical solutions in hundreds or even thousands of dimensions, a range traditional mesh-based methods cannot reach.","It gives a working algorithm for high-dimensional backward stochastic differential equations themselves, not just the PDEs they represent.","Several theoretical results cited in the review show that neural networks can approximate solutions of certain linear and semilinear high-dimensional PDEs with no curse of dimensionality: the number of parameters grows at most polynomially in the dimension and in the reciprocal of the target accuracy.","The same terminal-matching or residual-minimization idea reappears in least-squares methods (including physics-informed neural networks), the Deep Ritz method, and weak or Galerkin adversarial methods, so the original formulation seeded multiple research lines.","The review's own outlook places optimization error as the main unresolved piece: even for one-dimensional PDEs, a complete convergence proof for deep-learning PDE solvers is still open."],"supporting_citations":[{"why":"This is the original Deep BSDE paper; it supplies the algorithm, the stacked-subnetwork architecture, and the first high-dimensional numerical demonstrations.","marker":"[24]"},{"why":"This companion paper publicly demonstrated the method on high-dimensional nonlinear PDEs and made the field take notice.","marker":"[40]"},{"why":"This work establishes existence and up-to-equivalence uniqueness of solutions to BSDEs, the foundation of the PDE-BSDE equivalence.","marker":"[72]"},{"why":"This work extends the equivalence to forward-backward SDEs and quasilinear parabolic PDEs, supporting the reverse direction in which solving the BSDE gives the PDE solution.","marker":"[73]"},{"why":"This work documented the first deep-learning adaptation to stochastic control problems linked to Hamilton-Jacobi-Bellman PDEs; the review cites it as the immediate precursor to the Deep BSDE method.","marker":"[35]"}],"fun_headline_variants":["Deep learning cracks high-dimensional PDEs with BSDE trick","BSDE meets neural nets: high-dim PDEs solved","Neural nets overcome curse of dimensionality in PDEs","BSDE trick: neural nets conquer high-dimensional PDEs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method rests on the assumption that the semilinear parabolic PDE and its associated backward stochastic differential equation are genuinely equivalent under the stated regularity conditions; if that equivalence fails, the mapping to a stochastic control problem collapses.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning cracks high-dimensional PDEs with BSDE trick","BSDE meets neural nets: high-dim PDEs solved","Neural nets overcome curse of dimensionality in PDEs","BSDE trick: neural nets conquer high-dimensional PDEs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000675,"raw_usage":{"total_tokens":3049,"prompt_tokens":900,"completion_tokens":2149,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":2083}},"tokens_in":516,"tokens_out":2149,"duration_ms":13925,"temperature":1.0,"reasoning_tokens":2083,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:23:36.066026+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Choose a 100-dimensional semilinear parabolic PDE with a known solution and a nonlinear term $f$ that violates the Lipschitz-type conditions needed for the BSDE uniqueness results the review relies on. Run the Deep BSDE training, increasing network size and training steps, and compare the recovered $u(0,x_0)$ with the true value; if the error does not tend to zero, the method's validity outside the assumed equivalence regime is disproved.","supporting_citations":[{"cited_title":"Deep learning-based numerical meth- ods for high-dimensional parabolic partial differential equations and backward stochastic differential equations","cited_arxiv_id":null,"evidence_quote":"This is the original Deep BSDE paper; it supplies the algorithm, the stacked-subnetwork architecture, and the first high-dimensional numerical demonstrations."},{"cited_title":"Solving high-dimensional partial differ- ential equations using deep learning","cited_arxiv_id":null,"evidence_quote":"This companion paper publicly demonstrated the method on high-dimensional nonlinear PDEs and made the field take notice."},{"cited_title":"Backward stochastic differential equations and quasilinear parabolic partial differential equations","cited_arxiv_id":null,"evidence_quote":"This work establishes existence and up-to-equivalence uniqueness of solutions to BSDEs, the foundation of the PDE-BSDE equivalence."},{"cited_title":"Forward-backward stochastic differential equa- tions and quasilinear parabolic PDEs","cited_arxiv_id":null,"evidence_quote":"This work extends the equivalence to forward-backward SDEs and quasilinear parabolic PDEs, supporting the reverse direction in which solving the BSDE gives the PDE solution."},{"cited_title":"Deep learning approximation for stochastic control problems","cited_arxiv_id":null,"evidence_quote":"This work documented the first deep-learning adaptation to stochastic control problems linked to Hamilton-Jacobi-Bellman PDEs; the review cites it as the immediate precursor to the Deep BSDE method."}],"review_version":1}