{"id":"e8b24c06-8c4c-406e-868c-ac0af00b548e","arxiv_id":"2507.10607","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Neural Expectation Operators are defined as solutions of BSDEs with neural-network drivers; the paper proves well-posedness under local Lipschitz and quadratic growth assumptions and gives architectural constraints for convexity and monotonicity.","lead":"This paper proposes to model uncertainty by learning the 'driver' term of a backward stochastic differential equation with a neural network, and proves that such equations have unique solutions under quadratic-growth and monotonicity conditions. It extends the theory to coupled systems and large interacting populations, and applies it to a portfolio choice problem.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.7's existence proof fails to close: Step 2.2's Cauchy estimate cannot control the quadratic z-difference, so the central well-posedness claim is not established as written.","rationale":"The reader's weakest-assumption analysis identifies monotonicity in y as the main structural restriction. I do not dispute that monotonicity is essential to the comparison-based proof. However, the most load-bearing failure I find is independent of that assumption: the existence half of Theorem 3.7 rests on an approximation argument whose Cauchy step (Step 2.2) is not justified. The quadratic growth of the driver means the difference Δf_s between two approximating solutions is not controlled by |δZ_s|, so Gronwall cannot eliminate it from the stability estimate. This is a concrete proof gap in the central claim. The theorem is likely true via known quadratic-BSDE results, which is why I keep the conditional verdict rather than rejecting or marking the paper unverdictable; but the manuscript must either supply the missing estimate or replace the approximation argument with a correct one. The reader's rationale already notes 'unsupported stability estimates and hand-waved convergence steps'; my critique sharpens that to a specific false step, hence partial agreement.","tokens_in":29013,"tokens_out":16734,"duration_ms":201478,"concrete_test":"Independently re-derive the Step 2.2 stability estimate from the cited references for the admissible driver fθ(t,y,z)=α/2|z|² with d=1 and no y-dependence. Write the Itô formula for |δY|²: E∫|δZ|² ds = 2E∫ δY_s (fθ(Z^k_s)−fθ(Z^m_s)) ds = αE∫ δY_s (|Z^k_s|²−|Z^m_s|²) ds. Check whether this right-hand side can be bounded by a constant times E|ξ_k−ξ_m|² with a constant independent of the BMO norms of Z^k and Z^m; if no such estimate follows, the manuscript's Cauchy argument fails and must be replaced by the exponential-transform proof. A second check: locate the displayed inequality in Kobylanski (2000) or in El Karoui, Peng and Quenez (1997); the latter's stability estimates are for Lipschitz drivers, so their citation does not justify the quadratic case.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript's central claim is Theorem 3.7. Its proof is not valid as written. In Step 2.2 of Part 2, for approximate solutions (Y^k,Z^k) and (Y^m,Z^m), the paper invokes \"standard stability estimates for quadratic BSDEs\" and asserts E[sup|δY|^2 + ∫|δZ|^2] ≤ C_stab E[|ξ_k−ξ_m|^2 + (∫|Δf_s|ds)^2], then claims that uniform local Lipschitz continuity in y plus Gronwall reduces the right side to C' E|ξ_k−ξ_m|^2. The driver difference is Δf_s = fθ(s,Y^k_s,Z^k_s)−fθ(s,Y^m_s,Z^m_s). Under Assumption 3.3(ii), this term contains a pure quadratic-z contribution of order (α/2)(|Z^k|^2+|Z^m|^2); it is not proportional to |δZ_s|. Hence the L² norm of ∫|Δf|ds does not vanish merely because ξ_k→ξ_m in L²; it depends on the very quantity ∫|δZ|^2 whose convergence is the target of the estimate. The claimed Gronwall simplification is therefore unsupported, and the sequence is not shown to be Cauchy in S²×H². The standard Kobylanski route for quadratic BSDEs uses an exponential change of variable and BMO arguments, not this raw S²-stability inequality. This gap is load-bearing: even if one grants the monotonicity assumption (3.3(iv)), the existence half of Theorem 3.7 is not established by the manuscript. The theorem may still be true as a specialization of known quadratic-BSDE results, but the present proof does not supply the required argument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces \"Neural Expectation Operators,\" defined as solutions of backward stochastic differential equations whose drivers are neural networks, and develops a framework called \"Measure Learning\" for data-driven modeling of ambiguity. The central theoretical claim is Theorem 3.7, which asserts well-posedness in S^∞ × H^2_BMO for such BSDEs under a local-Lipschitz-in-y, quadratic-in-z, monotone driver and an exponentially integrable terminal condition. The paper also proposes constructive neural architectures satisfying the assumptions, derives axiomatic properties such as convexity and Jensen's inequality, extends the framework to fully coupled FBSDEs and mean-field systems with a propagation-of-chaos theorem and a central limit theorem, and illustrates the framework on a Merton portfolio problem.","tokens_in":29368,"tokens_out":6422,"duration_ms":77528,"significance":"If Theorem 3.7 is established with a complete proof, the paper provides a useful bridge between the deep theory of quadratic BSDEs and neural-network-based modeling of nonlinear expectations. The constructive architectural results in Proposition 3.5, the candid discussion of the restrictive monotonicity assumption, and the explicit treatment of BMO spaces are valuable and will be of interest to researchers working at the interface of stochastic analysis and machine learning. However, the proofs of the main theorem and several extensions are incomplete as written, and the central well-posedness claim is not yet established by the manuscript.","major_comments":[{"comment":"The asserted stability estimate cannot be used to prove that the approximation sequence is Cauchy as written. For δY = Y^k − Y^m and δZ = Z^k − Z^m, the driver difference Δf_s = f_θ(s,X_s,Y^k_s,Z^k_s) − f_θ(s,X_s,Y^m_s,Z^m_s) contains, under Assumption 3.3(ii), a quadratic contribution roughly of order (α/2)(|Z^k_s|^2 − |Z^m_s|^2). This term is not controlled by |δZ_s| and does not vanish merely because ξ_k → ξ_m in L^2. The displayed inequality bounding E[sup|δY|^2 + ∫|δZ|^2] by C_stab E[|ξ_k−ξ_m|^2 + (∫|Δf_s|ds)^2] is therefore not compatible with the subsequent claim that Gronwall's lemma reduces the bound to C' E|ξ_k−ξ_m|^2. A complete proof needs a Kobylanski-type exponential transformation and BMO estimates, or an explicit citation of a theorem that applies to this class of drivers. As written, the existence half of Theorem 3.7 is not established.","section":"§3.3, Step 2.2 of Theorem 3.7"},{"comment":"The convergence of the driver integrals is asserted in a single sentence: 'The stability properties of quadratic BSDEs ... ensure' that E[∫|f_k(s,...,Y^k,Z^k) − f_θ(s,...,Y,Z)|ds] → 0. This is a nontrivial point because f has quadratic growth in z and convergence in S^2 × H^2 together with uniform S^∞ and BMO bounds does not by itself imply L^1 convergence of the driver. The proof must supply an argument that controls ∫|Z^k_s|^2 ds using the BMO property, or invoke a precise theorem that yields this convergence. Without this, the passage to the limit in the integral equation is incomplete.","section":"§3.3, Step 2.3 of Theorem 3.7"},{"comment":"The convex-analysis inequality used in the proof is false as stated. For a convex h and a C^2 convex φ, the inequality h(φ(y),φ'(y)z) ≥ φ'(y)h(y,z) − (1/2)φ''(y)|z|^2 is not a consequence of convexity. For example, take h(y,z)=−y+z^2, which is convex and non-increasing in y, and φ(y)=e^y. At y=−10 and z=0, the left-hand side is −e^{−10}, while the right-hand side is 10e^{−10}. Thus the proof of Jensen's inequality for Neural Expectations is invalid. The statement may still be true under additional conditions, but it requires a correct proof or a clear retraction.","section":"§4.3, Proposition 4.3"},{"comment":"The BSDE stability step in the proof of the law of large numbers repeats the difficulty identified in Step 2.2 of Theorem 3.7: the displayed stability inequality for quadratic BSDEs and the subsequent Gronwall reduction are not justified for a quadratic driver. In addition, the proof asserts that Y^N and ar Y are uniformly bounded 'by a comparison argument analogous to Theorem 3.7,' but the mean-field dependence of f_θ requires an extra a priori estimate or exponential-moment condition involving sup_s W_2(µ^N_s, µ_s), which is not established. The theorem may be true, but the proof as written does not close these gaps.","section":"§6.2, Theorem 6.2"},{"comment":"The proof of the central limit theorem is only a sketch. Tightness of the fluctuation processes (U^{i,N}, V^{i,N}, Z^{i,N}) and uniqueness of the limiting linear McKean-Vlasov FBSDE are asserted to follow from 'standard but technical' arguments without being supplied. Since the backward driver f_θ has quadratic growth in z, the passage from the finite-N fluctuation BSDEs to the limiting equation requires estimates that are not provided. This gap is load-bearing for the CLT claim.","section":"§6.5, Theorem 6.5"}],"minor_comments":[{"comment":"The text contains stray LaTeX artifacts such as '/bracehtipupleft/bracehtipdownright/bracehtipdownleft/bracehtipupright' that should be removed before publication.","section":"§4.3.1"},{"comment":"In part (a), N_1 is described as 'any network satisfying the quadratic growth condition,' but the proposition should also require the continuity and progressive measurability needed for Assumption 3.3(i).","section":"§3.2, Proposition 3.5(a)"},{"comment":"The smallness condition is stated with an unspecified constant C; to be verifiable, it should be expressed explicitly in terms of the Lipschitz constants and the horizon T, as in the reference to Delarue (2002).","section":"§5.1, Assumption 5.1(iv)"},{"comment":"The proof is explicitly heuristic ('A formal proof relies on comparison theorems', 'A full analysis shows'), yet the proposition is stated without qualification; it should either be proved rigorously or labeled as a heuristic/conjecture.","section":"§8.4, Proposition 8.2"},{"comment":"The implicit equation z = (∇_x u)^T σ_θ(...,z) is introduced before the contraction argument that guarantees a unique solution; the exposition would be clearer if the smallness condition were stated before the construction.","section":"§5, Part 2 of Theorem 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's value rests almost entirely on Theorem 3.7, and the proof gap in Step 2.2 is the key issue. The central claim is likely true as a specialization of known quadratic-BSDE theory, so this is not a clear reject; a revision that replaces the unsupported stability estimates with a complete Kobylanski-type argument, or that explicitly imports a precise existence theorem, would make the manuscript viable. I would not object on scope grounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper proposes a sensible bridge between quadratic BSDE theory and neural-network parameterized drivers, and the architectural constructions are correct as far as they go. But the proof of the main well-posedness theorem has a load-bearing gap, and the applications overclaim what the theorem delivers. This deserves a serious referee, but as written it is not a solid paper.\n\nWhat is actually new: the observation that sign-constrained ReLU networks can satisfy the monotonicity and local Lipschitz hypotheses of the quadratic BSDE theory, and the idea of treating the resulting operator as a learnable object. Proposition 3.5 is straightforward but correct. The axiomatic properties (dynamic consistency, convexity via ICNN) are standard consequences of the BSDE framework, stated cleanly.\n\nThe soft spots are significant. Theorem 3.7's proof is not valid as written. The stability estimate in Step 2.2 is asserted, not derived, and the stress-test concern is correct: the driver difference contains a pure quadratic z-term, so the claimed Cauchy estimate does not close. The standard Kobylanski route uses exponential transforms and BMO arguments, not the raw S^2 inequality invoked here. The theorem may be true as a specialization of known results, but this paper does not supply the proof. The mean-field LLN and CLT sections are similarly sketchy; the stability estimates are hand-waved and the CLT is a derivation sketch. Finally, the Merton application claims all conditions of Theorem 3.7 hold, but the terminal utility U(X_T) does not satisfy the required exponential integrability for any γ≠0, so the application is not covered by the theorem. That is a clear mismatch.\n\nThe paper is honest about its limitations, and the writing is clear. It would be useful as a roadmap for someone who wants to connect quadratic BSDEs with neural driver parameterization, but the central theorem needs a repaired proof and the applications need either verification of hypotheses or explicit conjecture labels.\n\nRecommendation: send to peer review, but flag the main proof gap and the Merton integrability issue for the referee to check.","headline":"A useful bridge idea between quadratic BSDEs and neural drivers, but the main theorem's proof has a real gap and the Merton application overclaims; worth refereeing, not citable as is.","tokens_in":29894,"tokens_out":2783,"would_cite":false,"duration_ms":32557,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60H10","93E20","60K35","91G80","35K59","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes that a neural-network driver satisfying quadratic growth in the martingale component, uniform local Lipschitz dependence on the value, and monotonicity in that value, together with an exponentially integrable…","keywords":["Backward Stochastic Differential Equations","Non-linear Expectation","Neural Networks","Quadratic BSDEs","Mean-Field Systems","Propagation of Chaos","Stochastic Control","Ambiguity"],"falsifier":"Take the separable architecture from Proposition 3.5(a), for instance $f_\\theta(y,z)=N_2(y)+\\tfrac{\\alpha}{2}z^2$ with $N_2$ a smooth bounded non-increasing function and terminal condition $\\xi=e^{\\beta_0 W_T}$ with $\\beta_0>\\alpha$, then run the paper's truncation scheme for two different approximating sequences; if the two limits differ in $S^2\\times H^2$, or leave $S^\\infty\\times H^2_{\\mathrm{BMO}}$, then Theorem 3.7 is false.","tokens_in":28763,"feed_emoji":"🧠","tokens_out":8128,"duration_ms":89847,"temperature":0.7,"pith_summary":"The paper introduces 'Measure Learning': instead of committing to a single probability measure or a hand-picked driver, it represents model ambiguity by a backward stochastic differential equation (BSDE) whose driver is a neural network, and calls the resulting solution operator a Neural Expectation Operator. The central result, Theorem 3.7, gives conditions under which such an equation is well-posed, meaning it has one unique solution, even though neural drivers are not globally Lipschitz. The paper then shows constructively that these conditions are achievable by concrete network designs: separable or bounded-interaction architectures, sign constraints that enforce monotonicity, and input-convex networks that enforce convexity. If the theorem is right, ambiguity-aware models in finance, control, and mean-field systems can be learned from data with a rigorous guarantee that the learned expectation is a genuine mathematical object.","feed_headline":"Unique solutions for neural-network ambiguity models","feed_subtitle":"Theorem 3.7 makes quadratically growing BSDE drivers well-posed, at the price of one monotonicity constraint.","key_machinery":"The central object is the Neural BSDE, whose driver $f_\\theta$ is parameterized by a neural network. Four components carry the argument: the quadratic-growth condition $|f_\\theta(t,x,y,z)|\\leq K(1+\\|x\\|^p+|y|)+\\tfrac{\\alpha}{2}\\|z\\|^2$, the uniform local Lipschitz condition in $y$ whose constant depends only on the radius $R$ and not on $t,x,z$, the exponential integrability $E[e^{\\beta_0|\\xi|}]<\\infty$ with $\\beta_0>\\alpha$, and the space of BMO martingales that makes Girsanov changes of measure available. The proof machinery is the comparison and domination theory of quadratic BSDEs: a dominating bounded BSDE supplies uniform bounds, monotonicity in $y$ lets the comparison principle compare solutions, and BMO control of $Z$ closes the estimates. On the architecture side, monotonicity is enforced by a sign-constrained feed-forward network and convexity by an input-convex network.","core_discovery":"The paper proves that for a terminal condition with $E[e^{\\beta_0|\\xi|}]<\\infty$ where $\\beta_0>\\alpha$, a neural driver $f_\\theta$ that is continuous, at most quadratically growing in $z$, uniformly locally Lipschitz in $y$, and non-increasing in $y$, the BSDE $Y_s=\\xi+\\int_s^T f_\\theta(r,X_r,Y_r,Z_r)\\,dr-\\int_s^T Z_r\\,dW_r$ has a unique solution $(Y,Z)$ in $S^\\infty\\times H^2_{\\mathrm{BMO}}$. Uniqueness comes from the comparison principle for quadratic BSDEs, which the monotonicity assumption makes available; existence comes from truncating the terminal condition and the $y$-dependence, obtaining uniform $S^\\infty$ and BMO bounds, and passing to the limit. From this foundation the paper derives the axiomatic properties of the resulting expectation, including dynamic consistency, monotonicity, normalization under a structural driver condition, and convexity when the driver is an input-convex network. It extends the well-posedness to fully coupled forward-backward systems via the four-step scheme, and proves a law of large numbers and a central limit theorem for interacting-particle limits under learned ambiguity.","pith_inferences":["If the monotonicity assumption could be relaxed, the class of learnable ambiguity models would grow substantially; a natural test is to run the paper's truncation scheme on a non-monotone but still quadratically growing driver and look for multiple limit points.","The normalization condition $f_\\theta(t,x,y,0)=0$ is not enforced by architecture, and the proposed regularization suggests that learned models will trade exact normalization against fit, which is a falsifiable identifiability question.","The linearized fluctuation system in Theorem 6.5 could be solved numerically through the decoupled PDE representation mentioned in Remark 6.6, offering a scalable way to compute distributional uncertainty around mean-field limits.","Wealth-dependent investment in the Merton application is a testable signature of learned ambiguity: with $0<\\gamma<1$ the allocation fraction should fall with wealth, while with $\\gamma<0$ it should rise."],"forward_implications":["Every driver that meets Assumptions 3.1 and 3.3 defines a unique non-linear expectation $E_\\theta$, so the statistical learning problem for $\\theta$ is a well-posed optimization rather than a heuristic.","ReLU and GeLU activations are compatible with the theory, since the uniform local Lipschitz condition replaces the classical global Lipschitz assumption.","Risk-averse convex expectations can be built into the architecture, giving data-driven convex risk measures that are well-posed by construction.","Fully coupled neural forward-backward systems are well-posed under the stated Lipschitz and hybrid conditions, and their mean-field limits satisfy propagation of chaos and a central limit theorem.","In the Merton application, learned ambiguity makes the optimal allocation uniformly smaller than the classical Merton fraction and decreasing in the ambiguity parameter $\\theta$.","In the Merton application, learned ambiguity makes the optimal allocation uniformly smaller than the classical Merton fraction and decreasing in the ambiguity parameter $\\theta$."],"supporting_citations":[{"why":"Supplies the quadratic BSDE well-posedness, comparison principles, and dominating-solution estimates that Theorem 3.7 adapts.","marker":"Kobylanski (2000)"},{"why":"Grounds the classical global-Lipschitz well-posedness that the paper's result extends.","marker":"Pardoux and Peng (1990)"},{"why":"Provides BSDE stability estimates, comparison arguments, and the financial risk-measure interpretation used throughout.","marker":"El Karoui, Peng and Quenez (1997)"},{"why":"Gives $L^p$ solutions and bounded dominating solutions under exponential integrability, used for the a priori bounds.","marker":"Briand et al. (2003)"},{"why":"Supplies the convex-generator inequality behind the Jensen inequality for the Neural Expectation.","marker":"Briand and Hu (2008)"},{"why":"Provides the four-step-scheme extension to quadratic-growth FBSDEs used in Theorem 5.2.","marker":"Delarue (2002)"},{"why":"Original four-step scheme that connects FBSDEs to quasilinear PDEs in the coupled case.","marker":"Ma, Protter and Yong (1994)"},{"why":"Canonical mean-field theory used for propagation of chaos and the central limit theorem.","marker":"Carmona and Delarue (2018)"},{"why":"Supplies the input-convex neural network architecture used to enforce convexity.","marker":"Amos, Xu and Kolter (2017)"}],"fun_headline_variants":["Neural BSDEs: well-posed for quadratic growth","Quadratic BSDEs meet ReLU networks","Ambiguity modeling with neural expectations","Well-posed neural BSDEs without global Lipschitz","Measure Learning: neural non-linear expectations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole proof leans on the driver being non-increasing in the value variable $y$; if that monotonicity property is dropped, the comparison theorem that delivers existence and uniqueness no longer applies, so the central claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["Neural BSDEs: well-posed for quadratic growth","Quadratic BSDEs meet ReLU networks","Ambiguity modeling with neural expectations","Well-posed neural BSDEs without global Lipschitz","Measure Learning: neural non-linear expectations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000572,"raw_usage":{"total_tokens":2754,"prompt_tokens":1043,"completion_tokens":1711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":1638}},"tokens_in":659,"tokens_out":1711,"duration_ms":16364,"temperature":1.0,"reasoning_tokens":1638,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:55:15.646639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the separable architecture from Proposition 3.5(a), for instance $f_\\theta(y,z)=N_2(y)+\\tfrac{\\alpha}{2}z^2$ with $N_2$ a smooth bounded non-increasing function and terminal condition $\\xi=e^{\\beta_0 W_T}$ with $\\beta_0>\\alpha$, then run the paper's truncation scheme for two different approximating sequences; if the two limits differ in $S^2\\times H^2$, or leave $S^\\infty\\times H^2_{\\mathrm{BMO}}$, then Theorem 3.7 is false.","supporting_citations":[{"cited_title":", Delyon , Bernard B","cited_arxiv_id":null,"evidence_quote":"Gives $L^p$ solutions and bounded dominating solutions under exponential integrability, used for the a priori bounds."},{"cited_title":"Hu , Ying Y","cited_arxiv_id":null,"evidence_quote":"Supplies the convex-generator inequality behind the Jensen inequality for the Neural Expectation."},{"cited_title":"( 2002 )","cited_arxiv_id":null,"evidence_quote":"Provides the four-step-scheme extension to quadratic-growth FBSDEs used in Theorem 5.2."},{"cited_title":", Xu , Lei L","cited_arxiv_id":null,"evidence_quote":"Supplies the input-convex neural network architecture used to enforce convexity."}],"review_version":1}