REVIEW 3 major objections 5 minor 29 references
The paper proves that, under smoothness and uniform ellipticity, the value function of an expectation-constrained stochastic control problem is fully characterized by an interior Hamilton-Jacobi-Bellman equation plus a Dirichlet condition o
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-07-31 23:00 UTC pith:NJHRRC4R
load-bearing objection Useful and mostly sound paper that deserves refereeing; the central Dirichlet characterization is conditional on a non-trivial continuous-feedback selection assumption that the numerics do not verify. the 3 major comments →
Optimal Control with Expectation Constraint in a Smooth Boundary Case
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is Theorem 2.13: under smoothness of the terminal constraint G, Hölder continuity of the running cost g, and uniform ellipticity of the diffusion σ, the boundary trace V̄(t,x) = V(t,x,ϖ(t,x)) is a viscosity subsolution of the reduced Hamilton-Jacobi-Bellman equation −max_{u∈U(t,x)}(L^u_X φ + f) = 0 with terminal condition φ(T,x) = F(x), and, under an additional continuous-selection assumption, a viscosity supersolution as well. Here U(t,x) is the set of controls attaining the minimum in the HJB equation that defines w, and L^u_X is the linear parabolic operator generated by drift μ and diffusion σ. Since w is smooth and the boundary is absorbing, this gives a proper Dir
What carries the argument
The engine of the argument is the martingale representation of the expectation constraint: the inequality E[G(X_T)+∫g] ≤ p is rewritten as the existence of a martingale P^α with P^α(T) ≥ G(X_T)+∫g, so the constraint becomes the geometric state constraint P^α ≥ w(·,X) on the whole time interval. The paper exploits the fact that under uniform ellipticity the boundary function w is smooth (Proposition 2.8), and therefore the boundary is absorbing: once the martingale touches w, the optimal continuation keeps it there. The key identity is that on the boundary the martingale control is forced to be α = Dϖ σ and the state control must belong to the argmin set U(t,x) = {u : L^u_X ϖ + g = 0} of the
Load-bearing premise
The load-bearing premise is Assumption 2.10: for every starting point and every boundary-optimal control, there must exist a globally defined continuous feedback control that stays in the boundary-argmin set and yields a unique strong SDE solution; if this selection does not exist, the boundary value function may fail to be a viscosity supersolution and the bounded-control approximation may not converge.
What would settle it
Construct a smooth, uniformly elliptic one-dimensional example in which the argmin set U(t,x) of the boundary HJB equation switches between two isolated controls and admits no continuous selection, then compute the boundary trace V̄ by Monte Carlo and check the supersolution inequality of (2.10) at the switching locus; failure would show Assumption 2.10 is necessary. Separately, run the bounded-control scheme with N large and the ε-relaxation in a problem where U(t,x) is not a singleton; if lim_{ε↓0} lim_{N→∞} V^N(t,x,p+ε) differs from V(t,x,p), the singleton hypothesis in Proposition 3.1 is e
If this is right
- If comparison holds for the reduced equation (2.10)-(2.11), then the boundary trace V̄ is its unique solution, so the value function on the whole domain is determined by the interior HJB equation plus this Dirichlet data; numerical schemes can be targeted at this complete PDE system.
- The convergence of the bounded-control approximations V^N to V (up to a small relaxation of the constraint) means one can solve a compact-control HJB equation with comparison and read off the original value function for constraint levels arbitrarily close to the admissible set.
- For degenerate or non-smooth problems, the small-noise approximation V^ε → V pointwise provides a systematic route to regularize the problem before applying numerical methods.
- In the asset-liability toy model, the algorithm produces a total normalized error of about 0.069%, with controls of bang-bang type: sell and share losses when the asset is below book value, buy and keep profits when it is at or above book value.
Where Pith is reading between the lines
- An implicit consequence the paper does not spell out: because V is concave in the constraint level p (used in the proof of Lemma 3.3), a dual or Legendre-transform formulation may hold, potentially simplifying numerics; this is not established in the paper.
- The singleton assumption on U(t,x) in Proposition 3.1 looks stronger than necessary: since the convergence statement allows a small relaxation p+ε of the constraint, one might expect the result to survive with measurable selections whenever the closed-loop SDE can be solved; this remains an open extension.
- The absorbing-boundary mechanism likely persists in problems where w is only C^{1,2} on the reachable region, so the small-noise approximation suggests a general recipe—regularize coefficients, solve the boundary problem, then take the limit—though the paper proves convergence without rates.
- The toy ALM example is deliberately simplistic (no transaction costs, constant book value, linear drifts), so the qualitative behavior found there should not be read as a general policy prescription for real insurance portfolios.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a stochastic optimal control problem with a terminal constraint in expectation. Following Bouchard et al. and Bouchard–Nutz, it reformulates the constraint via a martingale representation and identifies the state domain D={p≥w(t,x)}. Assuming uniform ellipticity and smooth data (Assumptions 2.6–2.7), the paper proves that the boundary is smooth (Proposition 2.8) and that the trace of the value function on the boundary solves a reduced HJB equation (Theorem 2.13), with the supersolution property conditional on a continuous-feedback selection assumption (Assumption 2.10). It then provides two approximation results: bounded-martingale controls converge after an epsilon bump (Proposition 3.1), and degenerate problems can be regularized by adding noise (Proposition 3.5). A three-step deep-learning/PINN algorithm is proposed and applied to a toy ALM model, with PDE residuals reported as error estimates.
Significance. Read carefully, the conditional results are a solid contribution: the paper isolates a smooth-boundary setting in which the value function on the endogenous boundary is characterized by a reduced PDE, it gives a bounded-control approximation for which comparison can hold, and it demonstrates a numerical pipeline on a nontrivial toy model. The proof of Proposition 2.8 is self-contained and uses standard parabolic estimates, and the convergence arguments in Section 3 are nontrivial and mostly convincing under the stated assumptions. However, the breadth of the claims in the abstract is wider than what is proved, because Assumption 2.10 can fail in simple uniformly elliptic problems, and the numerical section works with a model for which that assumption is not verified. The paper would be a useful reference if the scope were narrowed, the assumptions stated more explicitly, and the numerical error measures reported as residuals rather than independent errors.
major comments (3)
- [Assumption 2.10; Theorem 2.13; §6.2] The supersolution part of Theorem 2.13 and the convergence result in Proposition 3.1 rest on Assumption 2.10, which requires a continuous feedback selector of the multifunction U(t,x). This is not a consequence of Assumptions 2.6–2.7. Concrete counterexample: d=1, U={-1,1}, μ(x,u)=u, σ=1, ϖ(x)=e^{-x^2}, g(x,u)=-1/2 ϖ''(x)+|ϖ'(x)|. Then (2.7) holds with U(x)={-sign(ϖ'(x))} for x≠0 and U(0)={-1,1}; no continuous selection exists. This satisfies the smoothness and ellipticity assumptions, so the abstract's unconditional claim of a proper Dirichlet condition in the uniformly elliptic case is too strong. The paper should either remove the unconditional wording, prove a more robust selection theorem for the cases of interest, or show that the ALM model satisfies Assumption 2.10.
- [Proposition 3.1; §4.2; §5.3] The bounded-control convergence result (Proposition 3.1) additionally requires U(t,x) to be a singleton. This is a substantial restriction: it excludes bang-bang problems in which the minimizer is unique except on a switching surface. Section 5.3 explicitly states that the ALM model's optimal controls are bang-bang, yet Section 4.2 assumes 'only one feasible control' and the numerical algorithm is run on the original model, not on a regularized version with ε|u|^2. Assumption 2.10 is never verified for the ALM model. Consequently, the numerical section does not demonstrate the theory for the model it solves; it only shows the algorithm is implementable under an additional assumption. The authors should either verify Assumption 2.10 (and the singleton condition) for the ALM model, or add a control regularization as in Section 3.2 and solve the regularized problem, and state clearly that t
- [§4.2–§4.3, Eq. (4.7)–(4.13); §5.4] The reported error E^V is not an independent numerical-error estimate. It is the normalized residual of the PDE used in the loss function, evaluated with the trained control network. Since the same residual is being minimized during training and the control used in ε^u_V is the estimated one, E^V reflects training success, not accuracy relative to the true value function. The claim in the abstract that the numerical resolution is 'complemented by an estimation of the numerical error' is therefore overstated. An independent benchmark (e.g., a Monte Carlo value computed by a different scheme, or a comparison on a problem with known solution) would be needed; absent that, the phrase should be softened to 'residual monitoring' or similar.
minor comments (5)
- [§4.2] Typo: 'feasable' should be 'feasible'.
- [§4.2–§4.3] Notation conflict: \hat V is used in Section 4.2 for the boundary estimate and again in Section 4.3 for the interior estimate. Suggest \hat V^b and \hat V^i.
- [Eq. (4.13)] ε^T_V is defined with arguments (x,p) but used with (X̂...,P̂...); align notation. Also, the symbols ε^υ_V and ε^u_V appear inconsistent in Section 4.3.
- [Remark 2.12] The definitions of \bar V_* and \bar V^* as 'envelopes' on the boundary are easy to confuse with the full-space semicontinuous envelopes; consider naming them \hat V_* and \hat V^*.
- [§6.2, step a.1] The proof uses h_n=√γ_n; if γ_n=0 for some n, h_n is not positive. This is fixable by a standard perturbation or by passing to a subsequence with γ_n>0, but it should be stated.
Circularity Check
Central boundary-PDE characterization is derived honestly from published dynamic-programming/martingale-representation inputs plus an explicit selection assumption; no fitted parameter is renamed as a theorem. The only circularity-like element is the numerical 'error' statistic, which is the same PINN residual used for training and is not an independent benchmark.
specific steps
-
fitted input called prediction
[Section 4.2 eqs. (4.6)-(4.8) and Section 4.3 eqs. (4.11)-(4.13)]
"Similar to the previous section, we also propose an error measure for the joint training of ν̂ϖ and V̂ defined by E^V := δV / (1/J Σ_{j=1}^J |V̂(t0, X̂^{ν̂ϖ,j,j}_{t0})|) where δV := 1/J Σ_{j=1}^J ( |ε^T_V(X̂^{ν̂ϖ,j,j}_T)| + 1/K Σ_{k=0}^{K-1} (|ε^u_V|+|ε^V_V|)(t_k, X̂^{ν̂ϖ,j,j}_{t_k}) ) ... ε^V_V(t,x):=L^{ν̂ϖ(t,x)}_X V̂(t,x) + f(x,ν̂ϖ(t,x))"
The reported error E^V is the same residue that the PINN loss (4.6) minimizes: the loss contains |L^{ν̂ϖ}_X n^V_θ + f|^2 and |n^V_θ(T)-F|^2, which are exactly ε^V_V and ε^T_V in (4.8); the interior error (4.12)-(4.13) likewise mirrors the loss terms of (4.11) (PDE, boundary, terminal). The control ν̂ϖ is itself trained in the previous step to drive the same residual down. Therefore the error statistic is an in-sample value of the fitted objective, not an independent benchmark or an out-of-sample test of Theorem 2.13. This is a validation circularity, not part of the derivation of the PDEs, so it does not affect the central claims.
full rationale
The paper's claimed derivation chain is largely self-contained conditional on its published inputs. The function w/ϖ is defined as the value of an inf-control problem, its smoothness is proved in Proposition 2.8 from uniform ellipticity, and the boundary trace V̅ is then shown in Theorem 2.13 to satisfy (2.10)-(2.11): the subsolution direction is a restriction of the interior HJB characterization quoted from [8, Thm 4.2] to test functions independent of p, and the supersolution direction uses the explicit continuous-selection Assumption 2.10 plus the DPP from [5]. None of these steps fits a parameter to the conclusion: the boundary PDE is not assumed as the definition of V̅, and Assumption 2.10 is stated as an assumption, not derived from the target result. The cited [5,6,8] are published peer-reviewed theorems with proofs (including by an overlapping author), so under the review rules they count as independent support. The footnote admitting a flaw in [5, Thm 3.1] is a limitation of a different route, and the present proof does not use that theorem. The convergence results (Prop 3.1, 3.5) are conditional on assumptions (singleton U, Assumption 2.10) that are explicit and not disguised as predictions. The only quasi-circular element is in the numerical section: the reported error measures (4.7)-(4.8), (4.12)-(4.13) are the same PDE residuals used as training losses (4.6), (4.11), so they are in-sample residuals rather than independent error estimates. This does not invalidate the theorem chain, but it should not be read as an external validation. Overall: no significant circularity in the central derivation; one minor non-load-bearing validation statistic.
Axiom & Free-Parameter Ledger
axioms (7)
- standard math Martingale representation theorem for Brownian functionals
- standard math Weak geometric dynamic programming principle
- standard math Interior viscosity characterization of Bouchard-Nutz [8, Theorem 4.2]
- domain assumption Assumption 2.6: G in C^2_b, g bounded and Holder
- domain assumption Assumption 2.7: uniform ellipticity of sigma sigma^T
- domain assumption Assumption 2.10: existence of continuous argmin feedback u_hat and unique strong solution of the closed-loop SDE
- domain assumption Singleton set U(t,x) in Proposition 3.1 and Remark 2.14
read the original abstract
As in Bouchard et al. (2010) and Bouchard and Nutz (2014), we study a utility maximization problem with expectation constraint. We first consider a uniformly elliptic case in which the endogenous state boundary associated with the constraint in expectation is proved to be smooth. This allows one to derive a proper Dirichlet condition for the value function of the optimal control problem on this boundary. We then propose a new truncation argument in the martingale representation of the expectation constraint. This leads to an approximating sequence of auxiliary systems of PDEs for which comparison holds. Convergence to the initial optimal control problem is proved. In the degenerate case, we propose another approximation which consists in adding a small noise term to recover uniformly ellipticity. Convergence is also proved. To the best of our knowledge, it is the first time that a full analysis is performed for such control problems, so as to open the doors to the use of numerical schemes. Numerical resolution in a toy example is performed using neural networks. It is complemented by an estimation of the numerical error, also performed by using a neural network approach.
Figures
Reference graph
Works this paper leans on
-
[1]
Diffusive Limit Approximation of Pure-Jump Optimal Stochastic Control Problems.Journal of Optimization Theory and Applications, 196(1):147–176, January 2023
Marc Abeille, Bruno Bouchard, and Lorenzo Croissant. Diffusive Limit Approximation of Pure-Jump Optimal Stochastic Control Problems.Journal of Optimization Theory and Applications, 196(1):147–176, January 2023
2023
-
[2]
Anagnostopoulos, Juan Diego Toscano, Nikolaos Stergiopulos, and George Em Karniadakis
Sokratis J. Anagnostopoulos, Juan Diego Toscano, Nikolaos Stergiopulos, and George Em Karniadakis. Residual-based attention in physics-informed neural net- works.Computer Methods in Applied Mechanics and Engineering, 421:116805, 2024
2024
-
[3]
Fully nonlinear parabolic equations in two space variables, February
Ben Andrews. Fully nonlinear parabolic equations in two space variables, February
-
[4]
A stochastic target formulation for optimal switching problems in finite horizon.Stochastics: An International Journal of Probability and Stochastics Processes, 81(2):171–197, 2009
Bruno Bouchard. A stochastic target formulation for optimal switching problems in finite horizon.Stochastics: An International Journal of Probability and Stochastics Processes, 81(2):171–197, 2009
2009
-
[5]
Optimal Control under Stochastic Target Constraints.SIAM Journal on Control and Optimization, 48(5):3501–3531, January 2010
Bruno Bouchard, Romuald Elie, and Cyril Imbert. Optimal Control under Stochastic Target Constraints.SIAM Journal on Control and Optimization, 48(5):3501–3531, January 2010
2010
-
[6]
Stochastic Target Problems with Controlled Loss.SIAM Journal on Control and Optimization, 48(5):3123–3150, Jan- uary 2010
Bruno Bouchard, Romuald Elie, and Nizar Touzi. Stochastic Target Problems with Controlled Loss.SIAM Journal on Control and Optimization, 48(5):3123–3150, Jan- uary 2010
2010
-
[7]
First time to exit of a con- tinuous Itˆ o process: General moment estimates andL 1-convergence rate for discrete time approximations.Bernoulli, 23(3):1631–1662, August 2017
Bruno Bouchard, Stefan Geiss, and Emmanuel Gobet. First time to exit of a con- tinuous Itˆ o process: General moment estimates andL 1-convergence rate for discrete time approximations.Bernoulli, 23(3):1631–1662, August 2017. Publisher: Bernoulli Society for Mathematical Statistics and Probability
2017
-
[8]
Weak Dynamic Programming for Generalized State Constraints.SIAM Journal on Control and Optimization, 50(6):3344–3373, January 2012
Bruno Bouchard and Marcel Nutz. Weak Dynamic Programming for Generalized State Constraints.SIAM Journal on Control and Optimization, 50(6):3344–3373, January 2012
2012
-
[9]
Weak dynamic programming principle for viscosity solutions.SIAM Journal on Control and Optimization, 49(3):948–962, 2011
Bruno Bouchard and Nizar Touzi. Weak dynamic programming principle for viscosity solutions.SIAM Journal on Control and Optimization, 49(3):948–962, 2011
2011
-
[10]
Deep learning method based on physics informed neural network with resnet block for solving fluid flow problems.Water, 13(4), 2021
Chen Cheng and Guang-Tao Zhang. Deep learning method based on physics informed neural network with resnet block for solving fluid flow problems.Water, 13(4), 2021
2021
-
[11]
User’s guide to viscos- ity solutions of second order partial differential equations.Bulletin of the American mathematical society, 27(1):1–67, 1992
Michael G Crandall, Hitoshi Ishii, and Pierre-Louis Lions. User’s guide to viscos- ity solutions of second order partial differential equations.Bulletin of the American mathematical society, 27(1):1–67, 1992
1992
-
[12]
A closed-form solution to the problem of super-replication under transaction costs.Finance and stochastics, 3:35–54, 1999
Jaksa Cvitani´ c, Huyen Pham, and Nizar Touzi. A closed-form solution to the problem of super-replication under transaction costs.Finance and stochastics, 3:35–54, 1999
1999
-
[13]
Deep learning approximation for stochastic control prob- lems.ArXiv, abs/1611.07422, 2016
Jiequn Han and Weinan E. Deep learning approximation for stochastic control prob- lems.ArXiv, abs/1611.07422, 2016. 40
Pith/arXiv arXiv 2016
-
[14]
A new formulation of state constraint problems for first-order pdes.SIAM Journal on Control and Optimization, 34(2):554–571, 1996
Hitoshi Ishii and Shigeaki Koike. A new formulation of state constraint problems for first-order pdes.SIAM Journal on Control and Optimization, 34(2):554–571, 1996
1996
-
[15]
Viscosity solutions of second order fully nonlinear elliptic equations with state constraints.Indiana University Mathematics Journal, pages 493– 519, 1994
Markos A Katsoulakis. Viscosity solutions of second order fully nonlinear elliptic equations with state constraints.Indiana University Mathematics Journal, pages 493– 519, 1994
1994
-
[16]
Strong solutions of stochastic equations with singular time dependent drift.Probability theory and related fields, 131(2):154–196, 2005
Nicolai V Krylov and Michael R¨ ockner. Strong solutions of stochastic equations with singular time dependent drift.Probability theory and related fields, 131(2):154–196, 2005
2005
-
[17]
Nonlinear elliptic equations with singu- lar boundary conditions and stochastic control with state constraints: 1
Jean-Michel Lasry and Pierre-Louis Lions. Nonlinear elliptic equations with singu- lar boundary conditions and stochastic control with state constraints: 1. the model problem.Mathematische Annalen, 283:583–630, 1989
1989
-
[18]
Lieberman.Second order parabolic differential equations
Gary M. Lieberman.Second order parabolic differential equations. World Scientific, New Jersey Singapore, repr., [rev. ed.] edition, 2005
2005
-
[19]
McClenny and Ulisses M
Levi D. McClenny and Ulisses M. Braga-Neto. Self-adaptive physics-informed neural networks.Journal of Computational Physics, 474:111722, February 2023
2023
-
[20]
Duality and approximation of stochastic optimal control problems under expectation constraints.SIAM Journal on Control and Optimization, 59(5):3231–3260, 2021
Laurent Pfeiffer, Xiaolu Tan, and Yu-Long Zhou. Duality and approximation of stochastic optimal control problems under expectation constraints.SIAM Journal on Control and Optimization, 59(5):3231–3260, 2021
2021
-
[21]
Raissi, P
M. Raissi, P. Perdikaris, and G.E. Karniadakis. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations.Journal of Computational Physics, 378:686–707, 2019
2019
-
[22]
Prajit Ramachandran, Barret Zoph, and Quoc V. Le. Swish: a Self-Gated Activation Function, October 2017. arXiv:1710.05941 [cs.NE] version: 1
Pith/arXiv arXiv 2017
-
[23]
Mete Soner and Nizar Touzi
H. Mete Soner and Nizar Touzi. Dynamic programming for stochastic target problems and geometric flows.Journal of the European Mathematical Society, 4(3):201–236, September 2002
2002
-
[24]
Optimal control with state-space constraint i.SIAM Journal on Control and Optimization, 24(3):552–561, 1986
Halil Mete Soner. Optimal control with state-space constraint i.SIAM Journal on Control and Optimization, 24(3):552–561, 1986
1986
-
[25]
Optimal control with state-space constraint
Halil Mete Soner. Optimal control with state-space constraint. ii.SIAM journal on control and optimization, 24(6):1110–1122, 1986
1986
-
[26]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need, August 2023. arXiv:1706.03762 [cs.CL]
Pith/arXiv arXiv 2023
-
[27]
On strong solutions and explicit formulas forsolutions of stochastic integral equations.Mathematics of the USSR-Sbornik, 39(3):387, 1981
Alexander Ju Veretennikov. On strong solutions and explicit formulas forsolutions of stochastic integral equations.Mathematics of the USSR-Sbornik, 39(3):387, 1981. 41
1981
-
[28]
An expert’s guide to training physics-informed neural networks.ArXiv, abs/2308.08468, 2023
Sifan Wang, Shyam Sankaran, Hanwen Wang, and Paris Perdikaris. An expert’s guide to training physics-informed neural networks.ArXiv, abs/2308.08468, 2023
Pith/arXiv arXiv 2023
-
[29]
Stochastic flows of sdes with irregular coefficients and stochastic transport equations.Bulletin des sciences mathematiques, 134(4):340–378, 2010
Xicheng Zhang. Stochastic flows of sdes with irregular coefficients and stochastic transport equations.Bulletin des sciences mathematiques, 134(4):340–378, 2010. 42
2010
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.