REVIEW 2 major objections 4 minor 30 references
Backward Stochastic Control System with Entropy Regularization
T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper proves that in the entropy-regularized backward linear-quadratic control problem with a standard normal reference measure, the unique optimal relaxed control is Gaussian and is given explicitly by its covariance and mean.
desk verdict First entropy-regularized maximum principle and LQ result for backward stochastic control; correct in architecture but needs a dimensions-fix in the Riccati equation and a cleanup of Remark 5.3. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the admissible set $\mathcal{A}$ of measure-valued relaxed controls with finite relative entropy, the flat (linear functional) derivative $\delta F/\delta m$ used to differentiate the cost and generator with respect to measures, and the stochastic Hamiltonian system (5.3), which couples the controlled BSDE for $(Y^\mu,Z^\mu)$, the adjoint SDE for $P^\mu$, and the first-order optimality condition that pins down the control density. The existence proof for the LQ case runs through the decoupling ansatz $Y_t^\mu = \Theta_t P_t^\mu + \varphi_t$, which turns the Hamiltonian system into a Riccati equation for $\Theta$ and a linear BSDE for $\varphi$; the Gaussian form of the control then follows by completing the square in the exponent of the Gibbs density.
What would settle it
Take $n=2$, $p=1$ and set $B_t$ to a nonzero two-dimensional column vector with $A=C=H=N=0$. Then the Riccati equation (5.8) as printed adds a vector to $2\times2$ matrix terms, so no $2\times2$ solution $\Theta$ can exist; testing whether the corrected term $B_t(R_t+(\sigma^2/2)I_p)^{-1}B_t^\top$ gives a unique solution would settle whether the claimed uniqueness is recoverable.
Extended reading notes
Core claim
The central claim is that entropy-regularized exploratory control of BSDEs is not merely a formal device: in the linear-quadratic case with standard normal reference measure, the optimal relaxed control exists, is unique, and is Gaussian. The paper derives this by combining a convex-variation maximum principle with the decoupling ansatz $Y_t^\mu = \Theta_t P_t^\mu + \varphi_t$, which reduces the stochastic Hamiltonian system to a Riccati equation for $\Theta$ and a linear BSDE for $\varphi$. The resulting optimal policy has covariance $\Sigma_t^\mu = (\sigma^2/2)(R_t + (\sigma^2/2)I_p)^{-1}$ and mean $v_t^\mu = -(R_t + (\sigma^2/2)I_p)^{-1}B_tP_t^\mu$, and as $\sigma \to 0$ it concentrates on the strict control $v_t^\mu$ with exploration cost $\sigma^2 p T/4$. The paper also gives the implicit form of the optimal control for the general non-LQ problem as a fixed point of a Gibbs-type density, and proves the maximum principle by convex variation for systems with random coefficients.
Load-bearing premise
The proof that an optimal control exists in the LQ case depends on the decoupling assumption $Y = \Theta P + \varphi$ and on the Riccati equation having a unique bounded solution; the printed equation has a term whose dimensions do not match the others, so this existence claim rests on a correction or a fuller derivation.
Editorial extensions
If this is right
- For the LQ problem with standard normal reference, the optimal exploration policy is fully explicit: a Gaussian with covariance $(\sigma^2/2)(R_t + (\sigma^2/2)I_p)^{-1}$, so exploration shrinks as the entropy weight $\sigma$ decreases or the running cost weight $R$ increases.
- The cost of exploration is $\sigma^2 pT/4$, and as $\sigma \to 0$ the relaxed optimal control converges weakly to a Dirac measure at the strict control, recovering the classical backward LQ optimum.
- The necessary and sufficient maximum principle gives a concrete fixed-point equation for the optimal control density in the non-LQ case, which can serve as the basis for successive-approximation algorithms.
- With convexity of the Hamiltonian, any control satisfying the variational inequality (3.17) is optimal, so the condition is necessary and sufficient in the LQ setting.
- The unique solvability of the stochastic Hamiltonian system provides a clean characterization of the optimal triple $(Y^\mu,Z^\mu,P^\mu)$ together with the Gaussian control density.
Reading between the lines
- A natural testable extension is to replace the standard normal reference by a general log-concave $e^{-U}$; the Gibbs form (4.7) suggests the optimal control remains Gaussian only when $U$ is quadratic, with covariance depending on the curvature of $U$.
- Because the maximum principle handles random coefficients, the same entropy-regularized backward framework could be applied to hedging or portfolio replication under Knightian uncertainty, where the relaxed control represents a family of beliefs rather than a single strategy.
- The decoupling method in Section 5 gives a template for proving existence and uniqueness in related exploratory backward problems, such as mean-field or time-inconsistent objectives; the key step would be a well-posed Riccati equation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies an entropy-regularized backward stochastic control problem in which the state is a controlled BSDE (1.1) and the cost functional contains a relative-entropy penalty with respect to a reference measure e^{-U}, as in (1.2). The control is measure-valued (relaxed), and the authors prove a stochastic maximum principle by the convex variation method (Theorem 3.7), a sufficient condition (Theorem 4.2), and an implicit form of the optimal relaxed control (4.7). In the linear-quadratic case, with standard normal reference measure, the paper claims that the stochastic Hamiltonian system (5.3) has a unique solution, that the optimal control is Gaussian with covariance Sigma_t^mu = (sigma^2/2)(R_t + (sigma^2/2)I_p)^{-1} and mean v_t^mu = -(R_t + (sigma^2/2)I_p)^{-1} B_t^T P_t^mu, and that this control is the unique optimal control (Theorem 5.5 and Corollary 5.6). The proof uses a decoupling ansatz Y_t = Theta_t P_t + phi_t and a Riccati equation for Theta.
Significance. If the linear-quadratic result is correct, the paper provides the first explicit form of an entropy-regularized optimal relaxed control for a backward stochastic system, thereby extending the exploratory-control methodology of Wang, Zariphopoulou and Zhou (2020) and Siska and Szpruch (2024) from forward to backward dynamics. The strengths of the paper are the detailed variation estimates, the derivation of the Gaussian form from the first-order condition rather than from an ad hoc ansatz, and the Itô-based uniqueness argument. However, the central LQ existence/uniqueness claim is not supported as printed because Eq. (5.8), the Riccati equation on which Theorem 5.5 and Corollary 5.6 depend, is dimensionally inconsistent. The issue appears to be a repairable transposition/dimensional typo, so the manuscript should be revised rather than rejected.
major comments (2)
- [§5, Eq. (5.8)] Eq. (5.8) is dimensionally inconsistent as printed. With B_t in L^infty(0,T; R^{n x p}) and R_t in S^p_+, the term (R_t + (sigma^2/2) I_p)^{-1} B_t is not an n x n matrix: the product is undefined if B_t is n x p, and it is p x n if B_t is read as B_t^T under the paper's transpose-omitting convention. Re-deriving the coefficient of P_t in the lambda equation from (5.3) and (5.5) gives B_t (R_t + (sigma^2/2) I_p)^{-1} B_t^T P_t, with one B on each side; the correct Riccati term should therefore be B_t (R_t + (sigma^2/2) I_p)^{-1} B_t^T, not (R_t + (sigma^2/2) I_p)^{-1} B_t. Since the claimed unique Theta, the explicit Gaussian formulas in Theorem 5.5, and the uniqueness argument in Corollary 5.6 all rest on this equation, the existence and uniqueness results are not proven as stated. The authors should correct the equation and either supply a proof or a precise citation for the existence of a bounded solution Theta for the corrected equation.
- [§5, Remark 5.3] Remark 5.3 takes U identically equal to 0, but Definition 2.1 requires e^{-U} to be a density on R^p; e^0 = 1 is not integrable over R^p, so the reference measure e^{-U} is not a probability density in that case. The cost-of-exploration computation in Remark 5.3 is therefore outside the framework defined in Section 2. This does not affect the main theorem, which uses the standard normal reference measure, but the remark should be corrected, for example by explicitly extending the definition to allow sigma-finite reference measures or by normalizing U appropriately.
minor comments (4)
- [§4, Eq. (4.4)] The assertion that the pointwise variational inequality (3.17) is equivalent to the minimization problem (4.4) requires convexity of H^sigma in m. As printed, the equivalence appears immediately after Theorem 4.2, but the text should explicitly state that this step uses Assumption 4.1; without that assumption, (3.17) is only a first-order variational inequality.
- [§5, Eqs. (5.5), (5.10)] In several places in Section 5, the notation B_t P_t should be B_t^T P_t (and P_t B_t should be P_t^T B_t). The paper states that transpose symbols are omitted unless necessary, but this convention is especially dangerous in Section 5, where matrix dimensions and quadratic forms are central; the dimensional error in Eq. (5.8) is partly a product of this convention. The authors should clarify the convention at the start of Section 5 or restore the transposes in all dimension-sensitive formulas.
- [§4, Eq. (4.5)] In Eq. (4.5), the argument l(t,Y_t^mu,Z_t^mu,P_t^mu,m) includes P_t^mu, but the running cost l does not depend on the adjoint variable; this should be l(t,Y_t^mu,Z_t^mu,m).
- [Throughout] There are numerous typographical issues, including 'theoretical depict' in the abstract and inconsistent spelling of author names in the references; a careful language and proofreading pass is needed.
Circularity Check
No circularity: the LQ Gaussian optimal control is derived from first-order stationarity, not assumed; imported lemmas are external, and no fitted input is renamed as a prediction.
full rationale
The paper's main derivation chain is self-contained rather than circular. In Section 3, the stochastic maximum principle is obtained by a convex-variation calculation: the variation equation (3.2), the adjoint equation (3.14), and the variation inequality in Propositions 3.5 and 3.6 are all proved in the text. The only imported technical tool there is Lemma 3.3, quoted from Siska and Szpruch [26], which is external prior work with no author overlap; it is used for a convexity/entropy inequality, not to pre-impose the form of the optimal control. In the LQ case, the Gaussian form of the optimal control in Eq. (5.5) is computed by completing the square in the first-order condition (5.4) with the standard normal reference measure; the Gaussian law is the output of the algebra, not an input or an ansatz. Existence and uniqueness in Theorem 5.5 proceed by constructing a candidate from the decoupling form Y=Theta P+phi, verifying that it satisfies the Hamiltonian system, and then proving uniqueness by applying Ito's formula to the difference of two solutions; the uniqueness argument does not invoke the decoupling ansatz or the Riccati equation, so the central uniqueness claim is independent of any imported theorem. The Riccati-existence statement is attributed to Lim and Zhou [20], an external source, not a self-citation, and the paper's authors do not rely on their own prior work for the load-bearing existence step. I do note the apparent dimensional inconsistency in the printed Riccati equation (5.8), where the term (R_t+sigma^2/2 I_p)^{-1}B_t cannot be an n x n matrix when B_t is n x p; this is a correctness or typographical gap affecting the existence leg of Theorem 5.5, but it is not circularity because it does not make any conclusion reduce by definition to its own input. No fitted parameter is renamed as a prediction, and no known result is merely relabeled. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (6)
- standard math Backward SDE solution existence and uniqueness for Lipschitz generators, from Pardoux and Peng [22].
- domain assumption Flat derivative calculus and convexity of the admissible control set, imported from Kerimkulov et al. [18] and Carmona-Delarue [4].
- domain assumption Entropy variational inequalities stated in Lemma 3.3, taken from Siska and Szpruch [26].
- domain assumption The reference measure e^{-U} is a density and satisfies coercivity and growth conditions in Section 4.
- domain assumption The Riccati equation (5.8) has a unique bounded solution, by citation to Lim and Zhou [20].
- domain assumption Strong duality and Slater's condition for the measure-valued convex optimization problem in Section 4.
Cite this review
Pith. "Pith review of Backward Stochastic Control System with Entropy Regularization." pith.science (2026). https://pith.science/paper/P45SPXMB
@misc{pith2026241113219,
author = {Pith},
title = {Pith review of: Backward Stochastic Control System with Entropy Regularization},
year = {2026},
howpublished = {\url{https://pith.science/paper/P45SPXMB}},
note = {Machine review of arXiv:2411.13219}
}
read the original abstract
The entropy regularization is inspired by information entropy from machine learning and the ideas of exploration and exploitation in reinforcement learning, which appears in the control problem to design an approximating algorithm for the optimal control. This paper is concerned with the optimal exploratory control for backward stochastic system, generated by the backward stochastic differential equation and with the entropy regularization in its cost functional. We give the theoretical depict of the optimal relaxed control so as to lay the foundation for the application of such a backward stochastic control system to mathematical finance and algorithm implementation. For this, we first establish the stochastic maximum principle by convex variation method. Then we prove sufficient condition for the optimal control and demonstrate the implicit form of optimal control. Finally, the existence and uniqueness of the optimal control for backward linear-quadratic control problem with entropy regularization is proved by decoupling techniques.
Reference graph
Works this paper leans on
- [28]
-
[26]
D. ˇSiˇska and L. Szpruch , Gradient flows for regularized stochastic control problems , SIAM Journal on Control and Optimization, 62 (2024), pp. 2036–20 70
work page 2024
-
[12]
X. Gao, Z. Q. Xu, and X. Y. Zhou , State-Dependent temperature control for Langevin diffu- sions, SIAM Journal on Control and Optimization, 60 (2022), pp. 12 50–1268
work page 2022
-
[1]
H. Becker and V. Mandrekar , On the existence of optimal random controls , J. Math. Mech., 18 (1968/69), pp. 1151–1166
work page 1968
-
[2]
R. Buckdahn, J. Li, S. Peng, and C. Rainer , Mean-field stochastic differential equations and associated pdes , The Annals of Probability, 45 (2017), pp. 824–878
work page 2017
-
[3]
R. Carmona and F. Delarue , Forward-backward stochastic differential equations and co n- trolled mckean cvlasov dynamics , The Annals of Probability, 43 (2015), pp. 2647–2700
work page 2015
-
[4]
, Probabilistic Theory of Mean Field Games with Applications I, Springer, 2018
work page 2018
-
[5]
M. Dai, Y. Dong, Y. Jia, and X. Y. Zhou , Learning merton s strategies in an incomplete mar- ket: recursive entropy regularization and biased gaussian exploration, arXiv:2312.11797, 43pp
Show all 30 references
-
[6]
Dokuchaev and X
N. Dokuchaev and X. Y. Zhou , Stochastic controls with terminal contingent conditions , Journal of Mathematical Analysis and Applications, 238 (19 99), pp. 143–165
-
[7]
Duffie and L
D. Duffie and L. G. Epstein , Stochastic differential utility , Econometrica, 60 (1992), pp. 353– 394
1992
-
[8]
El Karoui, D
N. El Karoui, D. H ˙u ˙u Nguyen, and M. Jeanblanc-Picqu ´e, Compactification methods in the control of degenerate diffusions: Existence of an optimal co ntrol, Stochastics, 20 (1987), pp. 169–219
1987
-
[9]
El Karoui, S
N. El Karoui, S. Peng, and M. C. Quenez , Backward stochastic differential equations in finance, Mathematical Finance, 7 (1997), pp. 1–71. 24
1997
-
[10]
Firoozi and S
D. Firoozi and S. Jaimungal , Exploratory LQG mean field games with entropy regularizatio n, Automatica, 139 (2022), Paper No. 110177, 12pp
2022
-
[11]
W. H. Fleming , Generalized solutions in optimal stochastic control , in Differential games and control theory, II (Proc. 2nd Conf., Univ. Rhode Island, Kin gston, R.I., 1976), vol. 30 of Lect. Notes Pure Appl. Math., Dekker, New York, 1977, pp. 147 –165
1976
-
[13]
D. A. Gomes and E. V aldinoci , Entropy penalization methods for Hamilton Jacobi equation s, Advances in Mathematics, 215 (2007), pp. 94–152
2007
-
[14]
K. Hu, Z. Ren, D. ˇSiˇska, and L. Szpruch , Mean-field Langevin dynamics and energy land- scape of neural networks , Annales de l’Institut Henri Poincar´ e, Probabilit´ es et Statistiques, 57 (2021), pp. 2043 – 2065
2021
-
[15]
Jia and X
Y. Jia and X. Y. Zhou , q-learning in continuous time , Journal of Machine Learning Research, 24 (2023), pp. 1–61
2023
-
[16]
Jordan, D
R. Jordan, D. Kinderlehrer, and F. Otto , The variational formulation of the Fokker– Planck Equation , SIAM Journal on Mathematical Analysis, 29 (1998), pp. 1–17
1998
-
[17]
Karnam, J
C. Karnam, J. Ma, and J. Zhang , Dynamic approaches for some time-inconsistent optimiza- tion problems , The Annals of Applied Probability, 27 (2017), pp. 3435–347 7
2017
-
[18]
Kerimkulov, D
B. Kerimkulov, D. ˇSiˇska, L. Szpruch, and Y. Zhang , Mirror descent for stochastic control problems with measure-valued controls , arXiv:2401.01198, 37pp
-
[19]
Klenke , Probability theory, Springer, 2020
A. Klenke , Probability theory, Springer, 2020
2020
-
[20]
Lim and X
A. Lim and X. Y. Zhou , Linear-Quadratic control of backward stochastic different ial equations, SIAM Journal on Control and Optimization, 40 (2001), pp. 450 –474
2001
-
[21]
Lions , Cours au coll ge de france: Th orie des jeu champs moyens
P. Lions , Cours au coll ge de france: Th orie des jeu champs moyens. , Available at http://www.college-de-france.fr/default/EN/all/equ[1]der/audiovideo.jsp, (2013)
2013
-
[22]
Pardoux and S
E. Pardoux and S. Peng , Adapted solution of a backward stochastic differential equa tion, System Control Letters, 14 (1990), pp. 55–61
1990
-
[23]
Peng , Probabilistic interpretation for systems of quasilinear p arabolic partial differential equations, Stochastics and Stochastics Reports, 37 (1991), pp
S. Peng , Probabilistic interpretation for systems of quasilinear p arabolic partial differential equations, Stochastics and Stochastics Reports, 37 (1991), pp. 61–74
1991
-
[24]
, Backward stochastic differential equations and applicatio ns to optimal control , Applied Mathematics and Optimization, 27 (1993), pp. 125–144
1993
-
[25]
Reisinger and Y
C. Reisinger and Y. Zhang , Regularity and stability of feedback relaxed controls , SIAM Jour- nal on Control and Optimization, 59 (2021), pp. 3118–3151
2021
-
[27]
Takahashi, Y
A. Takahashi, Y. Tsuchida, and T. Yamada , A new efficient approximation scheme for solving high-dimensional semilinear pdes: control variat e method for deep bsde solver , Journal of Computational Physics, 454 (2022), Paper No. 110 956, 39pp
2022
-
[29]
W ang and X
H. W ang and X. Y. Zhou , Continuous-time mean-variance portfolio selection: A rei nforce- ment learning framework , Mathematical Finance, 30 (2020), pp. 1273–1308
2020
-
[30]
Zhang and H
Q. Zhang and H. Zhao , Stationary solutions of SPDEs and infinite horizon BDSDEs , Journal of Functional Analysis, 252 (2007), pp. 171–219. 25
2007
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.