REVIEW 3 major objections 5 minor 21 references
This paper derives a covariance-controlled upper bound on the Kullback-Leibler divergence between a nonlinear system's true distribution and its Gaussian surrogate, and uses it to design convex steering policies with enforceable risk constr
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 19:32 UTC pith:EE4ZKVAZ
load-bearing objection Useful but over-claimed: the KLD rate bound is a heuristic substitution, not a verified upper bound; deserves peer review with revisions. the 3 major comments →
Approximate Relative Entropy Constraints for Nonlinear Covariance Steering Under Distribution Ambiguity
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that the KLD between the true nonlinear distribution and the Gaussian reference can be controlled through the covariance matrix, the decision variable of covariance steering. Using the Donsker–Varadhan variational identity, true expectations of risk indicators and mean-squared error are bounded by expressions that involve only the surrogate Gaussian plus the current KLD. The paper then bounds the KLD itself: after a Young's inequality on the Fokker–Planck flow, the time-rate of divergence is controlled by the surrogate expectation of the squared drift mismatch, which under a second-order Taylor truncation becomes a plain quadratic function of the covariance. Embe
What carries the argument
The Donsker–Varadhan variational formula (Eq. 6), which converts an exponential surrogate expectation into a supremum over distributions penalized by KLD, is the tool that turns unknown-distribution expectations into computable upper bounds. The second load-bearing object is the KLD-rate bound of Eqs. 20–35: a Young's inequality applied to the Fokker–Planck evolution of the two distributions produces a rate controlled by the second-order dynamics tensor and the covariance, making the divergence a function of the optimizer's own decision variable.
Load-bearing premise
The bound relies on replacing an expectation taken under the true nonlinear distribution by one taken under the Gaussian surrogate (the 'posit' after Eq. 29); if the true and surrogate distributions are not close in the optimized region, the derived bound is not an upper bound and the risk constraints may not hold.
What would settle it
Run the Monte Carlo trials, compute the empirical KLD between the true particle distribution and the Gaussian surrogate at each time node, and compare it to the integrated upper bound from the solved policy. If the empirical KLD exceeds the bound at any node, or if the true risk probability exceeds its computed upper bound, the central claim is refuted.
If this is right
- If the KLD-rate bound holds, risk constraints such as collision probability less than epsilon and MSE bounds become enforceable in a convex program without knowing the true distribution.
- Tighter entropy targets make the solution covariance align with the low-curvature directions of the dynamics (orthogonal to the Hessians), which matches the Monte Carlo observation that constrained policies reduce outlier samples.
- The framework removes the heuristic choice of ambiguity-set radius: the radius is replaced by an integrated bound computed from the trajectory's own nonlinearity measure.
- The approach extends beyond collision probability and MSE to any quantity whose exponential moment under the surrogate has a closed form, for example conditional value-at-risk.
- Because the rate bound is quadratic in the covariance, the entropy constraint is convexifiable in an SCvx loop, so existing covariance-steering solvers can be augmented directly.
Where Pith is reading between the lines
- The paper leaves implicit a direct validation of the KLD bound: comparing the empirical KLD computed from the Monte Carlo particles to the integrated upper bound over time would show how conservative (or not) the approximation actually is.
- The 'large-diffusion' prediction (larger process noise shrinks the rate bound) is counterintuitive and suggests a testable regime where noise actually helps the Gaussian surrogate; a parameter sweep over sigma could confirm this.
- The assumption that the user knows the true drift exactly limits the result to linearization error, not model error; an extension would fold parametric uncertainty into the drift mismatch term.
- Replacing the Euler integration of the KLD-rate bound with a more accurate quadrature may change the tightness of the risk bounds in high-nonlinearity segments.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses nonlinear covariance steering under distribution ambiguity by measuring the mismatch between the true state distribution and a Gaussian surrogate using the Kullback-Leibler divergence (KLD). It derives upper bounds on collision probability and mean-squared error via the Donsker-Varadhan variational formula, conditional on the KLD between the true and surrogate distributions. It then proposes an approximate upper bound on the time-derivative of this KLD that depends only on the covariance and the Hessians of the dynamics, which is compatible with covariance steering. The approach is embedded into a successive convex programming (SCvx) algorithm and demonstrated on an Earth-moon near-rectilinear halo orbit transfer, with Monte Carlo results showing improved terminal covariance consistency and MSE compliance.
Significance. The contribution is practically relevant: it offers a tractable way to enforce distributional robustness and risk constraints in nonlinear covariance steering, and the paper includes a complete algorithm and substantial numerical experiments (2500-trial Monte Carlo analyses) on a challenging astrodynamics problem. The variational bounds themselves are standard, and the central novelty is the covariance-controllable approximate KLD-rate bound. However, this bound is not proven; the key replacement of an expectation under the true distribution by one under the surrogate is stated as a 'posit' and is not validated empirically. Consequently, the claimed risk-sensitive guarantees are conditional on an unverified approximation. If the approximation holds in the optimized regime, the framework is a useful engineering tool; the theoretical standing of the 'upper bound' claims is weaker than the language suggests.
major comments (3)
- [Relative Entropy Steering, Eq. (20) and paragraph after Eq. (29)] The transition from Eq. (29) to Eq. (20) replaces E_{π_t} with E_{π̂_t} for the integrand ||D^{-1/2}(φ−f̂)||². The paper calls this a 'posit' and justifies it by assuming R(π_t||π̂_t) is small. However, the algorithm constrains R_ub, not the true KLD R, so the justification is self-referential: the controlled quantity is the bound, not the distance that would justify the substitution. This is load-bearing because Eqs. (13) and (18) use R_ub in place of R to enforce risk bounds. Without a quantitative error bound (e.g., in terms of R or total variation) or a direct Monte Carlo comparison between R_ub and an empirical estimate of R(π_t||π̂_t), Eq. (20) may not be an upper bound in the optimized region. Neither Fig. 7 nor Fig. 9 tests this step.
- [Simplification with a Gaussian Reference, Eqs. (30)–(35)] The second-order Taylor expansion in Eq. (30) neglects O(||δx||³) terms with no error control. Moreover, Eq. (32) evaluates the expectation under the Gaussian surrogate π̂_t, while the true distribution π_t is non-Gaussian. Thus Eq. (35) is an approximation even beyond the substitution in Eq. (20). The claim that Eq. (35) controls the KLD rate is therefore conditional on unquantified linearization-error terms and on the surrogate expectation being close to the true expectation. Please provide a bound on the neglected third-order terms, or demonstrate numerically in the optimized regime that they are negligible, e.g., by comparing the true drift mismatch with the second-order surrogate.
- [Convexification of the MSE Variational Upper Bound, Eq. (64)] The derivative formula for ∂ψ_k/∂λ_k contains the term −tr[(I−λ̄_k P̄_k)^{-1} P̄_k]. From the definition of ψ in Eq. (61), this term should be −tr[(I−2λ̄_k P̄_k)^{-1} P̄_k], which would cancel the first term. As written, the formula is inconsistent with Eq. (61) and makes the SCvx implementation irreproducible. Please correct the typo and verify the subsequent convergence results.
minor comments (5)
- [Eq. (24)] In the Fokker-Planck equation, the diffusion term is written as (1/2)∇·(D∇π_t), but it should be (1/2)∇·(D∇q_t) for a general distribution q_t.
- [Eqs. (14) and (19)] The limit statements as R→0 and λ*→0 are referenced to Ref. [17] without proof. A short derivation or a clear statement of the required regularity conditions would make the paper more self-contained.
- [Abstract and Introduction] The KLD-rate bound is repeatedly called an 'upper bound', but Eq. (20) is only an approximate bound. Please qualify the terminology (e.g., 'approximate upper bound') to avoid overclaiming.
- [After Eq. (35)] The claim that the large-diffusion regime makes the Gaussian surrogate 'more accurate' is based solely on the σ^{-2} dependence of the bound; this does not necessarily reflect the behavior of the true KLD. Please temper the interpretation or provide supporting evidence.
- [Fig. 8] The caption mentions 'red and blue points' but the legend uses 'iteration minimum' and 'empirical minimum'. Please align the caption and legend.
Circularity Check
Load-bearing KLD-rate bound rests on a self-referential posit and an unverified self-citation.
specific steps
-
other
[Relative Entropy Steering, Eq. 20 / paragraph after Eq. 29]
"Here, however, we posit that, provided R(π_t∥π̂_t) is constrained to be small, Eq. 20 provides a useful approximate upper bound for the distance time rate of change. That is, we can evaluate this expectation under our approximate distribution."
Eq. 29 gives the exact dR/dt ≤ ½E_π_t[||D^{-1/2}(φ−f̂)||²]. Eq. 20 replaces this with ½E_π̂_t[...]. The only stated justification is that R(π||π̂) is small. But the SCvx constraints (Eqs. 68d and 55) enforce smallness of R_ub, not of R. Thus R_ub's validity as an upper bound is made conditional on the very quantity R that R_ub is supposed to control, and the algorithm never enforces that condition. No error bound or empirical KLD-vs-R_ub check is supplied, so the 'approximate upper bound' is not shown to bound the true KLD in the optimized region.
-
self citation load bearing
[Relative Entropy Steering, Eq. 20 / after Eq. 29]
"The derivation of the inequality in Eq. 20 is introduced in our previous work, 17 and is similar to that originally provided in. 18 ... A formal treatment of this problem is discussed in our previous work. 17"
The load-bearing step—replacing E_π_t by E_π̂_t to turn the exact Eq. 29 inequality into the computable Eq. 20 constraint—is not derived in this manuscript. The only reference given for a 'formal treatment' is Ref. [17], a paper by the same two authors. The external Ref. [18] is cited for similarity but the manuscript does not claim it justifies the surrogate-expectation replacement. Hence the central premise of the method rests on an unverified self-citation rather than on a proof contained in or independently checkable from this paper.
full rationale
The exact variational bounds (Eqs. 6–18) are self-contained and the SCvx implementation plus Monte Carlo tests provide independent engineering content. However, the central transfer from Eq. 29 to Eq. 20—the step that converts an exact KLD-rate inequality into a covariance-only constraint—is not derived. It is introduced as a posit whose validity requires R small, while the optimizer only constrains R_ub, the very surrogate quantity whose legitimacy as a bound depends on R being small. The manuscript defers the formal treatment to the authors' own Ref. [17], and the Monte Carlo results in Figs. 7 and 9 check terminal covariance and MSE but never compare the empirical KLD to R_ub. This makes the claim that the true distribution is kept close to the surrogate partially circular: the controlled surrogate is used as its own justification. Because the remaining framework has substantial independent content and the paper is transparent about the approximation, the overall circularity is partial rather than complete.
Axiom & Free-Parameter Ledger
free parameters (1)
- γ (SCvx penalty weight)
axioms (7)
- ad hoc to paper Eπt[h] ≈ Eπ̂t[h] for h = ||D^{-1/2}(φ − f̂)||² (equivalently, Eq. 20 is a usable upper bound on dR/dt)
- domain assumption No model error: the true drift equals the user's nonlinear model φt = ft
- ad hoc to paper Second-order truncation of the linearization error (higher-order terms neglected)
- domain assumption Gaussian surrogate π̂t with covariance P_t
- domain assumption Spherical process noise D = σ² I
- domain assumption Initial true distribution equals the surrogate, π0 = π̂0
- standard math Donsker-Varadhan variational identity (Eq. 6) and the resulting bound Eq. 8
read the original abstract
Covariance steering provides an efficient framework for designing linear stochastic feedback policies, but its extension to nonlinear systems relies on a Gaussian surrogate obtained through local linearization. Because this surrogate may differ substantially from the true nonlinear state distribution, risk-sensitive quantities such as collision probability and mean-squared error may be inaccurately estimated. This work develops a distributionally robust covariance-steering framework based on the relative entropy, also known as the Kullback-Leibler divergence (KLD), to account for ambiguity in the propagated probability density function. Using a variational representation of exponential integrals, we derive computable upper bounds on risk-sensitive quantities over a KLD ambiguity set. We then formulate an upper bound on the time rate of change of the KLD between the true nonlinear distribution and a Gaussian reference surrogate. Under some assumptions, this bound is controlled by decision variables within a covariance-steering formulation. The resulting constraints are incorporated into a sequential convex programming algorithm to design stochastic guidance policies that keep the true distribution close to its Gaussian surrogate while enforcing bounds on risk-sensitive performance measures. The proposed approach is demonstrated on a challenging nonlinear spacecraft transfer between two near-rectilinear halo orbits.
Reference graph
Works this paper leans on
-
[1]
Optimal Covariance Control for Stochastic Systems Un- der Chance Constraints,
K. Okamoto, M. Goldshtein, and P. Tsiotras, “Optimal Covariance Control for Stochastic Systems Un- der Chance Constraints,”IEEE Control Systems Letters, V ol. 2, No. 2, 2018, pp. 266–271, 10.1109/LC- SYS.2018.2826038
arXiv 2018
-
[2]
K. Oguri and J. W. McMahon, “Robust Spacecraft Guidance Around Small Bodies Under Uncertainty: Stochastic Optimal Control Approach,”Journal of Guidance, Control, and Dynamics, V ol. 44, No. 7, 2021, pp. 1295–1313, 10.2514/1.G005426
-
[3]
Optimal Transport in Systems and Control,
Y . Chen, T. T. Georgiou, and M. Pavon, “Optimal Transport in Systems and Control,”Annual Review of Control, Robotics, and Autonomous Systems, V ol. 4, No. 1, 2021, pp. 89–113, 10.1146/annurev-control- 070220-100858
-
[4]
F. Liu and P. Tsiotras, “Optimal Covariance Steering for Continuous-Time Linear Stochastic Systems With Additive Martingale Noise,”IEEE Transactions on Automatic Control, V ol. 69, No. 4, 2024, pp. 2591–2597, 10.1109/TAC.2023.3331683
arXiv 2024
-
[5]
Optimal Covariance Steering for Discrete-Time Linear Stochastic Systems,
F. Liu, G. Rapakoulias, and P. Tsiotras, “Optimal Covariance Steering for Discrete-Time Linear Stochastic Systems,” 2024
2024
-
[6]
Non-Gaussian Distribution Steering in Nonlinear Dynamics with Conjugate Unscented Transformation,
D. C. Qi, K. Oguri, P. Singla, and M. R. Akella, “Non-Gaussian Distribution Steering in Nonlinear Dynamics with Conjugate Unscented Transformation,” 2025
2025
-
[7]
Non-Gaussian Chance-Constrained Trajectory Control Using Gaussian Mixtures and Risk Allocation,
S. Boone and J. McMahon, “Non-Gaussian Chance-Constrained Trajectory Control Using Gaussian Mixtures and Risk Allocation,”2022 IEEE 61st Conference on Decision and Control (CDC), 2022, pp. 3592–3597, 10.1109/CDC51059.2022.9993274
arXiv 2022
-
[8]
Chance-Constrained Gaussian Mixture Steering to a Terminal Gaussian Distribution,
N. Kumagai and K. Oguri, “Chance-Constrained Gaussian Mixture Steering to a Terminal Gaussian Distribution,” 2024
2024
-
[9]
Ambiguous Chance Constrained Problems and Robust Optimization,
E. Erdogan and G. Iyengar, “Ambiguous Chance Constrained Problems and Robust Optimization,” Mathematical Programming, V ol. 107, No. 1, 2006, pp. 37–61, 10.1007/s10107-005-0678-0
-
[10]
Wasserstein Tube MPC with Exact Uncertainty Propagation,
L. Aolaritei, M. Fochesato, J. Lygeros, and F. D ¨orfler, “Wasserstein Tube MPC with Exact Uncertainty Propagation,” 2023
2023
-
[11]
Distributionally Robust Density Control with Wasserstein Ambiguity Sets,
J. Pilipovsky and P. Tsiotras, “Distributionally Robust Density Control with Wasserstein Ambiguity Sets,” 2024
2024
-
[12]
A weak convergence approach to the theory of large deviations,
P. Dupuis and R. S. Ellis, “A weak convergence approach to the theory of large deviations,” 1997
1997
-
[13]
A. Dembo and O. Zeitouni,Large Deviations Techniques and Applications, V ol. 38 ofStochastic Mod- elling and Applied Probability. New York: Springer-Verlag, 2nd ed., 1998, 10.1007/978-1-4612-5320- 4
-
[14]
On Information and Sufficiency,
S. Kullback and R. A. Leibler, “On Information and Sufficiency,”The Annals of Mathematical Statistics, V ol. 22, No. 1, 1951, pp. 79–86, 10.1214/aoms/1177729694
arXiv 1951
-
[15]
Stochastic Atmosphere Modeling for Risk Adverse Aerocapture Guid- ance,
J. Ridderhof and P. Tsiotras, “Stochastic Atmosphere Modeling for Risk Adverse Aerocapture Guid- ance,”2020 IEEE Aerospace Conference, IEEE, 2020, 10.1109/AERO47225.2020.9172724
arXiv 2020
-
[16]
Chance-Constrained Covariance Steering in a Gaussian Random Field via Successive Convex Programming,
J. Ridderhof and P. Tsiotras, “Chance-Constrained Covariance Steering in a Gaussian Random Field via Successive Convex Programming,”Journal of Guidance, Control, and Dynamics, V ol. 45, No. 4, 2022, pp. 599–610, 10.2514/1.G005941
-
[17]
Relative Entropy-Bounded Ambiguous Chance Constraints for Ro- bust Planning in Nonlinear Systems,
T. N. Wolf and J. W. McMahon, “Relative Entropy-Bounded Ambiguous Chance Constraints for Ro- bust Planning in Nonlinear Systems,”Proceedings of the 2026 European Control Conference (ECC), Reykjav´ık, Iceland, July 2026
2026
-
[18]
Improved bounds for discretization of Langevin diffusions: Near-optimal rates without convexity,
W. Mou, N. Flammarion, M. J. Wainwright, and P. L. Bartlett, “Improved bounds for discretization of Langevin diffusions: Near-optimal rates without convexity,”Bernoulli, V ol. 28, No. 3, 2022, pp. 1577– 1601
2022
-
[19]
The moments of products of quadratic forms in normal variables,
J. R. Magnus, “The moments of products of quadratic forms in normal variables,”Statistica Neer- landica, V ol. 32, No. 4, 1978, pp. 201–210, https://doi.org/10.1111/j.1467-9574.1978.tb01399.x
arXiv 1978
-
[20]
Y . Mao, M. Szmuk, X. Xu, and B. Acikmese, “Successive Convexification: A Superlinearly Convergent Algorithm for Non-convex Optimal Control Problems,”arXiv e-prints, Apr. 2018, p. arXiv:1804.06539, 10.48550/arXiv.1804.06539
-
[21]
D. Malyuta, T. P. Reynolds, M. Szmuk, T. Lew, R. Bonalli, M. Pavone, and B. Ac ¸ıkmes ¸e, “Convex Optimization for Trajectory Generation: A Tutorial on Generating Dynamically Feasible Trajecto- ries Reliably and Efficiently,”IEEE Control Systems Magazine, V ol. 42, No. 5, 2022, pp. 40–113, 10.1109/MCS.2022.3187542. 19
arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.