REVIEW 1 major objections 1 cited by
Wasserstein robustness makes robust value functions Lipschitz when rewards and transitions are Lipschitz.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 00:20 UTC pith:DOHJJOXL
load-bearing objection Wasserstein robustness produces Lipschitz value functions from Lipschitz primitives where the non-robust case does not, but the claim needs checking on moment conditions for unbounded spaces. the 1 major comments →
Lipschitz Regularity in Wasserstein Robust Stochastic Optimal Control
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In the kernel-robust model the adversary selects next-state distributions inside a Wasserstein ball around the nominal kernel; in the noise-robust model it perturbs the driving noise inside a Wasserstein ball. Under Lipschitz assumptions on the reward and on the transition map, the resulting robust value function is Lipschitz continuous on the state space. The same assumptions do not guarantee Lipschitz continuity of the ordinary value function, so the Wasserstein constraint itself supplies the extra regularity.
What carries the argument
Wasserstein ball centered on the nominal transition kernel or driving noise, which defines the set of allowable adversarial perturbations and thereby regularizes the robust Bellman operator.
Load-bearing premise
Robustness is introduced exactly through Wasserstein balls around a nominal kernel or noise distribution, together with the infinite-horizon discounted criterion on Polish spaces.
What would settle it
An explicit pair of Lipschitz reward and transition functions on a Polish space such that the associated robust value function under either Wasserstein formulation fails to be Lipschitz.
If this is right
- Lipschitz robust value functions permit stable discretization and approximation on unbounded spaces.
- The regularity directly supports value-function estimation and learning algorithms for robust policies.
- Wasserstein robustness supplies both protection against misspecification and regularization of the Bellman fixed point.
- The same Lipschitz conclusion holds for both the kernel-robust and noise-robust formulations.
Where Pith is reading between the lines
- The regularization property may extend to other optimal-transport distances or to finite-horizon problems.
- Convergence rates of approximate dynamic programming could improve under these robust formulations.
- Similar Lipschitz inheritance might appear in distributionally robust versions of other sequential decision problems.
- The result suggests a route to robust continuous-state reinforcement learning with provable approximation guarantees.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims that for Wasserstein-robust stochastic optimal control on Polish state spaces under the infinite-horizon discounted criterion, Lipschitz continuity of the reward function and the nominal transition kernel (in either the kernel-robust formulation with a Wasserstein ball around the nominal kernel or the noise-robust formulation with a Wasserstein ball around the driving noise) implies that the robust value function is Lipschitz. This regularization effect is contrasted with the non-robust case, where the same assumptions do not generally yield a Lipschitz value function.
Significance. If the central claim holds, the result is significant because it identifies a concrete mechanism by which Wasserstein robustness induces regularity on the Bellman fixed point. This has direct implications for discretization, value-function approximation, and learning algorithms on general (possibly unbounded) state spaces, where Lipschitz continuity is a standard prerequisite for error bounds and convergence rates.
major comments (1)
- [Abstract and §1] Abstract and §1 (Introduction): the claim that Lipschitz assumptions on the model primitives suffice for Lipschitz robust value functions on possibly unbounded Polish spaces is load-bearing for the central result. The Wasserstein ball of positive radius around a nominal kernel on an unbounded space can contain conditional distributions whose first moments differ arbitrarily from the nominal; without explicit uniform integrability or growth controls on the admissible kernels, the worst-case expectation operator need not map Lipschitz functions to Lipschitz functions. The manuscript must either impose such conditions or demonstrate that they are unnecessary under the stated assumptions.
Simulated Author's Rebuttal
We thank the referee for the careful reading and for highlighting a potential subtlety regarding moment control on unbounded Polish spaces. We address the concern directly below.
read point-by-point responses
-
Referee: [Abstract and §1] Abstract and §1 (Introduction): the claim that Lipschitz assumptions on the model primitives suffice for Lipschitz robust value functions on possibly unbounded Polish spaces is load-bearing for the central result. The Wasserstein ball of positive radius around a nominal kernel on an unbounded space can contain conditional distributions whose first moments differ arbitrarily from the nominal; without explicit uniform integrability or growth controls on the admissible kernels, the worst-case expectation operator need not map Lipschitz functions to Lipschitz functions. The manuscript must either impose such conditions or demonstrate that they are unnecessary under the stated assumptions.
Authors: The Wasserstein-1 ball of radius ε around a nominal kernel P(·|s) consists exclusively of measures μ satisfying W_1(μ, P(·|s)) ≤ ε < ∞. By the very definition of the Wasserstein metric on a Polish space, any such μ necessarily possesses a finite first moment whenever the nominal does, and the moments cannot differ arbitrarily: for the standard Kantorovich–Rubinstein representation, |∫ f dμ − ∫ f dP| ≤ ε for every 1-Lipschitz f, which directly bounds the difference of first moments (taking f(x) = d(x, x_0)). Consequently, the admissible kernels inherit uniform integrability of the first moment from the nominal kernel. Moreover, the Lipschitz property of the worst-case expectation follows immediately from the same representation: for any L-Lipschitz function v, |E_μ v − E_ν v| ≤ L W_1(μ, ν) whenever the expectations exist. The paper works throughout under the standing hypothesis that all kernels appearing in the Wasserstein balls have finite first moments (required for the balls to be non-empty and the robust Bellman operator to be well-defined). No additional uniform-integrability assumption is therefore needed; the metric itself supplies the requisite control. We are happy to add a brief clarifying remark in §2 if the editor deems it useful, but the existing framework already renders the claim valid on unbounded spaces. revision: no
Circularity Check
No circularity; result derived from robustness model and contraction arguments
full rationale
The paper establishes Lipschitz regularity of the robust value function as a consequence of the Wasserstein ball formulation plus Lipschitz primitives on Polish spaces. The abstract explicitly contrasts this with the non-robust case where the same assumptions fail to yield Lipschitz value functions, indicating the regularization effect is a genuine consequence of the robustness construction rather than a definitional or fitted tautology. No self-citations, ansatzes, or renamings are invoked as load-bearing steps in the provided text. The derivation is therefore self-contained and independent of its inputs.
Axiom & Free-Parameter Ledger
axioms (2)
- standard math State space is a Polish space (possibly unbounded)
- domain assumption Infinite-horizon discounted reward criterion
read the original abstract
Robust Markov decision processes provide a principled framework for protecting sequential decision-making against transition-law misspecification and have attracted substantial recent research interest. As in non-robust stochastic optimal control, an important question is whether the robust value function is sufficiently regular for approximation and learning. This paper studies Lipschitz regularity of optimal value functions for Wasserstein robust stochastic optimal control on possibly unbounded Polish state spaces under an infinite-horizon discounted reward criterion. We consider two robustness formulations: a kernel-robust model, in which the adversary perturbs the next-state distribution within a Wasserstein ball around a nominal transition kernel, and a noise-robust model, in which the adversary perturbs the driving noise of the state transition dynamics. In the non-robust setting, Lipschitz rewards and Lipschitz transition dynamics do not, in general, imply a Lipschitz value function. In contrast, we show that in these Wasserstein robust formulations, Lipschitz assumptions on the model primitives yield Lipschitz robust value functions. Thus, Wasserstein robustness not only protects against misspecification but also regularizes the Bellman fixed point, providing stability relevant to discretization, value-function approximation, estimation, and learning.
Forward citations
Cited by 1 Pith paper
-
Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
The sample-based value function in discounted infinite-horizon stochastic control converges to a Gaussian process limit that solves a linear DP-type fixed-point equation.
Reference graph
Works this paper leans on
-
[1]
Asadi, K., Misra, D., and Littman, M. (2018). Lipschitz continuity in model-based reinforcement learning. In Dy, J. and Krause, A., editors,Proceedings of the 35th International Conference on Machine Learning, volume 80 ofProceedings of Machine Learning Research, pages 264–273. PMLR. 1, 3
2018
-
[2]
Bartl, D., Drapeau, S., Ob l´ oj, J., and Wiesel, J. (2021). Sensitivity analysis of wasserstein distributionally robust optimization problems.Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences, 477(2256):20210176. 3
2021
-
[3]
Bertsekas, D. P. and Shreve, S. E. (1978).Stochastic Optimal Control: The Discrete-Time Case, volume 139 ofMathematics in Science and Engineering. Academic Press, New York. 1, 2
1978
-
[4]
and Murthy, K
Blanchet, J. and Murthy, K. R. A. (2019). Quantifying distributional model risk via optimal transport. Mathematics of Operations Research, 44(2):565–600. 2, 3, 4
2019
-
[5]
and Prieto-Rumeau, T
Dufour, F. and Prieto-Rumeau, T. (2012). Approximation of Markov decision processes with general state space.Journal of Mathematical Analysis and Applications, 388(2):1254–1267. 1, 3
2012
-
[6]
and Kleywegt, A
Gao, R. and Kleywegt, A. J. (2023). Distributionally robust stochastic optimization with wasserstein distance.Mathematics of Operations Research, 48(2):603–655. 3
2023
-
[7]
and Kroer, C
Grand-Cl´ ement, J. and Kroer, C. (2021). First-order methods for wasserstein distributionally robust MDP. In Meila, M. and Zhang, T., editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 2010–2019. PMLR. 3
2021
-
[8]
Hinderer, K. (2005). Lipschitz continuity of value functions in Markovian decision processes.Mathematical Methods of Operations Research, 62(1):3–22. 1, 3
2005
-
[9]
Iyengar, G. N. (2005). Robust dynamic programming.Mathematics of Operations Research, 30(2):257–
2005
-
[10]
and Kuhn, D
Mohajerin Esfahani, P. and Kuhn, D. (2018). Data-driven distributionally robust optimization using the wasserstein metric: Performance guarantees and tractable reformulations.Mathematical Programming, 171(1–2):115–166. 2, 3 16
2018
-
[11]
M¨ uller, A. (1997). How does the value function of a Markov decision process depend on the transition probabilities?Mathematics of Operations Research, 22(4):872–885. 3
1997
-
[12]
and El Ghaoui, L
Nilim, A. and El Ghaoui, L. (2005). Robust control of Markov decision processes with uncertain transition matrices.Operations Research, 53(5):780–798. 3
2005
-
[13]
Srivastava, S. M. (1998).A course on Borel sets. Springer. 13
1998
-
[14]
(2009).Optimal Transport: Old and New, volume 338 ofGrundlehren der Mathematischen Wissenschaften
Villani, C. (2009).Optimal Transport: Old and New, volume 338 ofGrundlehren der Mathematischen Wissenschaften. Springer, Berlin, Heidelberg. 3
2009
-
[15]
Wang, S. (2026). Q-measure-learning for continuous state RL: Efficient implementation and convergence. 3
2026
-
[16]
Wang, S., Blanchet, J., and Glynn, P. (2026). Fast convergence of policy regret in learning stochastic optimal control. 1, 3
2026
-
[17]
Wang, S., Si, N., Blanchet, J., and Zhou, Z. (2025a). On the foundation of distributionally robust reinforcement learning. 1, 2
-
[18]
Wang, S., Si, N., Blanchet, J., and Zhou, Z. (2025b). Statistical learning of distributionally robust stochastic control in continuous state spaces. In Li, Y., Mandt, S., Agrawal, S., and Khan, E., editors, Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, volume 258 of Proceedings of Machine Learning Research, pa...
-
[19]
Wiesemann, W., Kuhn, D., and Rustem, B. (2013). Robust Markov decision processes.Mathematics of Operations Research, 38(1):153–183. 1, 3
2013
-
[20]
Yang, I. (2017). A convex optimization approach to distributionally robust Markov decision processes with wasserstein distance.IEEE Control Systems Letters, 1(1):164–169. 3
2017
-
[21]
Yang, I. (2021). Wasserstein distributionally robust stochastic control: A data-driven approach.IEEE Transactions on Automatic Control, 66(8):3863–3870. 3
2021
-
[22]
Zhou, Y., Song, Y., and Y¨ uksel, S. (2026). Robustness to model approximation, model learning from data, and sample complexity in wasserstein regular MDPs. 3 17
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.