REVIEW 2 major objections 4 minor 30 references
Zero-sum stochastic linear-quadratic games with Markov regime switching satisfy an exponential turnpike property.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Finite-horizon optimal feedback gains in zero-sum stochastic linear-quadratic games with regime switching converge exponentially to infinite-horizon gains, yielding a turnpike theorem for the optimal triple.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Genuine extension with an exponential Riccati engine, but Theorem 5.2's proof has a real gap in the state-difference estimate; repairable, and worth refereeing. the 2 major comments →
Turnpike properties for zero-sum stochastic linear quadratic differential games of Markovian regime switching system
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Under assumptions (A1) and (A2), for every initial pair (x,i) there exist K, mu>0, independent of T, such that E[|X_T(t)-X_infty(t)|^2+|u_{1,T}(t)-u_{1,infty}(t)|^2+|u_{2,T}(t)-u_{2,infty}(t)|^2] <= K|x|^2(e^{-mu(T-t)}+e^{-mu t}), where the subscript T denotes the finite-horizon optimal triple and the subscript infinity the infinite-horizon optimal triple. The proof first shows that the unique strongly regular solution P_T of the coupled differential Riccati equations converges to the unique solution P_infty of the coupled algebraic Riccati equations at rate K e^{-mu(T-t)} (Theorem 5.1), and then transfers this convergence, via the closed-loop feedback representations, into the state and con
What carries the argument
Coupled differential Riccati equations (CDREs) for the finite-horizon game and their algebraic counterparts (CAREs) for the infinite-horizon game. The CDRE solution P_T(t,i) encodes the value function V_T(x,i)=<P_T(t,i)x,x> and yields the open-loop saddle strategy in closed-loop form, u_T(t)=Theta_T(t,alpha_t)X_T(t), with Theta_T(t,i)=-N(t;P_T,i)^{-1}L(t;P_T,i)^top. The core argument proves that P_T converges exponentially to the unique CARE solution P_infty and that the feedback gains converge at the same rate; the stable dynamics under the infinite-horizon feedback then convert this Riccati convergence into the turnpike estimate via a Lyapunov/Gronwall step. The uniform convexity-concavity
Load-bearing premise
The whole argument leans on (A1), which requires that, uniformly over every horizon and every Markov regime, Player 1's control cost is bounded below by delta times its squared norm and Player 2's payoff is bounded above by -delta times its squared norm; if this uniform convexity-concavity fails, the Riccati equations may have no strongly regular solution and the exponential turnpike proof collapses.
What would settle it
Take a one-dimensional, two-regime system satisfying (A2), set R11(1)=0 so that (A1) fails, and check whether a unique finite-horizon saddle still exists while the claimed exponential estimate fails. Under (A1)-(A2), a more direct falsifier is numerical: fix a two-regime example, solve the CDREs (20) backward, simulate the two closed-loop systems, and look for any t, T, x for which the bound (73) is exceeded; a single such counterexample would disprove Theorem 5.2.
If this is right
- For any fixed kappa in (0,1/2), once T is large, the finite-horizon optimal state and controls on [kappa T,(1-kappa)T] are within O(e^{-mu kappa T}) of the infinite-horizon optimal triple, uniformly in the initial pair up to the factor |x|^2.
- The infinite-horizon stationary feedback can be used as an approximate solution of the finite-horizon game, and inequality (4) gives an explicit, computable error estimate for that approximation.
- The coupled algebraic Riccati equations (27) admit a unique solution satisfying the stabilizability condition (28), so the infinite-horizon problem is uniquely solvable and its value function is V_infty(x,i)=<P_infty(i)x,x>.
- The value functions differ by an exponentially small horizon-truncation penalty: |V_T(x,i)-V_infty(x,i)| <= K|x|^2 e^{-mu T}, directly from the Riccati convergence rate at t=0.
- The same exponential rate applies to the feedback gains, so the closed-loop approximation inherits the turnpike property.
Where Pith is reading between the lines
- Editorial inference: the proof's structure suggests the same template of exponential Riccati convergence plus stability of the infinite-horizon closed loop would yield turnpike estimates for regime-switching games with periodic or ergodic coefficients, with the stationary CARE replaced by the corresponding periodic or ergodic Riccati solution; the paper does not claim this extension.
- Because (A1) enters only through uniform invertibility of N, a testable extension is to let the convexity-concavity constant delta shrink with T and track how the turnpike rate depends on delta; the current theorem requires delta fixed.
- The explicit constants K and mu in the turnpike bound depend on the Lyapunov exponent of the stable closed-loop system under Theta_infty; estimating them from data would let an engineer decide in advance whether a given horizon T is long enough for the infinite-horizon strategy to be used—an inference not drawn in the paper.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the long-time behavior of zero-sum stochastic linear-quadratic differential games with Markov regime switching. Under a uniform convexity-concavity assumption (A1) and an L2-stability assumption (A2), it claims: (i) unique open-loop solvability of the finite- and infinite-horizon problems; (ii) convergence of the associated coupled differential Riccati equations (CDREs) to coupled algebraic Riccati equations (CAREs); (iii) exponential convergence of the Riccati solutions; and (iv) an exponential turnpike estimate for the optimal triple, i.e. the finite-horizon optimal state and controls are close to the infinite-horizon stationary feedback away from the temporal endpoints. The main tools are Hilbert-space operator representations of the cost, FBSDE stationarity conditions, Lyapunov/Gronwall arguments, and Riccati comparison.
Significance. If the main theorem is correct, this is a meaningful extension of the turnpike literature: it moves from deterministic or single-player stochastic LQ problems, and from zero-sum games without regime switching, to a Markovian regime-switching zero-sum SLQ setting, with explicit exponential rates. The paper also consolidates and extends the authors' previous solvability results for finite and infinite horizons. The operator-based uniqueness arguments and the Riccati convergence analysis are valuable in themselves. However, the proof of the central turnpike estimate (Theorem 5.2) contains a serious gap that must be repaired before the main claim can be regarded as established.
major comments (2)
- [§5, proof of Theorem 5.2 (Eq. (74))] The proof defines X_T = \bar X_T - \bar X_\infty, which satisfies the forced SDE (74) with X_T(0)=0. The text states that 'by (72) and Corollary 2.1, one has E|X_T(t)|^2 ≤ K|x|^2 e^{-\mu_3 t}'. Corollary 2.1 applies to a homogeneous L2-stable system with arbitrary initial condition, not to a forced process with zero initial condition. The claimed decay in t is essentially the turnpike estimate being proved, while the forcing term (\Theta_T-\Theta_\infty)\bar X_T is only known to decay in T-t. Without the e^{-\mu_3 t} bound, the line K|\Theta_\infty-\Theta_T| E[|\Sigma^{1/2}X_T||\bar X_T|] ≤ K|x|^2 e^{-2\mu_2(T-t)}e^{-\mu_3 t} is unsupported and the Gronwall step does not close. This is a load-bearing gap in the proof of (73).
- [§5, proof of Theorem 5.1 (Eqs. (66)-(69))] Equation (67) is stated as a linear bound |[\Theta-\Theta_T]^\top N_T[\Theta-\Theta_T]| ≤ K_1|\Sigma_T(t)|, but by (66) \Theta-\Theta_T is linear in \Sigma_T, so the left-hand side is quadratic in \Sigma_T. The subsequent inequality (69) uses |\Sigma_T(s)|^2, which is consistent only with a quadratic bound, not with the displayed (67). If (67) is taken literally, the comparison argument leading to h(t)≤\mu does not follow. Additionally, in Step 2 the replacement of |\Sigma_T(k)| by \rho is not explicitly justified; it can likely be obtained from Step 1 by choosing N large, but this needs to be written. Since (61) is used to derive (72) and then Theorem 5.2, these details matter.
minor comments (4)
- [§2, Proposition 2.1] Proposition 2.1 is stated under assumption (A1), but the estimate is a stability/regularity result for the linear SDE (9) and requires L2-stability of [A,C] (i.e., (A2)), not the uniform convexity-concavity condition. Corollary 2.1 inherits this mislabeling. Since (A2) is assumed in all main theorems, this is repairable, but the statement should be corrected.
- [§3, Theorem 3.3 proof] The system displayed after 'equation (43) is in turn equivalent to' repeats the first equation twice: the second line should read M^α_{21,∞}\bar u_{1,∞}+M^α_{22,∞}\bar u_{2,∞}+K^α_{2,∞}x=0.
- [§2, Eq. (39)] The formula for M^α_{ij,T} writes the adjoint term as (L^α_{j,T})^*S_i(α)^\top; for off-diagonal entries this should be (L^α_{i,T})^*S_j(α)^\top. As written, the off-diagonal blocks are not consistent with a self-adjoint M^α_T.
- [§4, proof of Theorem 4.1] The line '||\hat u_T(·)||≤KE|Y_∞(T)|^2' mixes the Hilbert-space norm with the expectation. It should read E||\hat u_T(·)||^2 ≤ K E|Y_∞(T)|^2.
Circularity Check
No definitional or self-citation circularity; the central turnpike proof is independent of the claimed conclusion, though it contains a non-circular proof gap in Theorem 5.2.
full rationale
Walking the derivation: (A1)-(A2) are assumptions, not the turnpike inequality (73). Theorems 3.1/3.2, Corollary 2.1, and Proposition 2.1 are imported from the authors' prior works [27,28,29], but those works establish solvability, FBSDE characterizations and stability estimates, not the exponential turnpike bound. Theorem 5.1 derives Riccati convergence from the CDREs by a Gronwall argument; Theorem 5.2 then uses (72) and a Lyapunov differential inequality for E<Sigma(alpha)X_T,X_T>. The target (73) does not appear among the inputs of any lemma; the proof is therefore not circular in the self-definitional or fitted-parameter sense. The one serious issue is the sentence in the proof of Theorem 5.2: 'by (72) and Corollary 2.1, one has ... E|X_T(t)|^2 <= K|x|^2 e^{-mu3 t}'. Corollary 2.1 applies to the homogeneous system [A,C] with arbitrary initial condition, whereas X_T is defined by (74) with X_T(0)=0 and nontrivial forcing terms. So the exponential-in-t bound on X_T is not supplied by the cited result; it is a proof gap. This is a correctness concern, not a circular reduction: the asserted bound is not the theorem's conclusion, is not a hidden restatement of an assumption, and is not shown to be equivalent to the quantity being estimated. Accordingly, the circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption A1: uniform convexity-concavity (10) for all T in (0,infinity] and all i in S
- domain assumption A2: L2-stability of the uncontrolled system [A,C]_alpha
- standard math Finite-state irreducible Markov chain alpha with generator Pi, independent of a one-dimensional Brownian motion
- domain assumption Stability and Lyapunov estimates from Wu-Li-Zhang [29, Prop. 2.2 and 2.5]
Cite this review
Pith. "Pith review of Turnpike properties for zero-sum stochastic linear quadratic differential games of Markovian regime switching system." pith.science (2026). https://pith.science/paper/B6UJ2OJI
@misc{pith2026250909358,
author = {Pith},
title = {Pith review of: Turnpike properties for zero-sum stochastic linear quadratic differential games of Markovian regime switching system},
year = {2026},
howpublished = {\url{https://pith.science/paper/B6UJ2OJI}},
note = {Machine review of arXiv:2509.09358}
}
read the original abstract
This paper investigates the long-time behavior of zero-sum stochastic linear-quadratic (SLQ) differential games within Markov regime-switching diffusion systems and establishes the turnpike property of the optimal triple. By verifying the convergence of the associated coupled differential Riccati equations (CDREs) along with their convergence rate, we show that, for a sufficiently large time horizon, the equilibrium strategy in the finite-horizon problem can be closely approximated by that of the infinite-horizon problem. Furthermore, this study enhances and extends existing results concerning zero-sum SLQ differential games over both finite and infinite horizons.
Reference graph
Works this paper leans on
-
[1]
A weak dynamic programming principle for zero-sum stochastic differential games with unbounded controls.SIAM Journal on Control and Optimization, 51(3): 2036–2080, 2013
Erhan Bayraktar and Song Yao. A weak dynamic programming principle for zero-sum stochastic differential games with unbounded controls.SIAM Journal on Control and Optimization, 51(3): 2036–2080, 2013
2036
-
[2]
On the turnpike property and the receding-horizon method for linear-quadratic optimal control problems.SIAM Journal on Control and Optimization, 58(2): 1077–1102, 2020
Tobias Breiten and Laurent Pfeiffer. On the turnpike property and the receding-horizon method for linear-quadratic optimal control problems.SIAM Journal on Control and Optimization, 58(2): 1077–1102, 2020
2020
-
[3]
Stochastic differential games and viscosity solutions of Hamilton- Jacobi-Bellman-Isaacs equations.SIAM Journal on Control and Optimization, 47(1):444–475, 2008
Rainer Buckdahn and Juan Li. Stochastic differential games and viscosity solutions of Hamilton- Jacobi-Bellman-Isaacs equations.SIAM Journal on Control and Optimization, 47(1):444–475, 2008
2008
-
[4]
An exponential turnpike theorem for dissipative discrete time optimal control problems.SIAM Journal on Control and Optimization, 52(3):1935–1957, 2014
Tobias Damm, Lars Grune, Marleen Stieler, and Karl Worthmann. An exponential turnpike theorem for dissipative discrete time optimal control problems.SIAM Journal on Control and Optimization, 52(3):1935–1957, 2014
1935
-
[5]
McGraw-Hill, New York, 1958
Robert Dorfman, Paul Anthony Samuelson, and Robert M Solow.Linear Programming and Eco- nomic Analysis. McGraw-Hill, New York, 1958
1958
-
[6]
On the existence of value functions of two-player, zero-sum stochastic differential games.Indiana University Mathematics Journal, 38(2):293–314, 1989
Wendell H Fleming and Panagiotis E Souganidis. On the existence of value functions of two-player, zero-sum stochastic differential games.Indiana University Mathematics Journal, 38(2):293–314, 1989
1989
-
[7]
Turnpike properties and strict dissipativity for discrete time linear quadratic optimal control problems.SIAM Journal on Control and Optimization, 56(2):1282– 1302, 2018
Lars Gr¨ une and Roberto Guglielmi. Turnpike properties and strict dissipativity for discrete time linear quadratic optimal control problems.SIAM Journal on Control and Optimization, 56(2):1282– 1302, 2018
2018
-
[8]
Zero-sum stochastic differential games and backward equations.Systems & Control Letters, 24(4):259–263, 1995
Saıd Hamad´ ene and Jean-Pierre Lepeltier. Zero-sum stochastic differential games and backward equations.Systems & Control Letters, 24(4):259–263, 1995
1995
-
[9]
Turnpike properties of optimal relaxed control problems.ESAIM: Control, Optimisation and Calculus of Variations, 25:74, 2019
Hongwei Lou and Weihan Wang. Turnpike properties of optimal relaxed control problems.ESAIM: Control, Optimisation and Calculus of Variations, 25:74, 2019
2019
-
[10]
Two-player zero-sum stochastic differential games with regime switching.Automatica, 114: 108819, 2020
Siyu Lv. Two-player zero-sum stochastic differential games with regime switching.Automatica, 114: 108819, 2020. 22
2020
-
[11]
Hongwei Mei, Rui Wang, and Jiongmin Yong. Turnpike property of stochastic linear-quadratic optimal control problems in large horizons with regime switching i: Homogeneous cases.arXiv preprint arXiv:2506.09337, 2025
Pith/arXiv arXiv 2025
-
[12]
Two-person zero-sum linear quadratic stochastic differential games by a Hilbert space method.Journal of Industrial and Management Optimization, 2(1):93–115, 2006
Libin Mou, Jiongmin Yong, et al. Two-person zero-sum linear quadratic stochastic differential games by a Hilbert space method.Journal of Industrial and Management Optimization, 2(1):93–115, 2006
2006
-
[13]
A model of general economic equilibrium.The Review of Economic Studies, 13(1): 1–9, 1945
J v Neumann. A model of general economic equilibrium.The Review of Economic Studies, 13(1): 1–9, 1945
1945
-
[14]
Long time versus steady state optimal control.SIAM Journal on Control and Optimization, 51(6):4242–4273, 2013
Alessio Porretta and Enrique Zuazua. Long time versus steady state optimal control.SIAM Journal on Control and Optimization, 51(6):4242–4273, 2013
2013
-
[15]
A mathematical theory of saving.The economic journal, 38(152):543–559, 1928
Frank Plumpton Ramsey. A mathematical theory of saving.The economic journal, 38(152):543–559, 1928
1928
-
[16]
The turnpike property in nonlinear optimal control—a geometric approach.Automatica, 134:109939, 2021
Noboru Sakamoto and Enrique Zuazua. The turnpike property in nonlinear optimal control—a geometric approach.Automatica, 134:109939, 2021
2021
-
[17]
Turnpike and dissipativity in generalized discrete-time stochastic linear-quadratic optimal control.SIAM Journal on Control and Optimization, 63(2):1432–1457, 2025
Jonas Schießl, Ruchuan Ou, Timm Faulwasser, Michael H Baumann, and Lars Gr¨ une. Turnpike and dissipativity in generalized discrete-time stochastic linear-quadratic optimal control.SIAM Journal on Control and Optimization, 63(2):1432–1457, 2025
2025
-
[18]
Two-person zero-sum stochastic linear-quadratic differential games.SIAM Journal on Control and Optimization, 59(3):1804–1829, 2021
Jingrui Sun. Two-person zero-sum stochastic linear-quadratic differential games.SIAM Journal on Control and Optimization, 59(3):1804–1829, 2021
2021
-
[19]
Linear quadratic stochastic differential games: Open-loop and closed-loop saddle points.SIAM Journal on Control and Optimization, 52(6):4082–4121, 2014
Jingrui Sun and Jiongmin Yong. Linear quadratic stochastic differential games: Open-loop and closed-loop saddle points.SIAM Journal on Control and Optimization, 52(6):4082–4121, 2014
2014
-
[20]
Jingrui Sun and Jiongmin Yong. Long-time behavior of zero-sum linear-quadratic stochastic differ- ential games.arXiv preprint arXiv:2406.02089, 2024
Pith/arXiv arXiv 2024
-
[21]
Turnpike properties for mean-field linear-quadratic optimal control problems.SIAM Journal on Control and Optimization, 62(1):752–775, 2024
Jingrui Sun and Jiongmin Yong. Turnpike properties for mean-field linear-quadratic optimal control problems.SIAM Journal on Control and Optimization, 62(1):752–775, 2024
2024
-
[22]
Turnpike properties for stochastic linear-quadratic optimal control problems with periodic coefficients.Journal of Differential Equations, 400:189–229, 2024
Jingrui Sun and Jiongmin Yong. Turnpike properties for stochastic linear-quadratic optimal control problems with periodic coefficients.Journal of Differential Equations, 400:189–229, 2024
2024
-
[23]
Turnpike properties for stochastic linear-quadratic optimal control problems.Chinese Annals of Mathematics, Series B, 43(6):999–1022, 2022
Jingrui Sun, Hanxiao Wang, and Jiongmin Yong. Turnpike properties for stochastic linear-quadratic optimal control problems.Chinese Annals of Mathematics, Series B, 43(6):999–1022, 2022
2022
-
[24]
The exponential turnpike property for periodic linear quadratic optimal control problems in infinite dimension.SIAM Journal on Control and Optimization, 63(4):2524–2546, 2025
Emmanuel Tr´ elat, Xingwu Zeng, and Can Zhang. The exponential turnpike property for periodic linear quadratic optimal control problems in infinite dimension.SIAM Journal on Control and Optimization, 63(4):2524–2546, 2025
2025
-
[25]
A pontryagin’s maximum principle for non-zero sum differential games of BSDEs with applications.IEEE Transactions on Automatic control, 55(7):1742–1747, 2010
Guangchen Wang and Zhiyong Yu. A pontryagin’s maximum principle for non-zero sum differential games of BSDEs with applications.IEEE Transactions on Automatic control, 55(7):1742–1747, 2010
2010
-
[26]
A partial information non-zero sum differential game of backward stochastic differential equations with applications.Automatica, 48(2):342–352, 2012
Guangchen Wang and Zhiyong Yu. A partial information non-zero sum differential game of backward stochastic differential equations with applications.Automatica, 48(2):342–352, 2012
2012
-
[27]
Fan Wu, Xun Li, Jie Xiong, and Xin Zhang. Stochastic linear-quadratic differential game with Markovian jumps in an infinite horizon.arXiv preprint arXiv:2408.12818, 2024. 23
Pith/arXiv arXiv 2024
-
[28]
Fan Wu, Xun Li, and Xin Zhang. Open-loop and closed-loop solvabilities for zero-sum stochas- tic linear quadratic differential games of Markovian regime switching system.arXiv preprint arXiv:2409.01973, 2024
Pith/arXiv arXiv 2024
-
[29]
Stochastic linear quadratic optimal control problems with regime- switching jumps in infinite horizon.SIAM Journal on Control and Optimization, 63(2):852–891, 2025
Fan Wu, Xun Li, and Xin Zhang. Stochastic linear quadratic optimal control problems with regime- switching jumps in infinite horizon.SIAM Journal on Control and Optimization, 63(2):852–891, 2025
2025
-
[30]
Zhiyong Yu. An optimal feedback control-strategy pair for zero-sum linear-quadratic stochastic differential game: The Riccati equation approach.SIAM Journal on Control and Optimization, 53 (4):2141–2167, 2015. 24
2015
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.