REVIEW 3 major objections 3 minor 34 references
Linear-quadratic Stochastic Stackelberg Differential Games with Affine Constraints
T0 review · 3 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Constrained leader-follower stochastic games get explicit feedback equilibria
desk verdict The problem and framework are a genuine extension of the constrained LQ-SSDG literature, but the paper's own examples fail to verify the key solvability assumption H3, so the submission is not ready as is. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The proof chain uses three linked devices. First, the follower's optimal strategy is put in feedback form $u_1^* = -E_1^{-1}B_1^\top(\phi_1 X^* + \psi_1)$ through the stochastic Riccati equation (8) and the BSDE (9). Second, the leader's relaxed problem is rewritten as the fully coupled FBSDE (17), and a four-step scheme reduces solvability to the nonstandard stochastic Riccati equation (19) for $(\phi_2,\tilde\phi_2)$, whose unique solution is assumed in (H3). Third, Lagrangian duality replaces the constrained leader problem by the dual problem (4.2), and the KKT system (32) ties the dual multiplier to the feedback gains, yielding the equilibrium pair (33).
What would settle it
In Example 5.1, evaluate the proposed $(\phi_2,\tilde\phi_2)$ at $t=1$ and compare with the terminal condition $\phi_2(1)=\operatorname{diag}(G_2,0_{n\times n})$; if the entries do not match, the example does not actually verify (H3).
Extended reading notes
Core claim
The paper's central claim is Theorem 4.5: under (H1)-(H4), the pair $(\hat u_1,\hat u_2)$ defined by (33) is a feedback Stackelberg equilibrium of Problem 2.1. The optimal strategies are affine in the augmented state $\hat Z$ and the dual multiplier $\lambda^*$, and $\lambda^*$ is characterized by the KKT conditions (32), with $\rho_i(\lambda^*)-a_i$ giving the slack of the $i$-th affine constraint. When the matrix $S_0$ in (23) is strictly positive definite, $\lambda^*$ is the unique maximizer of the dual problem (Problem 4.2).
Load-bearing premise
Assumption (H3), posed in Section 4 before Theorem 4.2, states that the nonstandard stochastic Riccati equation (19) for $(\phi_2,\tilde\phi_2)$ has a unique solution; all feedback formulas for the leader, the KKT system, and the final equilibrium depend on this solution, and the paper assumes it rather than proving a sufficient condition.
Editorial extensions
If this is right
- Constrained LQ Stackelberg games are reduced to solving a stochastic Riccati equation and a KKT system, so standard numerical solvers for Riccati equations and quadratic programs can be applied directly.
- Under (H1)-(H4) with $S_0\in\mathbb{S}^l_{++}$, the dual problem has a unique solution, so the KKT system (32) has a unique $\lambda^*$ and the feedback equilibrium is uniquely identified.
- The affine constraints can be checked through the explicit formula for $\rho(u_2^*(\lambda))$ in Proposition 4.1, giving a computable membership test for the leader's admissible set.
- The method generalizes the affine-constraint approach previously developed for single-player stochastic LQ control, pointing to a uniform treatment of constraints in hierarchical stochastic control.
Reading between the lines
- A natural follow-up is to test the Slater-condition verification on mean-field or jump-diffusion Stackelberg models, where the same FBSDE trick may bypass the usual verification bottleneck.
- The positive-definiteness of $S_0$ could be checked numerically before solving the game; if it fails, the dual maximizer may be nonunique, and the equilibrium selection would need extra criteria.
- If correct, the affine-constraint machinery could be combined with the approximation in Example 1.1 to handle quadratic and risk constraints in Stackelberg games.
- The new FBSDE rewriting of the constraints appears to be a transferable tool, likely useful in other hierarchical control problems where the Slater condition is the main obstruction.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper considers a non-zero-sum linear-quadratic stochastic Stackelberg differential game with affine constraints on the leader's admissible strategies, where the constraints depend on the state, the follower's best response, and the leader's control. It derives the follower's feedback strategy from a stochastic Riccati equation (SRE), reformulates the leader's problem through a fully coupled FBSDE and a nonstandard SRE, and constructs a Lagrangian dual. Under assumptions (H1)-(H4) it proves strong duality, gives a KKT condition, and states a feedback Stackelberg equilibrium in Theorem 4.5. Two examples with indefinite coefficients are presented as illustrations.
Significance. The paper's architecture—constraint reformulation via FBSDEs, dual problem, KKT conditions, and a positivity condition for uniqueness—is a reasonable extension of known techniques and, if the hypotheses can be guaranteed, would provide a useful recipe for constrained leader-follower LQ games. It is a strength that the main result is stated as a conditional theorem with explicit hypotheses rather than as a formal derivation from weaker assumptions. However, the paper's only nontrivial instances verifying the key hypothesis (H3) are the Section 5 examples, and these examples contain checkable algebraic errors; moreover, H3 itself is assumed with no sufficient condition. Thus the current manuscript does not demonstrate that its framework applies to any non-degenerate model. The contribution would be significant after the examples are corrected and H3 is supported.
major comments (3)
- [Section 5, Eq. (8)] The claimed solution (φ1,eφ1)=(1,0) does not solve SRE (8). For the data of (34)-(35) one has A=C=1, B1=1/2, E1=4, D1=-2, G1=1, hence S1=1/16 and the drift bracket in (8) equals D1+φ1A+A⊤φ1+C⊤eφ1+eφ1C+C⊤φ1C−φ1S1φ1 = -2+1+1+0+0+1−1/16 = 15/16 ≠0. Therefore dφ1 = −15/16 dt, not zero, and the feedback formula (10) in the examples is not established. This is not a typographical slip in the final formula only; the same φ1 is used in the construction of ψ1 and in the leader's SRE data.
- [Section 5, Eq. (19)] The proposed φ2 = diag(-1,1−e^{-t}) violates the terminal condition of (19): φ2(1) must equal diag(G2,0)=diag(-1,0), but its (2,2)-entry at t=1 is 1−e^{-1}≠0. Since (19) is the nonstandard SRE whose solvability is exactly assumption (H3), and since (21), (32), and (33) all depend on that solution, neither Example 5.1 nor Example 5.2 verifies the hypotheses of Theorem 4.5. The printed Stackelberg equilibrium formulas in these examples therefore do not follow from the stated assumptions.
- [Section 4, Assumption (H3)] The paper assumes existence and uniqueness of a solution to the stochastic Riccati equation (19) without providing any sufficient condition. The matrix F in (17) is indefinite (with −S1 on the off-diagonal) and the terminal data are singular, so this assumption is not covered by the standard Riccati theory invoked for (8). Because the leader's feedback strategy (21), the KKT system (32), and the equilibrium (33) all require φ2, the class of models for which the main theorem is applicable remains unspecified. A proof of H3 under checkable conditions, or at least corrected examples with a verifiable solution, is needed.
minor comments (3)
- [Section 5, Example 5.1] The text says that eJ(λ*,u*_2(λ*)) is 'strictly concave' in a and reaches its minimum at a=0. With λ*=2a/(e^{-2}-1), the expression eJ=-1-2aλ*-(1-e^{-2})λ*^2/2 equals -1+2a^2/(1-e^{-2}), which is strictly convex in a, not concave; the stated 'minimum at a=0' is consistent with convexity, so the wording should be corrected.
- [Section 5, Example 5.2] With λ* = max(2a/(e^{-2}-1),0), the value function equals -1 for a≥0 and is strictly convex in a for a<0; the minimum is attained for all a≥0, not 'when a≤0' as stated. The sentence about strengthening constraints should be revised accordingly.
- [Problem 4.1] The admissible space in the infimum is written as L^2_F([s,T],R^n), but u2 is R^m-valued; it should be L^2_F([s,T],R^m).
Circularity Check
No significant circularity: the derivation is a hypothesis-driven KKT/duality argument whose inputs are not defined in terms of the target equilibrium.
full rationale
The paper's central claim, Theorem 4.5, is conditional on assumptions (H1)-(H4). In particular, (H3) is an explicit existence hypothesis for the nonstandard stochastic Riccati equation (19); the paper states 'Thus we need the following assumption for ensuring the solvability of (19)' and does not derive the existence of (φ2, eφ2) from the conclusion it is later used to prove. The feedback formulas (21) and (33) are obtained by the four-step scheme and Itô calculus from the state equation and cost functionals, not by fitting parameters to the claimed equilibrium. The dual problem (Problem 4.2) is built from the Lagrangian of the leader's constrained problem, and strong duality (Theorem 4.3) is proved through a convex separation argument under the Slater condition (H4); it is not an equivalence inserted by definition. The positive-definiteness condition S0 ∈ S^l_{++} is presented as a sufficient condition for uniqueness in Remark 4.1 and Theorem 4.4, not as a forced ansatz. Self-citations to the authors' earlier work, e.g. [7] in the introduction, are motivational ('Gou et al. [7] showed that many expectation-type constraints ... can be approximately captured by finite many affine constraints') and are not load-bearing for the main theorem; the existence and uniqueness results actually invoked for the Riccati equation and BSDEs are [22, Theorem 6.3] and [29, Proposition 3.2, Theorem 1.25], which are external references. The paper also flags its own limitations honestly: Remark 4.3 notes possible non-uniqueness of the feedback Stackelberg equilibrium, and Section 6 states that 'the positive definiteness of S0 given by (23) also calls for further research.' The reviewer-flagged issue in Section 5, that the proposed φ2 = diag(-1, 1 - e^{-t}) may violate the terminal condition of (19) (and similarly the companion φ1 claim), is a correctness or verification defect in the illustrative examples, not a circularity: even if the examples fail to exhibit an instance satisfying (H3), the main theorem remains a conditional statement depending on (H1)-(H4). No equation in the paper reduces by construction to a fitted value, a renamed known result, or a self-citation chain. Hence the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption H1: uniform coercivity of the follower's cost along zero-initial-state perturbations (finds ϵ1>0).
- domain assumption H2: uniform coercivity of the leader's relaxed cost along zero-initial-state perturbations (finds ϵ2>0).
- ad hoc to paper H3: unique solvability of the nonstandard stochastic Riccati equation (19) for (φ2, eφ2).
- domain assumption H4: Slater condition (25) plus linear independence of the equality-constraint directions.
- standard math Standard solvability of linear FBSDEs and Riccati equations from [22], [29], and the usual hypothesis on the filtered probability space.
Cite this review
Pith. "Pith review of Linear-quadratic Stochastic Stackelberg Differential Games with Affine Constraints." pith.science (2026). https://pith.science/paper/XLH6HCRJ
@misc{pith2026241218802,
author = {Pith},
title = {Pith review of: Linear-quadratic Stochastic Stackelberg Differential Games with Affine Constraints},
year = {2026},
howpublished = {\url{https://pith.science/paper/XLH6HCRJ}},
note = {Machine review of arXiv:2412.18802}
}
read the original abstract
This paper investigates the non-zero-sum linear-quadratic stochastic Stackelberg differential games with affine constraints, which depend on both the follower's response and the leader's strategy. With the help of the stochastic Riccati equations and the Lagrangian duality theory, the feedback expressions of optimal strategies of the follower and the leader are obtained and the dual problem of the leader's problem is established. Under the Slater condition, the equivalence is proved between the solutions to the dual problem and the leader's problem, and the KKT condition is also provided for solving the dual problem. Then, the feedback Stackelberg equilibrium is provided for the linear-quadratic stochastic Stackelberg differential games with affine constraints, and a new positive definite condition is proposed for ensuring the uniqueness of solutions to the dual problem. Finally, two non-degenerate examples with indefinite coefficients are provided to illustrate and to support our main results.
Figures
Reference graph
Works this paper leans on
-
[1]
Linear quadratic stochastic control problems with stochastic terminal constraint
Peter Bank and Moritz Voß. Linear quadratic stochastic control problems with stochastic terminal constraint. SIAM Journal on Control and Optimization , 56(2):672–699, 2018
work page 2018
-
[2]
Fr´ ed´ eric Bonnans and Alexander Shapiro.Perturbation Analysis of Optimization Problems
J. Fr´ ed´ eric Bonnans and Alexander Shapiro.Perturbation Analysis of Optimization Problems . Springer, New York, 2013
work page 2013
-
[3]
On dynamic programming principle for stochastic control under expectation constraints
Yuk-Loong Chow, Xiang Yu, and Chao Zhou. On dynamic programming principle for stochastic control under expectation constraints. Journal of Optimization Theory and Applications , 185(3):803–818, 2020
work page 2020
-
[4]
A feedback Stackelberg game of cooperative advertising in a durable goods oligopoly
Anshuman Chutani and Suresh P Sethi. A feedback Stackelberg game of cooperative advertising in a durable goods oligopoly. In Dynamic Games in Economics , pages 89–114. Springer, Berlin, 2014
work page 2014
-
[5]
Contract Theory in Continuous-time Models
Jakˇ sa Cvitanic and Jianfeng Zhang. Contract Theory in Continuous-time Models . Springer, Berlin, 2013
work page 2013
-
[6]
Xinwei Feng, Ying Hu, and Jianhui Huang. Backward Stackelberg differential game with constraints: a mixed terminal-perturbation and linear-quadratic approach. SIAM Journal on Control and Optimiza- tion, 60(3):1488–1518, 2022
work page 2022
-
[7]
Stochastic linear-quadratic control problems with affine constraints
Zhun Gou, Nanjing Huang, Xianjun Long, and Jianhao Kang. Stochastic linear-quadratic control problems with affine constraints. Systems & Control Letters , 191:105887, 2024
work page 2024
-
[8]
A linear-quadratic mean-field stochastic Stackelberg differential game with random exit time
Zhun Gou, Nanjing Huang, and Minghui Wang. A linear-quadratic mean-field stochastic Stackelberg differential game with random exit time. International Journal of Control , 96(3):731–745, 2023
work page 2023
Show all 34 references
-
[9]
Mathematical programs with geometric constraints in Banach spaces: enhanced optimality, exact penalty, and sensitivity
Lei Guo, Jane J Ye, and Jin Zhang. Mathematical programs with geometric constraints in Banach spaces: enhanced optimality, exact penalty, and sensitivity. SIAM Journal on Optimization , 23(4):2295–2319, 2013
2013
-
[10]
J. Ye Jane. Necessary and sufficient optimality conditions for mathematical programs with equilibrium constraints. Journal of Mathematical Analysis and Applications , 307(1):350–369, 2005. 18
2005
-
[11]
The computation of approximate feedback stackelberg equilibria in multiplayer nonlinear constrained dynamic games
Jingqi Li, Somayeh Sojoudi, Claire J Tomlin, and David Fridovich-Keil. The computation of approximate feedback stackelberg equilibria in multiplayer nonlinear constrained dynamic games. SIAM Journal on Optimization, 34(4):3723–3749, 2024
2024
-
[12]
Andrew E. B. Lim and Xun Yu Zhou. Stochastic optimal LQR control with integral quadratic constraints and indefinite control weights. IEEE Transactions on Automatic Control , 44(7):1359–1369, 1999
1999
-
[13]
Forward-backward Stochastic Differential Equations and their Applications
Jin Ma and Jiongmin Yong. Forward-backward Stochastic Differential Equations and their Applications. Springer, Berlin, 1999
1999
-
[14]
A two-stage stochastic Stackelberg model for microgrid operation with chance constraints for renewable energy generation uncertainty
Yolanda Matamala and Felipe Feijoo. A two-stage stochastic Stackelberg model for microgrid operation with chance constraints for renewable energy generation uncertainty. Applied Energy, 303:117608, 2021
2021
-
[15]
A stochastic goodwill model depending on quality level and advertising
Jun Meng, Minghui Wang, Benzhang Yang, and Nanjing Huang. A stochastic goodwill model depending on quality level and advertising. Optimization, 72(10):2463–2497, 2023
2023
-
[16]
Linear-quadratic stochastic Stackelberg differential games for jump-diffusion systems
Jun Moon. Linear-quadratic stochastic Stackelberg differential games for jump-diffusion systems. SIAM Journal on Control and Optimization , 59(2):954–976, 2021
2021
-
[17]
Linear-quadratic mean-field type Stackelberg differential games for stochastic jump-diffusion systems
Jun Moon. Linear-quadratic mean-field type Stackelberg differential games for stochastic jump-diffusion systems. Mathematical Control & Related Fields , 12(2), 2022
2022
-
[18]
Linear-quadratic stochastic leader-follower differential games for markov jump-diffusion models
Jun Moon. Linear-quadratic stochastic leader-follower differential games for markov jump-diffusion models. Automatica, 147:110713, 2023
2023
-
[19]
Stochastic Stackelberg equilibria with applications to time-dependent newsvendor models
Bernt Øksendal, Leif Sandal, and Jan Ubøe. Stochastic Stackelberg equilibria with applications to time-dependent newsvendor models. Journal of Economic Dynamics and Control , 37(7):1284–1299, 2013
2013
-
[20]
Integral reinforcement learning-based dynamic event-triggered safety control for multiplayer Stackelberg-Nash games with time-varying state constraints
Chunbin Qin, Tianzeng Zhu, Kaijun Jiang, and Yinliang Wu. Integral reinforcement learning-based dynamic event-triggered safety control for multiplayer Stackelberg-Nash games with time-varying state constraints. Engineering Applications of Artificial Intelligence , 133:108317, 2024
2024
-
[21]
Radio resource sharing and edge caching with latency constraint for local 5G operator: Geometric programming meets stackelberg game
Tachporn Sanguanpuak, Dusit Niyato, Nandana Rajatheva, and Matti Latva-Aho. Radio resource sharing and edge caching with latency constraint for local 5G operator: Geometric programming meets stackelberg game. IEEE Transactions on Mobile Computing , 20(2):707–721, 2019
2019
-
[22]
Indefinite stochastic linear-quadratic optimal control problems with random coefficients: Closed-loop representation of open-loop optimal controls
Jingrui Sun, Jie Xiong, and Jiongmin Yong. Indefinite stochastic linear-quadratic optimal control problems with random coefficients: Closed-loop representation of open-loop optimal controls. The Annals of Applied Probability , 31(1):460–499, 2021
2021
-
[23]
Market Structure and Equilibrium
Heinrich Von Stackelberg. Market Structure and Equilibrium . Springer, Berlin, 2010
2010
-
[24]
Zero-sum stochastic linear-quadratic Stackelberg differential games with jumps
Fan Wu, Jie Xiong, and Xin Zhang. Zero-sum stochastic linear-quadratic Stackelberg differential games with jumps. Applied Mathematics & Optimization , 89(1):29, 2024
2024
-
[25]
On continuous-time constrained stochastic linear- quadratic control
Weiping Wu, Jianjun Gao, Junguo Lu, and Xun Li. On continuous-time constrained stochastic linear- quadratic control. Automatica, 114:108809, 2020
2020
-
[26]
Stochastic linear-quadratic Stackelberg differential game with asymmetric informational uncertainties: Robust optimization approach
Na Xiang and Jingtao Shi. Stochastic linear-quadratic Stackelberg differential game with asymmetric informational uncertainties: Robust optimization approach. arXiv preprint arXiv:2407.05728 , 2024. 19
2024 arXiv
-
[27]
Stackelberg game theory based model to guide users’ energy use behavior, with the consideration of flexible resources and consumer psychology, for an integrated energy system
Haoran Yan, Hongjuan Hou, Min Deng, Lengge Si, Xi Wang, Eric Hu, and Rhonin Zhou. Stackelberg game theory based model to guide users’ energy use behavior, with the consideration of flexible resources and consumer psychology, for an integrated energy system. Energy, 288:129806, 2024
2024
-
[28]
A leader-follower stochastic linear quadratic differential game.SIAM Journal on Control and Optimization , 41(4):1015–1041, 2002
Jiongmin Yong. A leader-follower stochastic linear quadratic differential game.SIAM Journal on Control and Optimization , 41(4):1015–1041, 2002
2002
-
[29]
Stochastic optimal control–A concise introduction
Jiongmin Yong. Stochastic optimal control–A concise introduction. Mathematical Control & Related Fields, 12(4):1039–1136, 2022
2022
-
[30]
Stochastic linear quadratic optimal control problems with expectation-type linear equality constraints on the terminal states
Haisen Zhang and Xianfeng Zhang. Stochastic linear quadratic optimal control problems with expectation-type linear equality constraints on the terminal states. Systems & Control Letters , 177:105551, 2023
2023
-
[31]
Global solutions of stochastic Stackelberg differential games under convex control constraint
Liangquan Zhang and Wei Zhang. Global solutions of stochastic Stackelberg differential games under convex control constraint. Systems & Control Letters , 156:105020, 2021
2021
-
[32]
A linear quadratic Stackelberg game of backward stochastic differential equations with partial information
Yueyang Zheng and Jingtao Shi. A linear quadratic Stackelberg game of backward stochastic differential equations with partial information. In 2020 39th Chinese Control Conference (CCC) , pages 966–971. IEEE, 2020
2020
-
[33]
A linear-quadratic partially observed Stackelberg stochastic differential game with application
Yueyang Zheng and Jingtao Shi. A linear-quadratic partially observed Stackelberg stochastic differential game with application. Applied Mathematics and Computation , 420:126819, 2022
2022
-
[34]
Stackelberg game-based decentralised sup- ply chain coordination considering both futures and spot markets
Suli Zou, Zhongjing Ma, Peng Wang, and Xiangdong Liu. Stackelberg game-based decentralised sup- ply chain coordination considering both futures and spot markets. International Journal of Control , 93(12):2804–2813, 2020. 20
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.