REVIEW 3 major objections 4 minor 32 references
Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper shows that in zero-sum stochastic linear-quadratic dynamic games with partial and asymmetric information, best-response dynamics converge in value within a few iterations, so low-dimensional belief-state controllers can closely…
desk verdict Useful recursive best-response construction for asymmetric LQG games, with a suggestive numerical phenomenon, but the low-order approximation claim outruns what Theorem 1 proves. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the augmented belief-state system built at each best-response step: the opponent's previous controller is folded into the dynamics, output map, and cost weighting, so the acting player must estimate a state of dimension $(2k-1)n$ or $2kn$. The carrier of the argument is the pair of algebraic Riccati equations for control and estimation—equations (22) and (24) for the minimizer, (30) and (32) for the maximizer—which define each best response, together with the Lyapunov equations for the controllability and observability Gramians and their Hankel singular values. Theorem 1, using Cholesky factors of a Cauchy matrix built from the eigenvalues of the augmented dynamics matrix and a bilinear transformation, converts rapid eigenvalue decay into a quantitative low-rank approximation bound, which is what supports the claim that low-order belief dynamics are nearly as good as the full hierarchy.
What would settle it
Run the best-response recursion on an open-loop-stable system whose dynamics matrix has eigenvalues very close to the unit circle, so the Cholesky decay factors are large; if the game value still changes materially after many iterations, or if the Hankel singular values fail to fall below the tolerance used in the low-order approximation, the central claim fails. A sharper counterexample would be any instance where at some iteration the pair $(A^k_i, B^k_i)$ is not stabilizable or $(A^k_i, C^k_i)$ is not detectable, making the Riccati solution and hence the approximate value undefined.
Extended reading notes
Core claim
The central claim is that in an infinite-horizon zero-sum linear-quadratic dynamic game where each player observes a private noisy linear measurement, best-response iteration within the class of pure linear output-feedback strategies produces controllers of increasing internal dimension—one dimension per order of belief—yet the value stabilizes after a few iterations. The paper derives explicit recursions: the minimizer at iteration $k$ faces an augmented state of dimension $(2k-1)n$, and the maximizer one of dimension $2kn$, with the gains and costs determined by Riccati equations (22), (24), (30), and (32). It then argues that the higher-order belief dynamics become increasingly hard to control and observe, using controllability and observability Gramians and Hankel singular values, and Theorem 1 bounds the low-rank approximation error of the controllability Gramian by Cholesky estimates of eigenvalue decay. The paper concludes that low-order belief dynamics approximate the infinite-dimensional equilibrium strategies with bounded error.
Load-bearing premise
The load-bearing premise is that at every best-response iteration the augmented system is stabilizable and detectable, so that the Riccati equations (22), (24), (30), and (32) admit stabilizing solutions even though the minimizer's state cost is indefinite and the maximizer's input cost is negative definite; the paper states this only as 'certain conditions' without specifying them.
Editorial extensions
If this is right
- In systems where the decay condition in Theorem 1 holds, a controller of fixed internal dimension built from the first few belief orders achieves cost within an explicit error bound of the infinite-dimensional equilibrium.
- Best-response iteration can be stopped once the game value stops changing, giving finite-order strategies with near-equilibrium performance and avoiding the infinite belief hierarchy.
- The Gramian and Hankel singular value decay rates provide a computable criterion for when higher-order beliefs are unnecessary: when the Cholesky ratio $\delta^k_l/\delta^k_1$ falls below a tolerance, the remaining belief directions contribute little.
- The result extends the practical reach of LQG control to asymmetric-information games by showing that the infinite regress of beliefs can be truncated without much performance loss.
- If the paper is right, the Nash equilibrium of the original game is well approximated by a finite-dimensional linear output-feedback strategy pair.
Reading between the lines
- Beyond the paper, an extension would be to run the same recursion on N-player nonzero-sum LQG games, where value convergence can fail even if the decay mechanism is similar; the numerical evidence suggests but does not prove such a result.
- The decay rates observed in the paper suggest a model-selection heuristic: choose the internal controller dimension at the iteration where the Hankel singular value curve flattens, connecting this game-solving method to balanced-truncation model reduction.
- A testable boundary condition is to construct systems with eigenvalues near the unit circle, where the Cholesky decay factors degrade, and check whether value convergence slows; this would expose how general the rapid-decay phenomenon is.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies infinite-horizon zero-sum linear-quadratic dynamic games (LQDGs) with partial and asymmetric information. It proposes best-response dynamics in which each player's strategy is a linear output-feedback controller whose internal state (belief state) is an augmented state of dimension an integer multiple of the system state. The paper derives the best-response Riccati equations and estimator gains at each iteration, then analyzes controllability/observability Gramians and Hankel singular values of the augmented belief dynamics, proving a Gramian approximation bound (Theorem 1) via Cholesky factors. Numerical experiments show that the value appears to converge within a few iterations and that Gramian eigenvalues and Hankel singular values decay rapidly, leading the authors to conclude that low-order belief feedback can closely approximate Nash equilibrium strategies.
Significance. If the central claim were fully established, the paper would provide a tractable route to approximate equilibrium computation in asymmetric-information stochastic games, giving a concrete model-reduction perspective on higher-order beliefs. The explicit algebraic derivation of both players' best responses is a useful contribution and the numerical setup is clearly described. The paper is honest that value convergence is observed rather than proven, and the comparison of Cholesky estimates with actual Gramian eigenvalue decay rates is a legitimate consistency check. However, the inference from Gramian approximation error to value suboptimality is currently unjustified, and the existence conditions for the indefinite Riccati equations are not stated. With those gaps filled, the paper would be a valuable contribution; as it stands, the main claim goes beyond what the theorems and experiments support.
major comments (3)
- [Section III-C.1 and III-C.2, Eqs. (22), (24), (30), (32)] The text states only that unique solutions of the Riccati equations exist 'under certain conditions' and that the closed-loop system remains stable at each iteration, with references [29], [30]. Because Q^1_k is indefinite and R2 is negative definite, the standard stabilizability/detectability hypotheses used for the LQG Riccati equation are not sufficient; one needs explicit conditions for the existence of stabilizing solutions to the associated indefinite algebraic Riccati equations, together with the definiteness condition R2 + (B^2_k)^T P^2_k B^2_k ≺ 0 for the maximizer. These conditions must be stated and verified at every iteration; otherwise the costs J^{1*}_k and J^{2*}_k, the gains K^i_k and L^i_k, and the Gramians in Eqs. (33)–(34) are not well defined.
- [Section IV-B, Theorem 1, and following paragraph] Theorem 1 bounds the infinity-norm distance between the true controllability Gramian W^i_{c,k} and a low-rank Cholesky-based approximation; it does not bound the difference in the players' cost functionals J^{1*}_k and J^{2*}_k, the distance of the truncated strategy pair to a best-response fixed point, or the value gap between the low-order and the full-order infinite-belief equilibrium. Therefore the sentence 'the strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies' does not follow from the theorem. Section V reports only Gramian eigenvalue and Hankel singular value decay (Table I), never the actual costs of truncated low-order strategies played against full-order best responses. To support the central claim, the paper needs either a Lipschitz/fixed-point argument that connects a Gramian approximation error to suboptimality in the objective (5), or a direct numerical comparison of the costs of low-order truncated strategies against full-order best responses.
- [Section V-A and Section VI] The conclusion that 'this point of convergence corresponds to a Nash equilibrium' is based on visual inspection of Figure 1. No convergence proof for the best-response iteration is provided, and no limiting value is characterized; even if the sequences J^{1*}_k and J^{2*}_k appear to converge numerically, the restricted strategy class (linear output feedback with internal dimension a multiple of n) means that convergence of these costs does not by itself imply that the limiting pair is a Nash equilibrium of (5). The claim should be softened to a conjecture, or supported by a proof of convergence or a finite-iteration error bound.
minor comments (4)
- [Eqs. (36) and (37)] The definitions of the Cholesky factors use inconsistent indexing: Eq. (36) writes lambda^i_{l,k}, while Eq. (37) writes lambda^i_k in the first factor and inside the product; the notation should be made uniform.
- [Theorem 1] The bound contains the undefined term 'mini k' and the approximation delta^i_{1,k} ≈ ||W^i_{c,k}||_2 is stated without precise constants; a precise definition of 'mini k' and a numerical verification of the bound itself would strengthen the result.
- [End of Section IV-B] The phrase 'analogous statements are easily obtained for the observability Gramian, and related bounds for the Hankel singular values' leaves the observability-side tool unproved; since the paper draws conclusions about both controllability and observability, a formal statement and proof (or at least a precise reference to a discrete-time analogue) should be included.
- [Throughout] There are several typos: 'maitrx' in Theorem 1, 'continous-time' in its proof, and 'the the' in the third paragraph of Section I.
Circularity Check
No significant circularity: the Gramian decay theorem is an external model-reduction bound applied to the belief dynamics, and the value-convergence claim rests on numerical observation rather than on a fitted or self-referential input.
full rationale
The derivation chain is not circular. Section III constructs each player's best response from standard LQG Riccati equations whose inputs are the opponent's previously fixed strategy and the common-knowledge system/cost parameters; the resulting costs are computed, not imposed. Section IV-B proves a Gramian approximation bound by applying Theorem 3.2 of Antoulas, Sorensen, and Zhou [28] under stated stabilizability/detectability and diagonalizability assumptions, with the Cholesky factors in Eq. (36) computed from the spectrum of A_i_k and then compared against actual Gramian eigenvalue decay in Fig. 2, rather than fitted to those decay values. The only self-citation is Ref. [27], which is used as a motivational pointer to network controllability studies and is not load-bearing for the game-theoretic or model-reduction claims. The paper's concluding leap from a bounded Gramian approximation error to 'strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies' is an unsupported inference that may be a correctness risk, but it is not a circular reduction: no equation defines the Nash-approximation claim in terms of the Gramian bound, and no fitted parameter is renamed as a prediction. Since the central numerical observations are self-contained and the formal result is cited from independent prior work, the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption All system and cost parameters are common knowledge across players.
- standard math For a fixed opponent strategy, the optimal strategy is a linear feedback of the Kalman filter state estimate (LQG separation).
- ad hoc to paper The algebraic Riccati equations (22), (24), (30), (32) have unique stabilizing solutions despite indefinite Q1_k and negative definite R2.
- domain assumption At each iteration the augmented matrices A_i_k are stable and diagonalizable, and the pairs (A_i_k, B_i_k) and (A_i_k, C_i_k) are controllable and observable.
- standard math The bilinear transformation preserves Gramians and the Cholesky estimates from Antoulas et al. apply to the discrete-time setting.
Cite this review
Pith. "Pith review of Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information." pith.science (2026). https://pith.science/paper/6E3EHLJJ
@misc{pith2026250106181,
author = {Pith},
title = {Pith review of: Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information},
year = {2026},
howpublished = {\url{https://pith.science/paper/6E3EHLJJ}},
note = {Machine review of arXiv:2501.06181}
}
read the original abstract
We analyze best response dynamics for finding a Nash equilibrium of an infinite horizon zero-sum stochastic linear quadratic dynamic game (LQDG) with partial and asymmetric information. We derive explicit expressions for each player's best response within the class of pure linear dynamic output feedback control strategies where the internal state dimension of each control strategy is an integer multiple of the system state dimension. With each best response, the players form increasingly higher-order belief states, leading to infinite-dimensional internal states. However, we observe in extensive numerical experiments that the game's value converges after just a few iterations, suggesting that strategies associated with increasingly higher-order belief states eventually provide no benefit. To help explain this convergence, our numerical analysis reveals rapid decay of the controllability and observability Gramian eigenvalues and Hankel singular values in higher-order belief dynamics, indicating that the higher-order belief dynamics become increasingly difficult for both players to control and observe. Consequently, the higher-order belief dynamics can be closely approximated by low-order belief dynamics with bounded error, and thus feedback strategies with limited internal state dimension can closely approximate a Nash equilibrium.
Figures
Reference graph
Works this paper leans on
-
[29]
Least squares stationary optimal control and the algebraic Riccati equation,
J. Willems, “Least squares stationary optimal control and the algebraic Riccati equation,” IEEE Transactions on Automatic Control , vol. 16, no. 6, pp. 621–634, Dec. 1971
work page 1971
-
[30]
The stable regulator problem and its inverse,
B. Molinari, “The stable regulator problem and its inverse,” IEEE Trans. on Automatic Control , vol. 18, no. 5, pp. 454–459, Oct. 1973
work page 1973
-
[1]
Bas ¸ar and G
T. Bas ¸ar and G. J. Olsder, Dynamic Noncooperative Game Theory, 2nd Edition. Society for Industrial and Applied Mathematics, 1998
1998
-
[2]
A multistage pursuit-evasion game that admits a gaussian random process as a maximin control policy,
B. Tamer and M. Max, “A multistage pursuit-evasion game that admits a gaussian random process as a maximin control policy,” Stochastics: An International Journal of Probability and Stochastic Processes, vol. 1, no. 1-4, pp. 25–69, 1973
work page 1973
-
[3]
Differential games with imperfect state information,
I. Rhodes and D. Luenberger, “Differential games with imperfect state information,” IEEE Transactions on Automatic Control , vol. 14, no. 1, pp. 29–38, 1969
work page 1969
-
[4]
Decomposition techniques for markov zero-sum games with nested information,
J. Zheng and D. A. Casta ˜n´on, “Decomposition techniques for markov zero-sum games with nested information,” in 52nd IEEE conference on decision and control . IEEE, 2013, pp. 574–581
work page 2013
-
[5]
A. Gupta, A. Nayyar, C. Langbort, and T. Basar, “Common informa- tion based markov perfect equilibria for linear-gaussian games with asymmetric information,” SIAM Journal on Control and Optimization , vol. 52, no. 5, pp. 3228–3260, 2014
work page 2014
-
[6]
Zero-sum stochastic games with asymmetric information,
D. Kartik and A. Nayyar, “Zero-sum stochastic games with asymmetric information,” in 2019 IEEE 58th Conference on Decision and Control (CDC). IEEE, 2019, pp. 4061–4066
work page 2019
Show all 32 references
-
[7]
Linear-quadratic gaussian games with asymmetric information: Belief corrections using the opponents actions,
B. Hambly, R. Xu, and H. Yang, “Linear-quadratic gaussian games with asymmetric information: Belief corrections using the opponents actions,” arXiv preprint arXiv:2307.15842 , 2023
2023 arXiv
-
[8]
Stochastic dynamic games in belief space,
W. Schwarting, A. Pierson, S. Karaman, and D. Rus, “Stochastic dynamic games in belief space,” IEEE Transactions on Robotics, vol. 37, no. 6, pp. 2157–2172, 2021
2021
-
[9]
Learning mixed strategies in trajectory games,
L. Peters, D. Fridovich-Keil, L. Ferranti, C. Stachniss, J. Alonso-Mora, and F. Laine, “Learning mixed strategies in trajectory games,” arXiv preprint arXiv:2205.00291, 2022
2022 arXiv
-
[10]
Learning equilibria in asymmetric auction games,
M. Bichler, N. Kohring, and S. Heidekr ¨uger, “Learning equilibria in asymmetric auction games,” INFORMS Journal on Computing , vol. 35, no. 3, pp. 523–542, 2023
2023
-
[11]
Dynamic games with asymmetric information and resource constrained players with applications to security of cyberphysical systems,
A. Gupta, C. Langbort, and T. Bas ¸ar, “Dynamic games with asymmetric information and resource constrained players with applications to security of cyberphysical systems,” IEEE Transactions on Control of Network Systems, vol. 4, no. 1, pp. 71–81, 2016
2016
-
[12]
Dynamic games for secure and resilient control system design,
Y . Huang, J. Chen, L. Huang, and Q. Zhu, “Dynamic games for secure and resilient control system design,” National Science Review , vol. 7, no. 7, pp. 1125–1141, 2020
2020
-
[13]
Games with incomplete information played by “bayesian
J. C. Harsanyi, “Games with incomplete information played by “bayesian” players, i–iii part i. the basic model,” Management science, vol. 14, no. 3, pp. 159–182, 1967
1967
-
[14]
The market for “lemons
G. A. Akerlof, “The market for “lemons”: Quality uncertainty and the market mechanism,” The quarterly journal of economics , vol. 84, no. 3, pp. 488–500, 1970
1970
-
[15]
Agreeing to disagree,
R. J. Aumann, “Agreeing to disagree,” The Annals of Statistics , vol. 4, no. 6, pp. 1236–1239, 1976
1976
-
[16]
Job market signaling,
M. Spence, “Job market signaling,” in Uncertainty in economics . Elsevier, 1978, pp. 281–306
1978
-
[17]
Credit rationing in markets with imperfect information,
J. E. Stiglitz and A. Weiss, “Credit rationing in markets with imperfect information,” The American economic review , vol. 71, no. 3, pp. 393– 410, 1981
1981
-
[18]
The theory of learning in games, economics learning and social evolution series,
D. Fudenberg and D. Levine, “The theory of learning in games, economics learning and social evolution series,” 1998
1998
-
[19]
Computing best-response strate- gies in infinite games of incomplete information,
D. Reeves and M. P. Wellman, “Computing best-response strate- gies in infinite games of incomplete information,” arXiv preprint arXiv:1207.4171, 2012
2012 arXiv
-
[20]
On the rate of convergence of continuous-time fictitious play,
C. Harris, “On the rate of convergence of continuous-time fictitious play,” Games and Economic Behavior , vol. 22, no. 2, pp. 238–259, 1998
1998
-
[21]
Best response dynamics for continuous zero-sum games,
J. Hofbauer and S. Sorin, “Best response dynamics for continuous zero-sum games,” Discrete and Continuous Dynamical Systems Series B, vol. 6, no. 1, p. 215, 2006
2006
-
[22]
W. H. Sandholm, Population games and evolutionary dynamics . MIT press, 2010
2010
-
[23]
Best response model predictive control for agile interactions between autonomous ground vehicles,
G. Williams, B. Goldfain, P. Drews, J. M. Rehg, and E. A. Theodorou, “Best response model predictive control for agile interactions between autonomous ground vehicles,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 2403–2410
2018
-
[24]
Game theoretic motion planning for multi-robot racing,
Z. Wang, R. Spica, and M. Schwager, “Game theoretic motion planning for multi-robot racing,” in Distributed Autonomous Robotic Systems: The 14th International Symposium . Springer, 2019, pp. 225–238
2019
-
[25]
Game- theoretic planning for self-driving cars in multivehicle competitive scenarios,
M. Wang, Z. Wang, J. Talbot, J. C. Gerdes, and M. Schwager, “Game- theoretic planning for self-driving cars in multivehicle competitive scenarios,” IEEE Transactions on Robotics , vol. 37, no. 4, pp. 1313– 1325, 2021
2021
-
[26]
Controllability metrics, limitations and algorithms for complex networks,
F. Pasqualetti, S. Zampieri, and F. Bullo, “Controllability metrics, limitations and algorithms for complex networks,” IEEE Transactions on Control of Network Systems , vol. 1, no. 1, pp. 40–52, 2014
2014
-
[27]
Performance bounds for optimal and robust feedback control in networks,
K. Ganapathy, J. Ruths, and T. Summers, “Performance bounds for optimal and robust feedback control in networks,” IEEE Transactions on Control of Network Systems , vol. 8, no. 4, pp. 1754–1766, 2021
2021
-
[28]
On the decay rate of hankel singular values and related issues,
A. C. Antoulas, D. C. Sorensen, and Y . Zhou, “On the decay rate of hankel singular values and related issues,” Systems & Control Letters , vol. 46, no. 5, pp. 323–342, 2002
2002
-
[31]
S. Boyd, L. El Ghaoui, E. Feron, and V . Balakrishnan, Linear matrix inequalities in system and control theory . SIAM, 1994
1994
-
[32]
All optimal Hankel-norm approximations of linear multivariable systems and their L, ∞ -error bounds†,
Keith. Glover, “All optimal Hankel-norm approximations of linear multivariable systems and their L, ∞ -error bounds†,” International Journal of Control , vol. 39, no. 6, pp. 1115–1193, June 1984
1984
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.