Pith. sign in

REVIEW 3 major objections 4 minor 32 references

Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper shows that in zero-sum stochastic linear-quadratic dynamic games with partial and asymmetric information, best-response dynamics converge in value within a few iterations, so low-dimensional belief-state controllers can closely…

desk verdict Useful recursive best-response construction for asymmetric LQG games, with a suggestive numerical phenomenon, but the low-order approximation claim outruns what Theorem 1 proves. read the letter →

arxiv 2501.06181 v2 pith:6E3EHLJJ submitted 2025-01-10 eess.SY cs.SYmath.OC

classification eess.SYcs.SYmath.OC MSC 91A2593E2093B11
keywords zero-sumgamesasymmetricinformationbestresponsedynamicshigher-orderbeliefsGramianeigenvaluedecayHankelsingularvalueslinearquadraticGaussianmodelreduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether players in an infinite-horizon zero-sum stochastic linear-quadratic dynamic game with partial and asymmetric information can settle on strategies that use only low-dimensional belief states. It derives the best-response recursion: with each round, a player's optimal controller tracks not only the state but also the opponent's belief, so the controller's internal dimension grows. The paper reports that the game value converges within a few rounds across numerical experiments, and attributes this to rapid decay of controllability and observability Gramian eigenvalues and Hankel singular values of the higher-order belief dynamics. If true, the practical consequence is that simple feedback strategies with limited internal state dimension can closely approximate a Nash equilibrium.

What carries the argument

The central object is the augmented belief-state system built at each best-response step: the opponent's previous controller is folded into the dynamics, output map, and cost weighting, so the acting player must estimate a state of dimension $(2k-1)n$ or $2kn$. The carrier of the argument is the pair of algebraic Riccati equations for control and estimation—equations (22) and (24) for the minimizer, (30) and (32) for the maximizer—which define each best response, together with the Lyapunov equations for the controllability and observability Gramians and their Hankel singular values. Theorem 1, using Cholesky factors of a Cauchy matrix built from the eigenvalues of the augmented dynamics matrix and a bilinear transformation, converts rapid eigenvalue decay into a quantitative low-rank approximation bound, which is what supports the claim that low-order belief dynamics are nearly as good as the full hierarchy.

What would settle it

Run the best-response recursion on an open-loop-stable system whose dynamics matrix has eigenvalues very close to the unit circle, so the Cholesky decay factors are large; if the game value still changes materially after many iterations, or if the Hankel singular values fail to fall below the tolerance used in the low-order approximation, the central claim fails. A sharper counterexample would be any instance where at some iteration the pair $(A^k_i, B^k_i)$ is not stabilizable or $(A^k_i, C^k_i)$ is not detectable, making the Riccati solution and hence the approximate value undefined.

Watch

Extended reading notes

Core claim

The central claim is that in an infinite-horizon zero-sum linear-quadratic dynamic game where each player observes a private noisy linear measurement, best-response iteration within the class of pure linear output-feedback strategies produces controllers of increasing internal dimension—one dimension per order of belief—yet the value stabilizes after a few iterations. The paper derives explicit recursions: the minimizer at iteration $k$ faces an augmented state of dimension $(2k-1)n$, and the maximizer one of dimension $2kn$, with the gains and costs determined by Riccati equations (22), (24), (30), and (32). It then argues that the higher-order belief dynamics become increasingly hard to control and observe, using controllability and observability Gramians and Hankel singular values, and Theorem 1 bounds the low-rank approximation error of the controllability Gramian by Cholesky estimates of eigenvalue decay. The paper concludes that low-order belief dynamics approximate the infinite-dimensional equilibrium strategies with bounded error.

Load-bearing premise

The load-bearing premise is that at every best-response iteration the augmented system is stabilizable and detectable, so that the Riccati equations (22), (24), (30), and (32) admit stabilizing solutions even though the minimizer's state cost is indefinite and the maximizer's input cost is negative definite; the paper states this only as 'certain conditions' without specifying them.

Editorial extensions

If this is right

  • In systems where the decay condition in Theorem 1 holds, a controller of fixed internal dimension built from the first few belief orders achieves cost within an explicit error bound of the infinite-dimensional equilibrium.
  • Best-response iteration can be stopped once the game value stops changing, giving finite-order strategies with near-equilibrium performance and avoiding the infinite belief hierarchy.
  • The Gramian and Hankel singular value decay rates provide a computable criterion for when higher-order beliefs are unnecessary: when the Cholesky ratio $\delta^k_l/\delta^k_1$ falls below a tolerance, the remaining belief directions contribute little.
  • The result extends the practical reach of LQG control to asymmetric-information games by showing that the infinite regress of beliefs can be truncated without much performance loss.
  • If the paper is right, the Nash equilibrium of the original game is well approximated by a finite-dimensional linear output-feedback strategy pair.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, an extension would be to run the same recursion on N-player nonzero-sum LQG games, where value convergence can fail even if the decay mechanism is similar; the numerical evidence suggests but does not prove such a result.
  • The decay rates observed in the paper suggest a model-selection heuristic: choose the internal controller dimension at the iteration where the Hankel singular value curve flattens, connecting this game-solving method to balanced-truncation model reduction.
  • A testable boundary condition is to construct systems with eigenvalues near the unit circle, where the Cholesky decay factors degrade, and check whether value convergence slows; this would expose how general the rapid-decay phenomenon is.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies infinite-horizon zero-sum linear-quadratic dynamic games (LQDGs) with partial and asymmetric information. It proposes best-response dynamics in which each player's strategy is a linear output-feedback controller whose internal state (belief state) is an augmented state of dimension an integer multiple of the system state. The paper derives the best-response Riccati equations and estimator gains at each iteration, then analyzes controllability/observability Gramians and Hankel singular values of the augmented belief dynamics, proving a Gramian approximation bound (Theorem 1) via Cholesky factors. Numerical experiments show that the value appears to converge within a few iterations and that Gramian eigenvalues and Hankel singular values decay rapidly, leading the authors to conclude that low-order belief feedback can closely approximate Nash equilibrium strategies.

Significance. If the central claim were fully established, the paper would provide a tractable route to approximate equilibrium computation in asymmetric-information stochastic games, giving a concrete model-reduction perspective on higher-order beliefs. The explicit algebraic derivation of both players' best responses is a useful contribution and the numerical setup is clearly described. The paper is honest that value convergence is observed rather than proven, and the comparison of Cholesky estimates with actual Gramian eigenvalue decay rates is a legitimate consistency check. However, the inference from Gramian approximation error to value suboptimality is currently unjustified, and the existence conditions for the indefinite Riccati equations are not stated. With those gaps filled, the paper would be a valuable contribution; as it stands, the main claim goes beyond what the theorems and experiments support.

major comments (3)
  1. [Section III-C.1 and III-C.2, Eqs. (22), (24), (30), (32)] The text states only that unique solutions of the Riccati equations exist 'under certain conditions' and that the closed-loop system remains stable at each iteration, with references [29], [30]. Because Q^1_k is indefinite and R2 is negative definite, the standard stabilizability/detectability hypotheses used for the LQG Riccati equation are not sufficient; one needs explicit conditions for the existence of stabilizing solutions to the associated indefinite algebraic Riccati equations, together with the definiteness condition R2 + (B^2_k)^T P^2_k B^2_k ≺ 0 for the maximizer. These conditions must be stated and verified at every iteration; otherwise the costs J^{1*}_k and J^{2*}_k, the gains K^i_k and L^i_k, and the Gramians in Eqs. (33)–(34) are not well defined.
  2. [Section IV-B, Theorem 1, and following paragraph] Theorem 1 bounds the infinity-norm distance between the true controllability Gramian W^i_{c,k} and a low-rank Cholesky-based approximation; it does not bound the difference in the players' cost functionals J^{1*}_k and J^{2*}_k, the distance of the truncated strategy pair to a best-response fixed point, or the value gap between the low-order and the full-order infinite-belief equilibrium. Therefore the sentence 'the strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies' does not follow from the theorem. Section V reports only Gramian eigenvalue and Hankel singular value decay (Table I), never the actual costs of truncated low-order strategies played against full-order best responses. To support the central claim, the paper needs either a Lipschitz/fixed-point argument that connects a Gramian approximation error to suboptimality in the objective (5), or a direct numerical comparison of the costs of low-order truncated strategies against full-order best responses.
  3. [Section V-A and Section VI] The conclusion that 'this point of convergence corresponds to a Nash equilibrium' is based on visual inspection of Figure 1. No convergence proof for the best-response iteration is provided, and no limiting value is characterized; even if the sequences J^{1*}_k and J^{2*}_k appear to converge numerically, the restricted strategy class (linear output feedback with internal dimension a multiple of n) means that convergence of these costs does not by itself imply that the limiting pair is a Nash equilibrium of (5). The claim should be softened to a conjecture, or supported by a proof of convergence or a finite-iteration error bound.
minor comments (4)
  1. [Eqs. (36) and (37)] The definitions of the Cholesky factors use inconsistent indexing: Eq. (36) writes lambda^i_{l,k}, while Eq. (37) writes lambda^i_k in the first factor and inside the product; the notation should be made uniform.
  2. [Theorem 1] The bound contains the undefined term 'mini k' and the approximation delta^i_{1,k} ≈ ||W^i_{c,k}||_2 is stated without precise constants; a precise definition of 'mini k' and a numerical verification of the bound itself would strengthen the result.
  3. [End of Section IV-B] The phrase 'analogous statements are easily obtained for the observability Gramian, and related bounds for the Hankel singular values' leaves the observability-side tool unproved; since the paper draws conclusions about both controllability and observability, a formal statement and proof (or at least a precise reference to a discrete-time analogue) should be included.
  4. [Throughout] There are several typos: 'maitrx' in Theorem 1, 'continous-time' in its proof, and 'the the' in the third paragraph of Section I.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Gramian decay theorem is an external model-reduction bound applied to the belief dynamics, and the value-convergence claim rests on numerical observation rather than on a fitted or self-referential input.

full rationale

The derivation chain is not circular. Section III constructs each player's best response from standard LQG Riccati equations whose inputs are the opponent's previously fixed strategy and the common-knowledge system/cost parameters; the resulting costs are computed, not imposed. Section IV-B proves a Gramian approximation bound by applying Theorem 3.2 of Antoulas, Sorensen, and Zhou [28] under stated stabilizability/detectability and diagonalizability assumptions, with the Cholesky factors in Eq. (36) computed from the spectrum of A_i_k and then compared against actual Gramian eigenvalue decay in Fig. 2, rather than fitted to those decay values. The only self-citation is Ref. [27], which is used as a motivational pointer to network controllability studies and is not load-bearing for the game-theoretic or model-reduction claims. The paper's concluding leap from a bounded Gramian approximation error to 'strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies' is an unsupported inference that may be a correctness risk, but it is not a circular reduction: no equation defines the Nash-approximation claim in terms of the Gramian bound, and no fitted parameter is renamed as a prediction. Since the central numerical observations are self-contained and the formal result is cited from independent prior work, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claims rest on standard LQG separation, common knowledge, and unproved well-posedness of indefinite-cost Riccati equations. No free parameters are fit to data, and no new physical or conceptual entities are introduced.

assumptions (5)
  • domain assumption All system and cost parameters are common knowledge across players.
    Assumption 1 in Section II; required for each player to reason about the opponent's best response and beliefs.
  • standard math For a fixed opponent strategy, the optimal strategy is a linear feedback of the Kalman filter state estimate (LQG separation).
    Invoked throughout Section III to derive best responses; standard but requires stabilizability and detectability.
  • ad hoc to paper The algebraic Riccati equations (22), (24), (30), (32) have unique stabilizing solutions despite indefinite Q1_k and negative definite R2.
    Stated in Section III-C.1 only 'under certain conditions' without precise hypotheses; this is load-bearing for all costs and Gramians.
  • domain assumption At each iteration the augmented matrices A_i_k are stable and diagonalizable, and the pairs (A_i_k, B_i_k) and (A_i_k, C_i_k) are controllable and observable.
    Required by Theorem 1 in Section IV-B to apply the Cholesky-based Gramian decay bounds.
  • standard math The bilinear transformation preserves Gramians and the Cholesky estimates from Antoulas et al. apply to the discrete-time setting.
    Used in the proof of Theorem 1, citing [28] and [32]; accepted background results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information." pith.science (2026). https://pith.science/paper/6E3EHLJJ

@misc{pith2026250106181,
  author       = {Pith},
  title        = {Pith review of: Best Response Convergence for Zero-sum Stochastic Dynamic Games with Partial and Asymmetric Information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6E3EHLJJ}},
  note         = {Machine review of arXiv:2501.06181}
}
read the original abstract

We analyze best response dynamics for finding a Nash equilibrium of an infinite horizon zero-sum stochastic linear quadratic dynamic game (LQDG) with partial and asymmetric information. We derive explicit expressions for each player's best response within the class of pure linear dynamic output feedback control strategies where the internal state dimension of each control strategy is an integer multiple of the system state dimension. With each best response, the players form increasingly higher-order belief states, leading to infinite-dimensional internal states. However, we observe in extensive numerical experiments that the game's value converges after just a few iterations, suggesting that strategies associated with increasingly higher-order belief states eventually provide no benefit. To help explain this convergence, our numerical analysis reveals rapid decay of the controllability and observability Gramian eigenvalues and Hankel singular values in higher-order belief dynamics, indicating that the higher-order belief dynamics become increasingly difficult for both players to control and observe. Consequently, the higher-order belief dynamics can be closely approximated by low-order belief dynamics with bounded error, and thus feedback strategies with limited internal state dimension can closely approximate a Nash equilibrium.

Figures

Figures reproduced from arXiv: 2501.06181 by the authors.

Figure 1
Figure 1. Optimal cost vs. best response iteration. The optimal [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. The eigenvalues of the minimizer’s controllability [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 28 canonical work pages

  1. [29]

    Least squares stationary optimal control and the algebraic Riccati equation,

    J. Willems, “Least squares stationary optimal control and the algebraic Riccati equation,” IEEE Transactions on Automatic Control , vol. 16, no. 6, pp. 621–634, Dec. 1971

  2. [30]

    The stable regulator problem and its inverse,

    B. Molinari, “The stable regulator problem and its inverse,” IEEE Trans. on Automatic Control , vol. 18, no. 5, pp. 454–459, Oct. 1973

  3. [1]

    Bas ¸ar and G

    T. Bas ¸ar and G. J. Olsder, Dynamic Noncooperative Game Theory, 2nd Edition. Society for Industrial and Applied Mathematics, 1998

  4. [2]

    A multistage pursuit-evasion game that admits a gaussian random process as a maximin control policy,

    B. Tamer and M. Max, “A multistage pursuit-evasion game that admits a gaussian random process as a maximin control policy,” Stochastics: An International Journal of Probability and Stochastic Processes, vol. 1, no. 1-4, pp. 25–69, 1973

  5. [3]

    Differential games with imperfect state information,

    I. Rhodes and D. Luenberger, “Differential games with imperfect state information,” IEEE Transactions on Automatic Control , vol. 14, no. 1, pp. 29–38, 1969

  6. [4]

    Decomposition techniques for markov zero-sum games with nested information,

    J. Zheng and D. A. Casta ˜n´on, “Decomposition techniques for markov zero-sum games with nested information,” in 52nd IEEE conference on decision and control . IEEE, 2013, pp. 574–581

  7. [5]

    Common informa- tion based markov perfect equilibria for linear-gaussian games with asymmetric information,

    A. Gupta, A. Nayyar, C. Langbort, and T. Basar, “Common informa- tion based markov perfect equilibria for linear-gaussian games with asymmetric information,” SIAM Journal on Control and Optimization , vol. 52, no. 5, pp. 3228–3260, 2014

  8. [6]

    Zero-sum stochastic games with asymmetric information,

    D. Kartik and A. Nayyar, “Zero-sum stochastic games with asymmetric information,” in 2019 IEEE 58th Conference on Decision and Control (CDC). IEEE, 2019, pp. 4061–4066

Show all 32 references
  1. [7]

    Linear-quadratic gaussian games with asymmetric information: Belief corrections using the opponents actions,

    B. Hambly, R. Xu, and H. Yang, “Linear-quadratic gaussian games with asymmetric information: Belief corrections using the opponents actions,” arXiv preprint arXiv:2307.15842 , 2023

  2. [8]

    Stochastic dynamic games in belief space,

    W. Schwarting, A. Pierson, S. Karaman, and D. Rus, “Stochastic dynamic games in belief space,” IEEE Transactions on Robotics, vol. 37, no. 6, pp. 2157–2172, 2021

  3. [9]

    Learning mixed strategies in trajectory games,

    L. Peters, D. Fridovich-Keil, L. Ferranti, C. Stachniss, J. Alonso-Mora, and F. Laine, “Learning mixed strategies in trajectory games,” arXiv preprint arXiv:2205.00291, 2022

  4. [10]

    Learning equilibria in asymmetric auction games,

    M. Bichler, N. Kohring, and S. Heidekr ¨uger, “Learning equilibria in asymmetric auction games,” INFORMS Journal on Computing , vol. 35, no. 3, pp. 523–542, 2023

  5. [11]

    Dynamic games with asymmetric information and resource constrained players with applications to security of cyberphysical systems,

    A. Gupta, C. Langbort, and T. Bas ¸ar, “Dynamic games with asymmetric information and resource constrained players with applications to security of cyberphysical systems,” IEEE Transactions on Control of Network Systems, vol. 4, no. 1, pp. 71–81, 2016

  6. [12]

    Dynamic games for secure and resilient control system design,

    Y . Huang, J. Chen, L. Huang, and Q. Zhu, “Dynamic games for secure and resilient control system design,” National Science Review , vol. 7, no. 7, pp. 1125–1141, 2020

  7. [13]

    Games with incomplete information played by “bayesian

    J. C. Harsanyi, “Games with incomplete information played by “bayesian” players, i–iii part i. the basic model,” Management science, vol. 14, no. 3, pp. 159–182, 1967

  8. [14]

    The market for “lemons

    G. A. Akerlof, “The market for “lemons”: Quality uncertainty and the market mechanism,” The quarterly journal of economics , vol. 84, no. 3, pp. 488–500, 1970

  9. [15]

    Agreeing to disagree,

    R. J. Aumann, “Agreeing to disagree,” The Annals of Statistics , vol. 4, no. 6, pp. 1236–1239, 1976

  10. [16]

    Job market signaling,

    M. Spence, “Job market signaling,” in Uncertainty in economics . Elsevier, 1978, pp. 281–306

  11. [17]

    Credit rationing in markets with imperfect information,

    J. E. Stiglitz and A. Weiss, “Credit rationing in markets with imperfect information,” The American economic review , vol. 71, no. 3, pp. 393– 410, 1981

  12. [18]

    The theory of learning in games, economics learning and social evolution series,

    D. Fudenberg and D. Levine, “The theory of learning in games, economics learning and social evolution series,” 1998

  13. [19]

    Computing best-response strate- gies in infinite games of incomplete information,

    D. Reeves and M. P. Wellman, “Computing best-response strate- gies in infinite games of incomplete information,” arXiv preprint arXiv:1207.4171, 2012

  14. [20]

    On the rate of convergence of continuous-time fictitious play,

    C. Harris, “On the rate of convergence of continuous-time fictitious play,” Games and Economic Behavior , vol. 22, no. 2, pp. 238–259, 1998

  15. [21]

    Best response dynamics for continuous zero-sum games,

    J. Hofbauer and S. Sorin, “Best response dynamics for continuous zero-sum games,” Discrete and Continuous Dynamical Systems Series B, vol. 6, no. 1, p. 215, 2006

  16. [22]

    W. H. Sandholm, Population games and evolutionary dynamics . MIT press, 2010

  17. [23]

    Best response model predictive control for agile interactions between autonomous ground vehicles,

    G. Williams, B. Goldfain, P. Drews, J. M. Rehg, and E. A. Theodorou, “Best response model predictive control for agile interactions between autonomous ground vehicles,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 2403–2410

  18. [24]

    Game theoretic motion planning for multi-robot racing,

    Z. Wang, R. Spica, and M. Schwager, “Game theoretic motion planning for multi-robot racing,” in Distributed Autonomous Robotic Systems: The 14th International Symposium . Springer, 2019, pp. 225–238

  19. [25]

    Game- theoretic planning for self-driving cars in multivehicle competitive scenarios,

    M. Wang, Z. Wang, J. Talbot, J. C. Gerdes, and M. Schwager, “Game- theoretic planning for self-driving cars in multivehicle competitive scenarios,” IEEE Transactions on Robotics , vol. 37, no. 4, pp. 1313– 1325, 2021

  20. [26]

    Controllability metrics, limitations and algorithms for complex networks,

    F. Pasqualetti, S. Zampieri, and F. Bullo, “Controllability metrics, limitations and algorithms for complex networks,” IEEE Transactions on Control of Network Systems , vol. 1, no. 1, pp. 40–52, 2014

  21. [27]

    Performance bounds for optimal and robust feedback control in networks,

    K. Ganapathy, J. Ruths, and T. Summers, “Performance bounds for optimal and robust feedback control in networks,” IEEE Transactions on Control of Network Systems , vol. 8, no. 4, pp. 1754–1766, 2021

  22. [28]

    On the decay rate of hankel singular values and related issues,

    A. C. Antoulas, D. C. Sorensen, and Y . Zhou, “On the decay rate of hankel singular values and related issues,” Systems & Control Letters , vol. 46, no. 5, pp. 323–342, 2002

  23. [31]

    S. Boyd, L. El Ghaoui, E. Feron, and V . Balakrishnan, Linear matrix inequalities in system and control theory . SIAM, 1994

  24. [32]

    All optimal Hankel-norm approximations of linear multivariable systems and their L, ∞ -error bounds†,

    Keith. Glover, “All optimal Hankel-norm approximations of linear multivariable systems and their L, ∞ -error bounds†,” International Journal of Control , vol. 39, no. 6, pp. 1115–1193, June 1984

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.