{"id":"fbd00dbe-8b35-4b2d-a88a-1c6062f9b934","arxiv_id":"2501.06181","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Best response dynamics in partially observed zero-sum linear quadratic games converge numerically after a few iterations, and low-order belief feedback strategies approximate the Nash equilibrium.","lead":"The paper studies two-player zero-sum games where players control a shared system but see different, noisy measurements. It finds that a few rounds of best-response play already get very close to the equilibrium, so simple low-memory strategies may be nearly optimal.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1 bounds Gramian approximation error, not value suboptimality, so the abstract's claim that low-order belief strategies closely approximate a Nash equilibrium is not established.","rationale":"The reader's weakest assumption is the well-posedness of the indefinite Riccati recursions. That is a real gap, but it is partially mitigated by the fact that the paper's numerical experiments do compute these quantities, and a failure would invalidate the recursion rather than the inference from Gramian decay. The most load-bearing issue for the paper's advertised conclusion is the missing value-suboptimality theorem: Theorem 1 is the only formal result invoked to support low-order approximation, and it concerns a controllability Gramian of an open-loop belief system. No inequality connects that Gramian error to the players' costs or to best-response deviation. The paper even states 'analogous statements are easily obtained for observability ... and related bounds for Hankel singular values,' but says nothing that bounds J. Therefore, even granting all Riccati solutions and value convergence, the abstract's central assertion does not follow. A direct truncation experiment would settle whether the phenomenon is real.","tokens_in":14721,"tokens_out":5087,"duration_ms":52515,"concrete_test":"Use the Section V system matrices (A, B1=B2, C1=C2, W, V, Q, R1, R2) and the converged best-response fixed point. For l = 1, 2, 3, construct l-dimensional balanced truncations of each player's converged belief filter, then simulate the truncated strategy pair (and each truncated strategy against the full-order opponent) over a long horizon with identical noise seeds, computing the average cost in (5). If the cost gap between truncated and full-order strategies does not shrink below, say, 1% of the full-order value as l increases, the claim that limited internal-state-dimension strategies closely approximate a Nash equilibrium is falsified. If the gap does shrink, that is direct evidence for the claim, independent of Gramian decay.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section IV-B, Theorem 1: for stable diagonalizable (A_i_k, B_i_k, C_i_k), the controllability Gramian W^i_{c,k} is within epsilon (in infinity-norm) of a rank-lmi Cholesky-based approximation. The text immediately concludes: 'the strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies.' That inference is not valid. A bound on the open-loop controllability Gramian of the augmented belief dynamics does not imply a bound on (i) the difference in cost functionals J^1_k and J^2_k between full and truncated strategies, (ii) the distance of the truncated strategy pair to the best-response fixed point, or (iii) the value gap to the true infinite-belief equilibrium. Standard balanced-truncation H-infinity bounds apply to a fixed LTI system with fixed inputs and outputs; here the belief dynamics embed the opponent's strategy, which changes when either player truncates, so no fixed-point or Lipschitz argument is supplied. Moreover, Section V reports only Gramian eigenvalue and Hankel singular value decay (Table I), never measuring the actual cost of truncated low-order strategies against full-order opponents. The paper's central approximation claim therefore rests on an unproved bridge from model-reduction metrics to game-value suboptimality.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies infinite-horizon zero-sum linear-quadratic dynamic games (LQDGs) with partial and asymmetric information. It proposes best-response dynamics in which each player's strategy is a linear output-feedback controller whose internal state (belief state) is an augmented state of dimension an integer multiple of the system state. The paper derives the best-response Riccati equations and estimator gains at each iteration, then analyzes controllability/observability Gramians and Hankel singular values of the augmented belief dynamics, proving a Gramian approximation bound (Theorem 1) via Cholesky factors. Numerical experiments show that the value appears to converge within a few iterations and that Gramian eigenvalues and Hankel singular values decay rapidly, leading the authors to conclude that low-order belief feedback can closely approximate Nash equilibrium strategies.","tokens_in":14960,"tokens_out":4704,"duration_ms":46961,"significance":"If the central claim were fully established, the paper would provide a tractable route to approximate equilibrium computation in asymmetric-information stochastic games, giving a concrete model-reduction perspective on higher-order beliefs. The explicit algebraic derivation of both players' best responses is a useful contribution and the numerical setup is clearly described. The paper is honest that value convergence is observed rather than proven, and the comparison of Cholesky estimates with actual Gramian eigenvalue decay rates is a legitimate consistency check. However, the inference from Gramian approximation error to value suboptimality is currently unjustified, and the existence conditions for the indefinite Riccati equations are not stated. With those gaps filled, the paper would be a valuable contribution; as it stands, the main claim goes beyond what the theorems and experiments support.","major_comments":[{"comment":"The text states only that unique solutions of the Riccati equations exist 'under certain conditions' and that the closed-loop system remains stable at each iteration, with references [29], [30]. Because Q^1_k is indefinite and R2 is negative definite, the standard stabilizability/detectability hypotheses used for the LQG Riccati equation are not sufficient; one needs explicit conditions for the existence of stabilizing solutions to the associated indefinite algebraic Riccati equations, together with the definiteness condition R2 + (B^2_k)^T P^2_k B^2_k ≺ 0 for the maximizer. These conditions must be stated and verified at every iteration; otherwise the costs J^{1*}_k and J^{2*}_k, the gains K^i_k and L^i_k, and the Gramians in Eqs. (33)–(34) are not well defined.","section":"Section III-C.1 and III-C.2, Eqs. (22), (24), (30), (32)"},{"comment":"Theorem 1 bounds the infinity-norm distance between the true controllability Gramian W^i_{c,k} and a low-rank Cholesky-based approximation; it does not bound the difference in the players' cost functionals J^{1*}_k and J^{2*}_k, the distance of the truncated strategy pair to a best-response fixed point, or the value gap between the low-order and the full-order infinite-belief equilibrium. Therefore the sentence 'the strategies based on low-order belief states provide a good approximation of the Nash equilibrium strategies' does not follow from the theorem. Section V reports only Gramian eigenvalue and Hankel singular value decay (Table I), never the actual costs of truncated low-order strategies played against full-order best responses. To support the central claim, the paper needs either a Lipschitz/fixed-point argument that connects a Gramian approximation error to suboptimality in the objective (5), or a direct numerical comparison of the costs of low-order truncated strategies against full-order best responses.","section":"Section IV-B, Theorem 1, and following paragraph"},{"comment":"The conclusion that 'this point of convergence corresponds to a Nash equilibrium' is based on visual inspection of Figure 1. No convergence proof for the best-response iteration is provided, and no limiting value is characterized; even if the sequences J^{1*}_k and J^{2*}_k appear to converge numerically, the restricted strategy class (linear output feedback with internal dimension a multiple of n) means that convergence of these costs does not by itself imply that the limiting pair is a Nash equilibrium of (5). The claim should be softened to a conjecture, or supported by a proof of convergence or a finite-iteration error bound.","section":"Section V-A and Section VI"}],"minor_comments":[{"comment":"The definitions of the Cholesky factors use inconsistent indexing: Eq. (36) writes lambda^i_{l,k}, while Eq. (37) writes lambda^i_k in the first factor and inside the product; the notation should be made uniform.","section":"Eqs. (36) and (37)"},{"comment":"The bound contains the undefined term 'mini k' and the approximation delta^i_{1,k} ≈ ||W^i_{c,k}||_2 is stated without precise constants; a precise definition of 'mini k' and a numerical verification of the bound itself would strengthen the result.","section":"Theorem 1"},{"comment":"The phrase 'analogous statements are easily obtained for the observability Gramian, and related bounds for the Hankel singular values' leaves the observability-side tool unproved; since the paper draws conclusions about both controllability and observability, a formal statement and proof (or at least a precise reference to a discrete-time analogue) should be included.","section":"End of Section IV-B"},{"comment":"There are several typos: 'maitrx' in Theorem 1, 'continous-time' in its proof, and 'the the' in the third paragraph of Section I.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a systems/control journal. The main gap is the unproved inference from Gramian model reduction to game-value suboptimality; if the authors cannot supply a proof, they should explicitly re-scope the central claim as a conjecture supported by numerical evidence. There is no apparent circularity in the Cholesky comparison, and the literature review is adequate. The derivation of best responses is a solid contribution, but the load-bearing existence conditions and the value-suboptimality bridge need to be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this paper is better than its abstract. What is new: a recursive best-response construction for infinite-horizon zero-sum LQDG with partial and asymmetric information, where each iteration augments the belief state, plus a numerical observation that the value converges in a few iterations while Gramian eigenvalues and Hankel singular values decay fast. The derivation is standard LQG separation applied recursively, but the bookkeeping is clean and the 1000-random-system study backs the phenomenon beyond a single example. The Cholesky bound in Theorem 1 is imported from Antoulas et al., and the authors mostly say so. There is no circularity: the Cholesky estimates are compared against actual decay as a consistency check, not fitted.\n\nThe central gap is the one the stress-test note names. Theorem 1 bounds the low-rank approximation error of an open-loop controllability Gramian, not the value gap between a truncated-strategy pair and the true equilibrium. The sentence in Section IV-B claiming that low-order belief strategies approximate Nash strategies does not follow from the theorem. Balanced truncation bounds apply to a fixed LTI system with fixed inputs and outputs; here the opponent's strategy changes when either player truncates, and no fixed-point or Lipschitz argument is supplied. The numerics report Gramian decay but never measure the actual cost incurred by low-order strategies against full-order opponents. That is a real soft spot, and it is load-bearing if the abstract's approximation claim is read as a theorem.\n\nThe second soft spot is well-posedness. The paper says unique Riccati solutions exist \"under certain conditions\" but never states the stabilizability/detectability or indefinite-cost conditions needed for equations (22), (24), (30), (32). With R2 negative and the Q1 indefinite, a stabilizing solution is not automatic. If those equations fail at some iteration, the cost sequence is undefined and the convergence claim collapses. The remark that R2 must be sufficiently negative is not quantified. This is fixable by stating explicit assumptions on the augmented pairs and the Riccati solutions before claiming convergence.\n\nThird, minor: no code, precise random-instance generation, or error bars are provided. For an empirical claim that matters, but it is minor compared with the first two issues.\n\nWho gets value: people working on belief-space dynamic games and LQG model reduction. This is a plausible lead, not a solved theorem. I would send it to review because the recursive construction is useful and the numerical phenomenon is worth checking. A serious referee should ask for a revised Section IV and conclusion that stops overclaiming, plus explicit well-posedness hypotheses.","headline":"Useful recursive best-response construction for asymmetric LQG games, with a suggestive numerical phenomenon, but the low-order approximation claim outruns what Theorem 1 proves.","tokens_in":15455,"tokens_out":1851,"would_cite":true,"duration_ms":18983,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A25","93E20","93B11"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that in zero-sum stochastic linear-quadratic dynamic games with partial and asymmetric information, best-response dynamics converge in value within a few iterations, so low-dimensional belief-state controllers can closely…","keywords":["zero-sum games","asymmetric information","best response dynamics","higher-order beliefs","Gramian eigenvalue decay","Hankel singular values","linear quadratic Gaussian","model reduction"],"falsifier":"Run the best-response recursion on an open-loop-stable system whose dynamics matrix has eigenvalues very close to the unit circle, so the Cholesky decay factors are large; if the game value still changes materially after many iterations, or if the Hankel singular values fail to fall below the tolerance used in the low-order approximation, the central claim fails. A sharper counterexample would be any instance where at some iteration the pair $(A^k_i, B^k_i)$ is not stabilizable or $(A^k_i, C^k_i)$ is not detectable, making the Riccati solution and hence the approximate value undefined.","tokens_in":14507,"feed_emoji":"🎲","tokens_out":8542,"duration_ms":71347,"temperature":0.7,"pith_summary":"This paper asks whether players in an infinite-horizon zero-sum stochastic linear-quadratic dynamic game with partial and asymmetric information can settle on strategies that use only low-dimensional belief states. It derives the best-response recursion: with each round, a player's optimal controller tracks not only the state but also the opponent's belief, so the controller's internal dimension grows. The paper reports that the game value converges within a few rounds across numerical experiments, and attributes this to rapid decay of controllability and observability Gramian eigenvalues and Hankel singular values of the higher-order belief dynamics. If true, the practical consequence is that simple feedback strategies with limited internal state dimension can closely approximate a Nash equilibrium.","feed_headline":"Few belief levels nearly reach Nash equilibrium in hidden-state games","feed_subtitle":"Value converges after a few belief rounds, so compact controllers can approximate equilibrium play.","key_machinery":"The central object is the augmented belief-state system built at each best-response step: the opponent's previous controller is folded into the dynamics, output map, and cost weighting, so the acting player must estimate a state of dimension $(2k-1)n$ or $2kn$. The carrier of the argument is the pair of algebraic Riccati equations for control and estimation—equations (22) and (24) for the minimizer, (30) and (32) for the maximizer—which define each best response, together with the Lyapunov equations for the controllability and observability Gramians and their Hankel singular values. Theorem 1, using Cholesky factors of a Cauchy matrix built from the eigenvalues of the augmented dynamics matrix and a bilinear transformation, converts rapid eigenvalue decay into a quantitative low-rank approximation bound, which is what supports the claim that low-order belief dynamics are nearly as good as the full hierarchy.","core_discovery":"The central claim is that in an infinite-horizon zero-sum linear-quadratic dynamic game where each player observes a private noisy linear measurement, best-response iteration within the class of pure linear output-feedback strategies produces controllers of increasing internal dimension—one dimension per order of belief—yet the value stabilizes after a few iterations. The paper derives explicit recursions: the minimizer at iteration $k$ faces an augmented state of dimension $(2k-1)n$, and the maximizer one of dimension $2kn$, with the gains and costs determined by Riccati equations (22), (24), (30), and (32). It then argues that the higher-order belief dynamics become increasingly hard to control and observe, using controllability and observability Gramians and Hankel singular values, and Theorem 1 bounds the low-rank approximation error of the controllability Gramian by Cholesky estimates of eigenvalue decay. The paper concludes that low-order belief dynamics approximate the infinite-dimensional equilibrium strategies with bounded error.","pith_inferences":["Beyond the paper, an extension would be to run the same recursion on N-player nonzero-sum LQG games, where value convergence can fail even if the decay mechanism is similar; the numerical evidence suggests but does not prove such a result.","The decay rates observed in the paper suggest a model-selection heuristic: choose the internal controller dimension at the iteration where the Hankel singular value curve flattens, connecting this game-solving method to balanced-truncation model reduction.","A testable boundary condition is to construct systems with eigenvalues near the unit circle, where the Cholesky decay factors degrade, and check whether value convergence slows; this would expose how general the rapid-decay phenomenon is."],"forward_implications":["In systems where the decay condition in Theorem 1 holds, a controller of fixed internal dimension built from the first few belief orders achieves cost within an explicit error bound of the infinite-dimensional equilibrium.","Best-response iteration can be stopped once the game value stops changing, giving finite-order strategies with near-equilibrium performance and avoiding the infinite belief hierarchy.","The Gramian and Hankel singular value decay rates provide a computable criterion for when higher-order beliefs are unnecessary: when the Cholesky ratio $\\delta^k_l/\\delta^k_1$ falls below a tolerance, the remaining belief directions contribute little.","The result extends the practical reach of LQG control to asymmetric-information games by showing that the infinite regress of beliefs can be truncated without much performance loss.","If the paper is right, the Nash equilibrium of the original game is well approximated by a finite-dimensional linear output-feedback strategy pair."],"supporting_citations":[{"why":"Supplies the dynamic game theory foundation: the zero-sum Nash equilibrium concept and the strategy framework the paper works within.","marker":"[1]"},{"why":"Provides the Cholesky-factor eigenvalue decay bounds (Theorem 3.2) that Theorem 1 adapts to bound Gramian approximation error.","marker":"[28]"},{"why":"Underpins the algebraic Riccati equation solutions used to compute the LQG best response at each iteration.","marker":"[29]"},{"why":"Supports the claim that the closed-loop system under the state-feedback strategy remains stable at each iteration.","marker":"[30]"},{"why":"Gives the bilinear transformation used to map discrete-time Gramians into the continuous-time setting of the Cholesky decay estimates.","marker":"[32]"},{"why":"Relates Hankel singular values to Gramian eigenvalue products, motivating their use as a decay metric for belief dynamics.","marker":"[31]"}],"fun_headline_variants":["Compact controllers nearly solve hidden-state games after few belief steps","Belief rounds converge fast: low-order strategies hit Nash","Few belief levels enough: value stabilizes in hidden-state games","Scale down belief states: few rounds suffice for Nash in hidden-info games"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that at every best-response iteration the augmented system is stabilizable and detectable, so that the Riccati equations (22), (24), (30), and (32) admit stabilizing solutions even though the minimizer's state cost is indefinite and the maximizer's input cost is negative definite; the paper states this only as 'certain conditions' without specifying them.","fun_headline_variants_meta":{"raw":{"variants":["Compact controllers nearly solve hidden-state games after few belief steps","Belief rounds converge fast: low-order strategies hit Nash","Few belief levels enough: value stabilizes in hidden-state games","Scale down belief states: few rounds suffice for Nash in hidden-info games"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":2996,"prompt_tokens":924,"completion_tokens":2072,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":2001}},"tokens_in":540,"tokens_out":2072,"duration_ms":15253,"temperature":1.0,"reasoning_tokens":2001,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:59.467627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the best-response recursion on an open-loop-stable system whose dynamics matrix has eigenvalues very close to the unit circle, so the Cholesky decay factors are large; if the game value still changes materially after many iterations, or if the Hankel singular values fail to fall below the tolerance used in the low-order approximation, the central claim fails. A sharper counterexample would be any instance where at some iteration the pair $(A^k_i, B^k_i)$ is not stabilizable or $(A^k_i, C^k_i)$ is not detectable, making the Riccati solution and hence the approximate value undefined.","supporting_citations":[{"cited_title":"On the decay rate of hankel singular values and related issues,","cited_arxiv_id":null,"evidence_quote":"Provides the Cholesky-factor eigenvalue decay bounds (Theorem 3.2) that Theorem 1 adapts to bound Gramian approximation error."},{"cited_title":"Least squares stationary optimal control and the algebraic Riccati equation,","cited_arxiv_id":null,"evidence_quote":"Underpins the algebraic Riccati equation solutions used to compute the LQG best response at each iteration."},{"cited_title":"The stable regulator problem and its inverse,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that the closed-loop system under the state-feedback strategy remains stable at each iteration."},{"cited_title":"All optimal Hankel-norm approximations of linear multivariable systems and their L, ∞ -error bounds†,","cited_arxiv_id":null,"evidence_quote":"Gives the bilinear transformation used to map discrete-time Gramians into the continuous-time setting of the Cholesky decay estimates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Relates Hankel singular values to Gramian eigenvalue products, motivating their use as a decay metric for belief dynamics."}],"review_version":1}