REVIEW 3 major objections 5 minor 1 cited by
On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ games
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Finite-horizon 'look-ahead' strategies approximate infinite-horizon feedback Nash equilibria, with an explicit cubic cost-gap bound.
desk verdict A clean conditional result: finite-horizon LQ game strategies approximate an infinite-horizon FNE with an explicit cubic cost-gap bound, but the main assumption (backward Riccati convergence) is unproven and the simulation does not match the theorem's hypotheses. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the coupled generalized discrete Riccati difference equations (3)--(5), plus their affine partners (4), (6), (7) in the finite-horizon case. At each backward step, given $P^j_{t+1}$, the feedback gains come from one linear system $H(P_{t+1})K_t = g(P_{t+1})$, and the same matrix $H(P_{t+1})$ also enters the affine terms; the cost matrices then update by simple quadratic formulas. Because the horizon $T$ only shifts the iteration count, the first-stage matrices $K^{i*}_1(T^i)$ are constant over time, which lets each player reuse the same gain at every stage. Letting the backward iteration run to $-\infty$ turns these recursions into coupled algebraic Riccati equations whose solution supplies the limiting FNE, and the proof of Theorem 3 bounds each term of the infinite cost sum by comparing $(\tilde F^*,\tilde G^*)$ with $(F^*,G^*)$ through telescoping powers and the stability margin $\lambda + b\epsilon < 1$.
What would settle it
Choose a small LQ game satisfying the invertibility and stability parts of Assumption 1 and run the backward Riccati iteration (3)--(5) from $P^i_1=0$; if the iterates never settle at finite limits, Assumption 1(ii) fails and the theorem's conclusion is not guaranteed. For any instance where Assumption 1 holds, simulate the finite-horizon strategy for growing common horizon and compare the empirical cost gap with the right-hand side of the Theorem 3 inequality; a violation would falsify the stated bound.
Extended reading notes
Core claim
The central claim is that the 'look $T^i$ steps ahead, move one step' strategy is a certified approximation of one infinite-horizon FNE. Formally, with $\epsilon = \max_i \|K^{i*}_1(T^i) - K^{i*}\|_2$ and $\theta_i(\epsilon)$ a cubic polynomial in $\epsilon$ built from fixed model parameters, Theorem 3 states that when $\|F^*\|_2 + (\sum_j \|B^j\|_2)\epsilon < 1$, the per-player cost gap satisfies $$|\tilde{J}^i(x_1) - J^i(x_1)| \le \tfrac12 \|x_1\|$_2^{2}$\, \theta_i(\epsilon)/(1-\delta^i),$$ and $\epsilon \to 0$ as $T_h = \min_i T^i \to \infty$. The limit matrices $(K^{i*},P^{i*})$ are the fixed point reached by the backward Riccati recursion and solve the coupled algebraic Riccati equations, so the limiting strategies form an FNE. A sequence of easily computed finite-horizon equilibria therefore approximates a generally hard infinite-horizon equilibrium, with a quantitative performance certificate.
Load-bearing premise
The whole guarantee rests on an assumed convergence: as the backward Riccati recursion is run further, the strategy and cost matrices must settle at finite limits, and the paper does not prove when that happens. The paper itself leaves parameter conditions for this convergence as an open question.
Editorial extensions
If this is right
- Replacing the coupled algebraic Riccati equations by a sequence of linear solves makes the finite-horizon FNE computable in practice; Algorithm 1 verifies the invertibility condition as it runs.
- As the shortest prediction horizon grows, the first-stage gains approach the limiting gains and each player's total cost approaches the infinite-horizon FNE cost.
- The cost gap is bounded by a cubic polynomial in the strategy-matrix distance, so small gain errors translate into small cost errors, and the bound vanishes as $T_h\to\infty$.
- The result covers heterogeneous discount factors $\delta^i\in(0,1]$, so players with different patience levels may use different look-ahead lengths.
Reading between the lines
- Editorial extension: if the Riccati iteration converges geometrically, then $\epsilon$ decays geometrically and the cubic bound implies geometric decay of the cost gap; the paper does not quantify the rate, but rate analysis would yield a horizon-selection rule.
- Editorial extension: the first-stage-control scheme could be tried on time-varying or nonlinear dynamics, but the telescoping-power argument is specific to LQ structure, so a different proof would be needed.
- Editorial extension: since the bound carries $1/(1-\delta^i)$, patient players need longer horizons (or tighter gain convergence) to reach the same relative accuracy; heterogeneous horizons could be calibrated to discount factors.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a receding-horizon approximation of infinite-horizon feedback Nash equilibria (FNEs) in discrete-time linear-quadratic games. For finite-horizon games with input/output/state dynamics, the authors derive a sufficient invertibility condition under which the unique FNE can be computed by solving a sequence of linear equations (Algorithm 1). For the infinite-horizon game, they consider a 'look-T^i-steps-ahead, move-one-step' strategy and prove (under Assumption 1) that the resulting total costs converge to the limiting FNE cost as all prediction horizons go to infinity, with an explicit upper bound on the cost gap in terms of the deviation of the strategy matrices. A numerical example with two players illustrates the convergence.
Significance. The paper addresses a computationally challenging problem: solving coupled algebraic Riccati equations for infinite-horizon LQ games. The proposed finite-horizon strategy is tractable, and the error bound is derived rather than fitted. The linear-equation algorithm for finite-horizon games is a useful practical contribution. However, the main theorem relies on a strong convergence assumption (Assumption 1(ii)) that is acknowledged as an open problem, and the numerical example does not directly match the theoretical setup; these issues substantially temper the significance unless addressed.
major comments (3)
- [Section 3, Assumption 1(ii)] Assumption 1(ii) is the load-bearing hypothesis: Lemma 2 identifies the infinite-horizon FNE with the limit of the backward Riccati recursion, and Theorem 3's limit and bound require epsilon -> 0, which follows only from this convergence. The paper does not prove Assumption 1(ii) and states in the Conclusions that parameter-based conditions guaranteeing it remain open. Since the hypothesis is exactly the difficult part of the problem, the central guarantee is conditional on an unverifiable condition. Please either prove the convergence for a nontrivial class of games, provide a checkable sufficient condition, or restructure the paper's claims to make the conditional nature explicit in the title and abstract.
- [Section 4, numerical example] The infinite-horizon theory in Section 3 assumes zero reference trajectories (l_i^t = 0), as in the cost (12). The numerical example, however, uses constant nonzero references l_1 = [1,1]^T and l_2 = [-1,-1]^T. Thus the simulation does not validate the theorem's hypotheses; it illustrates convergence for a different (affine) problem. Please either extend the theoretical analysis to cover constant affine terms or modify the example to the zero-reference case so that the simulation directly tests Theorem 3.
- [Theorem 3 and its bound] Theorem 3 states that the bound is a cubic polynomial theta_i(epsilon) = theta_i1 epsilon + theta_i2 epsilon^2 + theta_i3 epsilon^3 with coefficients determined by fixed parameters. However, the definitions of G_i1 and G_i2 given after the theorem involve M = max_{t in N+} (t-1)(lambda + b epsilon)^{t-2}, which depends on epsilon. Consequently, the coefficients are not fixed; they vary with epsilon, and M diverges as lambda + b epsilon -> 1. This makes the advertised 'parameter-free' cubic bound inaccurate. Please either provide a uniform upper bound on M over the admissible range of epsilon (which would need to be stated as an additional condition), or revise the statement to make the epsilon-dependence of the coefficients explicit.
minor comments (5)
- [Appendix A, proof of Lemma 2] In the proof, the limits are written as t -> +infinity but should be T -> +infinity (e.g., 'lim_{t->+infinity} K_i^*(T)'). This is a typographical error that may confuse readers.
- [Section 3, Fact 2] The identity 'K_i^*(T) = K_i^*(T0-T+t(T0))' is not properly defined; the subscripts are missing. Please clarify the notation so that the dependence on the stage index is explicit.
- [Section 3, Assumption 1] Assumption 1 begins with sequences indexed by t = 0, -1, -2, ... and an initial condition P_i^1 = 0. It would be helpful to explain that this index denotes time-to-go in the backward iteration, to avoid confusion with the physical time index used elsewhere.
- [Section 2, Lemma 1] The proof of Lemma 1 is omitted with a reference to standard results. A concise sketch or a more precise citation than 'modifying Corollary 6.1 in [4]' would improve accessibility.
- [Section 4, Figure 1] The caption states that the horizontal dashed lines represent the limiting FNE strategy matrices; since the paper does not prove convergence for this example, it would be helpful to state explicitly that these are the limits of the iterations as T -> infinity, not independently verified equilibria.
Circularity Check
No circularity: the convergence and cost-gap claims are conditional on an explicit, independently-stated assumption, and the key external theorem is cited from a disjoint author group.
full rationale
The paper's main result (Theorem 3) is a conditional statement: given Assumption 1, the cost gap is bounded by an explicit cubic function of epsilon, and epsilon tends to zero when the Riccati iteration converges. The convergence of the backward Riccati recursion (Assumption 1(ii)) is not proven; the paper explicitly states in Section 5 that parameter-based conditions guaranteeing it remain open. This is an unproved hypothesis and a limitation, not a circular reduction: the theorem does not define convergence into existence, and it does not fit any parameter to data. Lemma 2 uses Assumption 1(ii) to identify the limit of the Riccati recursion with an FNE; the FNE verification is delegated to Theorem 3.2 of Monti et al. [6], whose authors are disjoint from the present authors, so the load-bearing citation is external and not self-referential. The bound in Theorem 3 is derived by direct norm inequalities from the definition of epsilon, not by assuming the conclusion. The 'limiting FNE' is deliberately selected as the equilibrium corresponding to the iteration limit, so the paper claims approximation of that particular equilibrium, which is a legitimate model choice rather than a renaming of the conclusion. No fitted-input-called-prediction, no uniqueness-imported-from-authors, and no ansatz-smuggling pattern is present. Overall, the derivation is self-contained as a conditional theorem; the open convergence question is a correctness or completeness risk, not circularity.
Assumptions & free parameters
assumptions (5)
- ad hoc to paper Assumption 1(ii): the coupled generalized discrete Riccati difference equations converge as t -> -infinity to finite limits P_i^*, K_i^*.
- domain assumption Assumption 1(i): |H(P_{t+1})| != 0 for all backward iteration steps.
- domain assumption Assumption 1(iii): ||A + sum_i B_i K_i^*||_2 < 1.
- standard math Lemma 1, cited from Basar and Olsder (Corollary 6.1): the Riccati difference equations characterize the unique FNE.
- standard math Theorem 3.2 of Monti et al. [6]: classical LQR optimality for the single-player control problem with other players' strategies fixed.
Cite this review
Pith. "Pith review of On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ games." pith.science (2026). https://pith.science/paper/R562YQ3E
@misc{pith2026250619565,
author = {Pith},
title = {Pith review of: On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ games},
year = {2026},
howpublished = {\url{https://pith.science/paper/R562YQ3E}},
note = {Machine review of arXiv:2506.19565}
}
abstract
In infinite-horizon discrete-time linear-quadratic (LQ) dynamic games, computing feedback Nash equilibria (FNEs) remains computationally challenging. Motivated by this, we study a finite-horizon strategy for approximating one of the infinite-horizon FNEs. The finite-horizon strategy is as follows. Each player $i$ has an individual prediction horizon $T^i$. In the infinite-horizon game, at each stage, each player $i$ computes its control in the following way: player $i$ envisions an auxiliary $T^i$-stage game in which the same set of players play, computes the unique FNE of the auxiliary game using a standard method, and implements only the first-stage control. Our main result is, under suitable conditions, the total cost under these finite-horizon strategies converges to that under one of the infinite-horizon FNEs when all players' prediction horizons tend to infinity. Moreover, we derive an explicit cubic-polynomial upper bound on this cost gap with respect to the distance between the corresponding strategy matrices. This strategy is tractable and implementable, as it avoids the direct solution of the coupled algebraic Riccati equations (CARE) of infinite-horizon LQ games.
Figures
Forward citations
Cited by 1 Pith paper
-
Linear-Quadratic Discrete-Time Dynamic Games with Unknown Dynamics
Offline input/output data can replace model knowledge in computing feedback Nash equilibria of linear-quadratic discrete-time games, with a finite-horizon equivalence result and a convergent rolling-horizon approximation.
Reference graph
Works this paper leans on
- [1]
- [2]
-
[3]
Y. Li, G. Carboni, F. Gonzalez, D. Campolo, E. Burdet, Differential game theory for versatile physical human–robot interaction, Nat. Mach. Intell. 1 (2019) 36–43
work page 2019
- [4]
-
[5]
B. Nortmann, A. Monti, M. Sassano, T. Mylvaganam, Nash equilibria for linear quadratic discrete-time dynamic games via iterative and data-driven algorithms, IEEE Trans. Autom. Control 69 (2024) 6561– 6575
work page 2024
- [6]
-
[7]
B. Nortmann, M. Sassano, T. Mylvaganam, Feedback nash equilibria for scalar n-player linear quadratic dynamic games, Automatica 174 (2025) 112133
work page 2025
-
[8]
Siniscalchi, Structural rationality in dynamic games, Economet- rica 90 (2022) 2437–2469
M. Siniscalchi, Structural rationality in dynamic games, Economet- rica 90 (2022) 2437–2469
work page 2022
Show all 27 references
-
[9]
H. Sieg, C. Yoon, Estimating dynamic games of electoral competition to evaluate term limits in us gubernatorial elections, Am. Econ. Rev. 107 (2017) 182457
2017
-
[10]
Kremer, J
M. Kremer, J. Willis, Guns, latrines, and land reform: Dynamic pigouvian taxation, Am. Econ. Rev. 106 (2016) 8388
2016
-
[11]
Jang, K.-Y
I. Jang, K.-Y. Kang, Dynamic adverse selection and belief update in credit markets, J. Financ. Quant. Anal. 60 (2025) 19942025
2025
-
[12]
Zhang, Z
Y. Zhang, Z. Yang, Dynamic incentive contracts for esg investing, J. Corp. Financ. 87 (2024) 102614
2024
-
[13]
Miikkulainen, S
R. Miikkulainen, S. Forrest, A biological perspective on evolutionary computation, Nat. Mach. Intell. 3 (2021) 9–15
2021
-
[14]
M. A. Nowak, K. Sigmund, Evolutionary dynamics of biological games, Science 303 (2004) 793–799
2004
-
[15]
N. Liu, L. Guo, Adaptive stabilization of noncooperative stochastic differential games, SIAM J. Control Optim. 62 (2024) 1317–1342
2024
-
[16]
Sadana, P
U. Sadana, P. V . Reddy, G. Zaccour, Feedback nash equilibria in differential games with impulse control, IEEE Trans. Autom. Control 68 (2023) 4523–4538
2023
-
[17]
O. L. Costa, A. M. de Oliveira, Feedback linear quadratic nash equilibrium for discrete-time markov jump linear systems, Syst. Control Lett. 192 (2024) 105893
2024
-
[18]
Gravell, K
B. Gravell, K. Ganapathy, T. Summers, Policy iteration for linear quadratic games with stochastic parameters, IEEE Control Syst. Lett. 5 (2021) 307–312
2021
-
[19]
Z. Xu, T. Shen, M. Huang, Model-free policy iteration approach to nce-based strategy design for linear quadratic gaussian games, Automatica 155 (2023) 111162
2023
-
[20]
Y. Guan, G. Salizzoni, M. Kamgarpour, T. H. Summers, A policy iteration algorithm for n-player general-sum linear quadratic dynamic games, in: 2024 IEEE 63rd Conference on Decision and Control (CDC), 2024, pp. 1725–1730. doi:https://doi.org/10.1109/CDC56724. 2024.10886048
2024
-
[21]
V . Y. Glizer, V . Turetsky, Approximate solution of a zero-sum linear-quadratic differential game by the auxiliary parameter method: analytical/numerical study, Optim. 74 (2025) 4627–4667
2025
-
[22]
Laine, D
F. Laine, D. Fridovich-Keil, C.-Y. Chiu, C. Tomlin, The computation of approximate generalized feedback nash equilibria, SIAM J. Optim. 33 (2023) 294–318
2023
-
[23]
Kamalapurkar, J
R. Kamalapurkar, J. R. Klotz, W. E. Dixon, Concurrent learning- based approximate feedback-nash equilibrium solution of n-player nonzero-sum differential games, IEEE/CAA Journal of Automatica Sinica 1 (2014) 239–247
2014
-
[24]
Nortmann, T
B. Nortmann, T. Mylvaganam, Approximate nash equilibria for discrete-time linear quadratic dynamic games, IFAC-PapersOnLine 56 (2023) 1760–1765. 22nd IFAC World Congress
2023
-
[25]
Mylvaganam, M
T. Mylvaganam, M. Sassano, A. Astolfi, Constructive 𝜖-nash equilib- ria for nonzero-sum differential games, IEEE Trans. Autom. Control 60 (2015) 950–965
2015
-
[26]
Engwerda, Algorithms for computing nash equilibria in determin- istic lq games, Comput
J. Engwerda, Algorithms for computing nash equilibria in determin- istic lq games, Comput. Manag. Sci. 4 (2007) 113–140
2007
-
[27]
R. R. Bitmead, M. R. Gevers, I. R. Petersen, R. Kaye, Monotonicity and stabilizability- properties of solutions of the riccati difference equation: Propositions, lemmas, theorems, fallacious conjectures and counterexamples, Syst. Control Lett. 5 (1985) 309–315. Page 10 of 10
1985
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.