{"id":"3921502e-ac9b-48a5-a5d7-f96be95439fb","arxiv_id":"2506.22073","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Offline input/output data can replace model knowledge in computing feedback Nash equilibria of linear-quadratic discrete-time games, with a finite-horizon equivalence result and a convergent rolling-horizon approximation.","lead":"The paper develops data-driven algorithms for linear-quadratic dynamic games in which players do not know the system dynamics or the state. It shows that offline input/output records can stand in for model knowledge, with finite-horizon equilibrium characterizations and an infinite-horizon convergence analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 3.3's proof does not bridge the gap between state-feedback strategies in the known game and history-dependent strategies in the unknown game; feasible-set equality alone is insufficient.","rationale":"The reader's weakest-assumption analysis correctly identified the information-structure gap in Theorem 3.3's proof. I agree that feasible-set equivalence, without an argument relating the best-response maps across the two strategy spaces, does not prove equality of FNE sets. This is the most load-bearing concern because Theorem 3.3 underlies the data-driven finite-horizon characterization and is also invoked in Theorem 4.2 (via Remark 3.4) to connect unknown- and known-dynamics FNE matrices. The concern is partially mitigable—one can simulate any history-dependent deviation with a state-feedback strategy because the state is a function of the history—but that argument is absent from the manuscript, and the present proof is not rigorous. I additionally note a separate, concrete error: Theorem 4.2(3) claims an O(max_i ||K_i^*(T_i)−K_i^*||^2) total-cost error, yet Appendix D's proof yields only a first-order bound C·max_i ||K_i^*(T_i)−K_i^*||. This rate mismatch does not invalidate the convergence itself but contradicts the stated theorem. Both issues support retaining the reader's conditional verdict: the equivalence proof must be repaired (or the claim restricted), and the convergence-rate statement must be corrected.","tokens_in":28409,"tokens_out":19155,"duration_ms":209157,"concrete_test":"Attempt an explicit proof that any history-dependent deviation in the unknown game can be simulated by a state-feedback strategy in the known game that produces the same cost along the realized path, using that xt is a deterministic function of Ut−1. If the simulation succeeds, Theorem 3.3 holds but the proof must be amended; alternatively, search for a two-player, two-stage LQ parameterization where a state-feedback FNE of the known game is not an FNE when players may use full-history strategies. If such a counterexample exists, Theorem 3.3 is false.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Theorem 3.3 claims that the finite-horizon unknown-dynamics game and the known-dynamics game have the same set of feedback Nash equilibria. The proof in Appendix A(b) says, 'Since the cost functions for all players are the same in both settings, it suffices to prove the equivalence of their feasible sets.' This is not sufficient: the games have different information structures. In the known game (Definition 2.1), each player's strategy is a map from the current state xt, i.e., γi_t(xt)=ui_t. In the unknown game (Definition 3.2), strategies are maps from the full history col(uini,yini,u1,...,ut−1). Equality of feasible output-input trajectory sets does not imply equality of best-response correspondences over these different strategy spaces. A player in the history-dependent game could condition on past inputs in ways that a state-feedback strategy cannot replicate, potentially creating new equilibria or destroying known ones. Part (a) proves that any FNE of the unknown game is affine in the history, which narrows the strategy space, but it does not show that the enlarged history-dependent strategy space yields no additional equilibria, nor does the proof compare the first-order conditions of the two dynamic programs (equations (1-1)–(1-5) versus the standard Riccati equations under the state mapping (3.2)). Thus the central equivalence result is not established as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies N-player discrete-time linear-quadratic dynamic games in which the underlying input/output/state (i/o/s) dynamics and the state trajectory are unknown, and players have access only to an offline input/output dataset and an initial trajectory. The main finite-horizon results are: (Theorem 3.3) the set of feedback Nash equilibria of the unknown-dynamics game coincides with that of the corresponding known-dynamics game; (Theorem 3.7) necessary and sufficient conditions for existence and uniqueness of FNE in terms of coupled data-dependent equations; and (Theorem 3.8) an invertibility condition under which one FNE is obtained by solving linear equations backward. For the infinite-horizon game, Theorem 4.2 claims, under controllability and an assumption from the companion paper [11], that the 'watching T steps into the future and moving one step now' strategy yields convergence of feedback matrices and total costs, with rate O(max_i ||K_i^*(T_i)-K_i^*||^2). A two-player non-scalar numerical example illustrates the convergence behavior.","tokens_in":28654,"tokens_out":9297,"duration_ms":104655,"significance":"If the equivalence theorem were established, the paper would provide a genuinely data-driven approach to LQ dynamic games without system identification or online exploration, with finite-horizon characterizations and an explicit computational algorithm. The strengths include the behavioral-system formulation, the affine-form derivation for history-dependent equilibria, the invertibility-based linear-equation algorithm, and the non-scalar numerical demonstration. However, the central equivalence claim currently rests on an information-structure gap, and the infinite-horizon convergence rate is not supported by the proof as written. The paper is therefore of interest to the data-driven control and dynamic-game communities, but it requires substantial revision before the claims can be accepted.","major_comments":[{"comment":"The proof of Theorem 3.3 claims that equality of the feasible input-output trajectory sets is sufficient for equality of the FNE sets because the cost functions are the same in both settings. This is not sufficient: Definition 2.1 restricts players to current-state feedback strategies, whereas Definition 3.2 allows strategies that depend on the full history U_{t-1}. Equality of feasible trajectory sets says nothing about best-response correspondences or subgame-perfect behavior under these different information structures. The affine-form result in part (a) narrows the unknown-game strategy space, but it does not show that every affine history-dependent equilibrium is realizable by a state-feedback equilibrium, nor that no additional history-dependent equilibria exist. This gap is load-bearing for Theorem 3.3, Remark 3.4, and the applications in Theorem 4.2.","section":"Theorem 3.3 and Appendix A(b)"},{"comment":"Theorem 4.2(3) states that \\tilde{J}^i - J^i = O(max_i ||K_i^*(T_i)-K_i^*||^2), but the proof's final displayed inequality is |\\tilde{J}^i - J^i| <= C max_i ||K_i^*(T_i)-K_i^*||_2, which is linear in the matrix error. The squared rate is therefore not established and is contradicted by the proof as written. The rate must be corrected to O(epsilon) or a separate argument for the quadratic rate must be supplied.","section":"Theorem 4.2(3) and Appendix D"},{"comment":"Theorem 4.2 and its proof depend on Assumption 1, Proposition 1, Lemma 2, and Theorem 1 of the companion paper [11], but Assumption 1 is not stated or summarized in this manuscript. Because the convergence claims for the infinite-horizon game are entirely conditional on these unstated premises, the statement is not self-contained and the applicability of those conditions to the present unknown-dynamics setting is not verifiable from the text. Please state the assumption explicitly and confirm that the cited results apply to the games defined here.","section":"Theorem 4.2 and Assumption 1 of [11]"}],"minor_comments":[{"comment":"The symbol B is used both for the behavioral system and for the basis matrix B(B_{Tini}); please use distinct notation, such as \\mathcal{B} for the system, to avoid confusion.","section":"Notation, Section 2.2 and after Theorem 3.7"},{"comment":"The matrix B(B_{Tini}) appears in equation (1-1a) before it is formally introduced in the notation block after Theorem 3.7; please move the definition earlier or add a forward reference.","section":"Equation (1-1a)"},{"comment":"Remark 3.4 contains the identity L_i^*_t = L_i^*_t, which appears to be a typo; presumably one side should refer to the known-dynamics constant term and the other to the unknown-dynamics constant term.","section":"Remark 3.4"},{"comment":"The caption states 't = 1,...,100' while the horizontal axis of the figure is T = 1,...,50; please clarify which quantity varies in the plot.","section":"Figure 1 caption"},{"comment":"Algorithm 4.1 uses the same symbol J~i* both as an accumulated sum and as the output total cost; please distinguish the iteration variable from the final output for clarity.","section":"Section 5, Algorithm 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of eess.SY and addresses a timely problem. The main concern is the information-structure gap in Theorem 3.3, which is central: Theorem 4.2 inherits it. The rate discrepancy in Theorem 4.2 is easy to fix but must be corrected. The heavy reliance on the companion paper [11] is acceptable only if the assumptions are stated; otherwise the infinite-horizon claims remain unverifiable from this manuscript alone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution to the data-driven dynamic-games subfield. The finite-horizon result — necessary and sufficient data-based conditions for existence/uniqueness of feedback Nash equilibria when the state and dynamics are all unknown — is new as far as I can tell from the cited literature. The coupled equations (1-1)–(1-5), the invertibility condition, and the backward linear-equation algorithm are clean and plausible. The non-scalar numerical example is honest and supports the qualitative claims. Credit where due: the authors know the behavioral-systems machinery and use it carefully, and the finite-horizon characterization does not assume its own conclusion. The self-citation to the companion paper [11] is heavy but not inherently disqualifying; the transferred infinite-horizon convergence result is the kind of thing one might reasonably build on, provided the companion is sound.\n\nThe soft spot is Theorem 3.3, and it is load-bearing. The proof in Appendix A(b) says that since the costs are the same, it suffices to prove the feasible trajectory sets coincide. That is not sufficient. The known-dynamics game restricts strategies to maps from the current state; the unknown-dynamics game allows strategies on the full input/output history. Equality of feasible input-output trajectories does not imply equality of equilibrium strategy sets across those two different information structures. Part (a) of the proof does show that any FNE of the unknown game is affine in the history, which is a useful narrowing, but it does not show that the enlarged strategy space produces no additional equilibria, nor that the affine-coefficient equations are exactly the known-dynamics Riccati equations. The later relations (3.1)–(3.3) and Remark 3.4 appear to presuppose the very equivalence being proved. Appendix D then leans on Theorem 3.3 and on Assumption 1 of [11], which is never stated in this paper, so the infinite-horizon convergence result inherits both gaps. On the plus side, Theorem 3.7's coupled-equation characterization may be salvageable independently of the equivalence claim; the stress-test concern does not obviously sink that part.\n\nMinor issues: Algorithm 4.1's tolerance epsilon and iteration cap M are not tied to any theoretical stopping guarantee, and the numerical example uses nonzero reference trajectories even though Theorem 4.2 assumes zero references. The paper itself acknowledges this last point, so I treat it as a minor loose end, not a concealment.\n\nWho is this for? Researchers working on data-driven control, behavioral systems, and dynamic games. It deserves a serious referee — the finite-horizon machinery is valuable and the central equivalence theorem is plausible and likely repairable. I would accept for peer review and ask for a rigorous proof or a clearly stated restriction of Theorem 3.3, plus explicit statements of the assumptions imported from [11]. I would not cite it in my own work until that repair lands.","headline":"Genuinely new data-driven FNE characterization for finite-horizon LQ games with unknown i/o/s dynamics, but Theorem 3.3's equivalence claim is not proven as written.","tokens_in":29168,"tokens_out":4795,"would_cite":false,"duration_ms":62270,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A50","90C39"],"pacs":[],"model":"deepseek-v4-flash","headline":"A linear-quadratic dynamic game can be solved for its feedback Nash equilibria using only offline input/output data, with the same equilibria as the known-dynamics game and a convergent rolling-horizon approximation in the…","keywords":["linear-quadratic discrete-time games","feedback Nash equilibrium","unknown dynamics games","data-driven methods","behavioral system theory","Hankel matrix","offline data","finite-horizon strategy"],"falsifier":"Convertible check: take the two-player scalar example of Section 5, compute $K_i^*(T)$ by Algorithm 3.1 for increasing $T$, and compare the achieved infinite-horizon cost to the known-dynamics feedback-Nash-equilibrium cost; if the cost difference fails to vanish at the claimed quadratic rate, or if two offline datasets satisfying Assumption 3.1 produce different equilibrium strategies, the equivalence claim fails. More directly, finding any known-dynamics linear-quadratic game whose unique current-state feedback Nash equilibrium differs from a history-dependent equilibrium over the same feasible trajectory set would refute Theorem 3.3.","tokens_in":28209,"feed_emoji":"🎯","tokens_out":7415,"duration_ms":74102,"temperature":0.7,"pith_summary":"Imagine $N$ players competing over time in a linear system whose dynamics and even current state are unknown to them. The paper shows that if the players hold a single offline record of input/output trajectories that satisfies a rank condition, they do not need to know the system at all: the finite-horizon game with unknown dynamics has exactly the same feedback Nash equilibria as the corresponding known-dynamics game, and the equilibria have an affine form whose coefficients are computable from offline data. It provides coupled equations that are necessary and sufficient for existence and uniqueness of equilibria, and, under an invertibility condition, an algorithm that computes one equilibrium by solving linear equations stage by stage. For infinite horizons, where offline data cannot fully determine the feedback Nash equilibrium, the paper proves that playing the \"watch $T$ steps ahead, move one step now\" strategy converges to the known-dynamics infinite-horizon equilibrium cost, with an explicit quadratic-in-feedback-error convergence rate. A sympathetic reader would care because this removes the need for system identification, state observation, or coordinated online experimentation before play begins.","feed_headline":"Unknown dynamics? Offline data still yields same Nash equilibria","feed_subtitle":"Finite-horizon equilibria match the known-dynamics game, and a rolling horizon converges to the infinite-horizon cost.","key_machinery":"The argument is carried by the behavioral-system representation of the unknown plant: instead of identifying the matrices $A,B,C,D$, the paper treats the set of possible input/output trajectories as a behavior and uses the Hankel matrix of offline trajectories to span that behavior (Lemma 2.3). Each stage's optimization is rewritten in terms of the data-only matrix $G_t = Y_{f_t} M_t^\\dagger$, which predicts the current output from the historical data vector $U_t$; the equilibrium conditions then become linear equations in the feedback matrices $K_t^i$, grouped into the system $\\tilde H_t(P_{t+1})K_t = \\tilde g_t(P_{t+1})$. The invertibility of $\\tilde H_t$ is the mechanism that turns a game equilibrium into a backward recursion of linear solves whose coefficients depend only on offline data and cost parameters.","core_discovery":"The central claim is Theorem 3.3: for the discrete-time linear-quadratic game with unknown input/output/state dynamics, given initial data of length at least the system lag and offline trajectories whose Hankel matrix has full row rank, the finite-horizon game has the same set of feedback Nash equilibria as the same game with known dynamics, including the possibility that both sets are empty. Moreover every equilibrium strategy is affine in the observed history, $u_t^{i*} = K_t^{i*} U_{t-1} + L_t^{i*}$, and the equilibrium value is quadratic in the initial data; the coefficients are fixed by the offline data and the cost parameters. The paper then derives data-only coupled equations (1-1)--(1-5) that are necessary and sufficient for existence and uniqueness of the feedback Nash equilibrium, and shows that when each stage's coefficient matrix $\\tilde H_t(P_{t+1})$ is invertible, the equilibrium is computed by backward linear solves. In the infinite-horizon case (Theorem 4.2), under controllability plus Assumption 1 of the companion paper, the rolling finite-horizon strategy converges in total cost to the known-dynamics feedback-Nash-equilibrium cost, with error $O(\\max_{i\\in\\mathcal{N}}\\|K_i^*(T_i)-K_i^*\\|^2)$.","pith_inferences":["An implicit consequence is that any dataset satisfying the rank condition carries enough information for equilibrium play, regardless of how the inputs were excited; this suggests testing whether weaker rank conditions on the data matrix would also suffice.","The convergence result implies that players need not agree on a common prediction horizon: heterogeneous $T_i$ still yield convergence for every player, which may extend to online receding-horizon play and to noisy data through set-membership estimates.","The affine-in-history form suggests that history-dependent strategies in unknown-dynamics games are not artifacts but can reproduce current-state feedback equilibria, raising the testable question of whether the same data-only technique extends to games with coupled control costs or partial observations."],"forward_implications":["Players with no knowledge of the dynamics or the state can implement feedback-Nash-equilibrium strategies immediately from offline data, with no system identification step.","Knowing the system dynamics confers no strategic advantage: if one player knows the dynamics and another does not, both can achieve the same equilibrium performance given sufficient offline data.","Under the invertibility condition, computing a feedback Nash equilibrium reduces to solving $T$ sets of linear equations backward, avoiding coupled high-order matrix equations.","In the infinite-horizon unknown-dynamics game, the rolling-horizon strategy makes each player's total cost converge to the known-dynamics infinite-horizon equilibrium cost as the minimum prediction horizon grows, with error at most quadratic in the feedback-matrix error.","The same data-driven equations give necessary and sufficient conditions for existence and uniqueness of the feedback Nash equilibrium: if no solution to (1-1)--(1-5) exists, no equilibrium exists, and uniqueness is characterized by uniqueness of the strategy maps."],"supporting_citations":[{"why":"Supplies the definition of feedback Nash equilibrium and the known-dynamics linear-quadratic game formulation whose equilibria are compared.","marker":"[3]"},{"why":"Supplies the data-enabled predictive form $G_t = Y_{f_t} M_t^\\dagger$ that rewrites the game using offline input/output data.","marker":"[4]"},{"why":"Companion paper; provides Assumption 1 and the finite-horizon-strategy convergence results for known-dynamics games that Theorem 4.2 extends.","marker":"[11]"},{"why":"Provides the identifiability lemma (Lemma 2.3 here) ensuring a full-rank Hankel matrix spans all possible future trajectories.","marker":"[21]"},{"why":"Provides the lemma (Lemma 2.2 here) that a finite input/output trajectory of length at least the lag uniquely determines the state.","marker":"[22]"},{"why":"Establishes the behavioral input/output/state representation used to equate the feasible sets of the known- and unknown-dynamics games.","marker":"[35]"}],"fun_headline_variants":["Offline data recovers known-dynamics Nash equilibria","Unknown dynamics? Same Nash equilibria from offline data","Data-only FNE for unknown LQ games","Rolling horizon converges in unknown-dynamics games"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The finite-horizon equivalence rests on treating equality of the feasible input-output trajectory sets as sufficient for equality of the equilibrium strategy sets, even though the known-dynamics game restricts strategies to current-state feedback while the unknown-dynamics game allows strategies on the full history; the infinite-horizon convergence additionally rests on controllability and the unstated Assumption 1 of the companion paper.","fun_headline_variants_meta":{"raw":{"variants":["Offline data recovers known-dynamics Nash equilibria","Unknown dynamics? Same Nash equilibria from offline data","Data-only FNE for unknown LQ games","Rolling horizon converges in unknown-dynamics games"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1661,"prompt_tokens":1052,"completion_tokens":609,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":668,"completion_tokens_details":{"reasoning_tokens":547}},"tokens_in":668,"tokens_out":609,"duration_ms":6399,"temperature":1.0,"reasoning_tokens":547,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:14:11.360497+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Convertible check: take the two-player scalar example of Section 5, compute $K_i^*(T)$ by Algorithm 3.1 for increasing $T$, and compare the achieved infinite-horizon cost to the known-dynamics feedback-Nash-equilibrium cost; if the cost difference fails to vanish at the claimed quadratic rate, or if two offline datasets satisfying Assumption 3.1 produce different equilibrium strategies, the equivalence claim fails. More directly, finding any known-dynamics linear-quadratic game whose unique current-state feedback Nash equilibrium differs from a history-dependent equilibrium over the same feasible trajectory set would refute Theorem 3.3.","supporting_citations":[{"cited_title":"Coulson, J","cited_arxiv_id":null,"evidence_quote":"Supplies the data-enabled predictive form $G_t = Y_{f_t} M_t^\\dagger$ that rewrites the game using offline input/output data."},{"cited_title":"On finite-horizon approximation of an infinite-horizon feedback Nash equilibrium in discrete-time LQ games","cited_arxiv_id":"2506.19565","evidence_quote":"Companion paper; provides Assumption 1 and the finite-horizon-strategy convergence results for known-dynamics games that Theorem 4.2 extends."},{"cited_title":"Markovsky and F","cited_arxiv_id":null,"evidence_quote":"Provides the identifiability lemma (Lemma 2.3 here) ensuring a full-rank Hankel matrix spans all possible future trajectories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the behavioral input/output/state representation used to equate the feasible sets of the known- and unknown-dynamics games."}],"review_version":1}