REVIEW 2 major objections 4 minor 42 references
One gradient step per time update stabilizes unknown linear time-varying systems while tracking frozen-time LQR policies.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
One-step policy-gradient LQR updates with normalized sliding-window least-squares stabilize unknown slowly varying and piecewise-constant linear systems and track frozen-time optima on average.
T0 review reviewed 2026-07-12 challenge →
load-bearing objection Solid Automatica-style LTV extension of PGAC: one-step gradient updates plus normalized window LS, with real PES and average optimality-gap theorems for slow drift and switched modes. the 2 major comments →
Adaptive Linear Quadratic Control of Unknown Linear Time-Varying Systems via Policy Gradient Methods
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Under small enough variation (or long enough dwell), small enough normalized identification error, and a sufficiently small gradient step-size, one-step policy-gradient adaptive control produces a sequentially stable gain sequence for slowly varying unknown LTV systems and therefore practical exponential stability without dwell time; for piecewise-constant LTV systems the same controller yields within-mode decay plus interval-wise practical exponential stability under a dwell-time condition, together with explicit average frozen-time LQR optimality-gap bounds in both cases.
What carries the argument
Policy-gradient adaptive control (PGAC): at each instant form a normalized sliding-window least-squares model of the recent closed-loop trajectory, evaluate the LQR policy gradient of that frozen surrogate, and take exactly one descent step; sequential stability of the resulting gain sequence then supplies the closed-loop certificates.
Load-bearing premise
The closed-loop data must stay persistently exciting under the adapting gain itself; if the normalized excitation level collapses, the identification-error bounds fail and the stability thresholds become empty.
What would settle it
Run the algorithm on a slowly varying LTV plant whose continuous drift exceeds the paper's (non-constructive) variation threshold while keeping excitation, noise and step-size fixed; if the state still decays practically exponentially, the slow-variation claim is false. Alternatively, drop the probing signal so that the normalized Gramian falls below the assumed floor and check whether the identification error and subsequent stability guarantees remain valid.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes policy gradient adaptive control (PGAC) for unknown LTV systems: at each time the state-feedback gain is updated by one gradient step of a frozen-time LQR cost built from a normalized sliding-window least-squares model estimate. For slowly varying systems (Assumptions 2–4, 11), Theorem 16 shows sequential stability of the policy sequence and practical exponential stability of the closed loop without a dwell-time condition, plus an average frozen-time optimality-gap bound. For piecewise-constant systems (Assumptions 2, 3, 7, 11), Theorem 18 gives within-mode decay and interval-wise PES under a dwell-time condition, with a corresponding optimality-gap bound that depends on switching frequency. Identification-error lemmas (14, 17), cost boundedness by induction, sequential-stability arguments, and telescoping gap estimates are developed in the appendix; three numerical examples (continuous LTV, switched LTV, nonlinear planar quadrotor) illustrate the method.
Significance. The contribution is a useful middle ground between one-shot certainty-equivalence LQR/SDP updates and purely direct data-driven schemes: the first-order update is computationally light and the stepsize directly limits policy chatter under noisy estimates. Extending sequential-stability certificates to continuous slow variation without dwell time, and obtaining uniform constants for infinitely many switches in the piecewise-constant case, are genuine advances over the authors’ preliminary switched result and over several recent LTV data-driven methods. The normalized sliding-window estimator cleanly separates variation, noise, and excitation in the error bounds. The analysis is fully spelled out in the appendix and the numerical examples (including a nonlinear plant) support the claims. The PE and smallness conditions are standard for this literature and are disclosed.
major comments (2)
- Assumption 11 requires a uniform lower bound γ on the normalized data Gramian for every window while the gain is adapting. Lemmas 14 and 17 and all subsequent thresholds in Theorems 16 and 18 collapse if closed-loop excitation fails. The manuscript should either (i) give a concrete, checkable construction of e_t that preserves γ under the sequential-stability bounds already proved, or (ii) state explicitly that PE is an external hypothesis and discuss how γ can be monitored online. Without one of these, the certificates remain conditional on an assumption that is not automatically inherited from the closed-loop dynamics.
- The admissible ranges for δ, ε (or ε_sw) and η in Theorems 16 and 18 are expressed via existential constants ν_i that depend on (a_m, b_m, Q, R, K_t0) but are never made explicit or estimated. While this is common in sequential-stability arguments, a short remark (or a numerical illustration of how large η may be taken for the examples of Section 5) would make the results more usable and would clarify that the conditions are not vacuous for the simulated systems.
minor comments (4)
- In Algorithm 1 the data window used for identification is written (X_{t+1}, U_{t+1}, X_{t+2}) while the surrounding text uses the indexing of (12); a one-line clarification of the time shift would avoid confusion.
- Figures 2–5 would benefit from a short caption note on the magnitude of the residual set relative to e_m and w_m, so that the practical-stability claim is visually linked to the theory.
- A few typos remain (e.g., “D¨ orfler”, “Z¨ urich” in the author block; “the pol-icy update” line break in Section 4.2). A light copy-edit pass is enough.
- The discussion of DeePO and discounted LQR in Section 4.4 is interesting but could be shortened; the main theorems already stand without it.
Circularity Check
No significant circularity: LTV stability and optimality-gap certificates are derived from PE, small-variation/dwell-time, and stepsize assumptions via standard Lyapunov/gradient-dominance arguments, not by construction from fitted quantities or load-bearing self-citation.
full rationale
The derivation chain is self-contained against its stated assumptions. Identification errors (Lemmas 14, 17) follow from the normalized data equation plus uniform PE (Assumption 11) and variation bounds; they are not defined in terms of the stability claim. Policy updates use certainty-equivalence PG of the frozen-time LQR cost of the windowed estimate; sequential stability (Definition 15, from the literature) is then proved under small δ, ε (or ε_sw), and η (Lemmas 25–27, Theorems 16 and 18), and PES / interval-wise PES and average gaps C_t(K_t)−C_t^* follow as consequences. Gaps are measured against true frozen-time LQR costs of (A_t,B_t), not against a quantity defined by the fitted gain. Self-citations to the authors’ LTI PGAC work supply standard gradient formulas and perturbation lemmas; they do not force the LTV certificates by uniqueness or ansatz. Remark 12 explicitly breaks the ordinary-LS state-bound circularity by using normalization. No fitted-input-as-prediction, self-definitional, or renaming pattern appears in the load-bearing theorems.
Axiom & Free-Parameter Ledger
free parameters (3)
- policy gradient stepsize η
- sliding window length L
- probing signal bound e_m and excitation level γ
axioms (7)
- domain assumption Frozen-time controllability of (A_t, B_t) for every t (Assumption 2)
- domain assumption Uniform bounds on ||A_t||, ||B_t|| and process noise ||w_t|| (Assumption 3)
- domain assumption Slow variation ||Δ_t|| ≤ δ (Assumption 4) or bounded switching jumps (Assumption 7)
- domain assumption Bounded normalized PE of closed-loop data for all t (Assumption 11)
- domain assumption Initial policy K_t0 stabilizes the initial frozen-time pair
- standard math Policy gradient formula and gradient dominance for LQR (Lemma 1 / Fazel et al.)
- standard math Sequential stability implies practical exponential stability / ISS-type bounds (Cohen et al.)
Cite this review
Pith. "Pith review of Adaptive Linear Quadratic Control of Unknown Linear Time-Varying Systems via Policy Gradient Methods." pith.science (2026). https://pith.science/paper/HDB4UVAS
@misc{pith2026260703251,
author = {Pith},
title = {Pith review of: Adaptive Linear Quadratic Control of Unknown Linear Time-Varying Systems via Policy Gradient Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/HDB4UVAS}},
note = {Machine review of arXiv:2607.03251}
}
read the original abstract
Unknown linear time-varying (LTV) systems require the control policy to adapt from online closed-loop data as dynamics evolve. Existing methods usually update the policy by solving a one-shot optimization problem, which can be computationally demanding and sensitive to noisy model estimates. In this paper, we propose a policy gradient adaptive control (PGAC) method for LTV system control with unknown model parameters. Specifically, PGAC integrates online policy optimization into feedback by updating the state-feedback policy with one-step gradient descent of the linear quadratic regulator cost at each time instant. This incremental update is computationally light and naturally limits policy variations caused by noisy data. To explicitly compute the policy gradient online, we estimate local models from recent closed-loop trajectories using normalized sliding-window least-squares. We provide stability and convergence certificates of PGAC for two classes of LTV systems. For slowly time-varying systems, we prove that the closed-loop system achieves practical exponential stability without a dwell-time condition. For piecewise-constant LTV systems, we establish practical stability through a dwell-time contraction argument. We also provide average frozen-time optimality-gap bounds of the policy sequence for both classes. Finally, we validate the effectiveness of our method via numerical case studies of both LTV and nonlinear systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Derivative-free data-enabled optimal tracking (DF-DeeOT) controller for a PV grid-connected inverter.IEEE Access, 14:16406–16420, 2026
Said Al-Abri, Myada Shadoul, and Hassan Yousef. Derivative-free data-enabled optimal tracking (DF-DeeOT) controller for a PV grid-connected inverter.IEEE Access, 14:16406–16420, 2026
2026
-
[2]
Courier Corporation, 2007
Brian DO Anderson and John B Moore.Optimal control: linear quadratic methods. Courier Corporation, 2007
2007
-
[3]
Marcell Bartos, Johannes K¨ ohler, Florian D¨ orfler, and Melanie N Zeilinger. Stability of certainty-equivalent adaptive LQR for linear systems with unknown time-varying parameters.arXiv preprint arXiv:2511.08236, 2025
Pith/arXiv arXiv 2025
-
[4]
Data-driven model predictive control with stability and robustness guarantees.IEEE Transactions on Automatic Control, 66(4):1702–1717, 2020
Julian Berberich, Johannes K¨ ohler, Matthias A M¨ uller, and Frank Allg¨ ower. Data-driven model predictive control with stability and robustness guarantees.IEEE Transactions on Automatic Control, 66(4):1702–1717, 2020
2020
-
[5]
Data-driven predictive control in a stochastic setting: a unified framework.Automatica, 152:110961, 2023
Valentina Breschi, Alessandro Chiuso, and Simone Formentin. Data-driven predictive control in a stochastic setting: a unified framework.Automatica, 152:110961, 2023
2023
-
[6]
Online linear quadratic control
Alon Cohen, Avinatan Hasidim, Tomer Koren, Nevena Lazic, Yishay Mansour, and Kunal Talwar. Online linear quadratic control. InInternational Conference on Machine Learning, pages 1029–1038. PMLR, 2018
2018
-
[7]
Learning linear-quadratic regulators efficiently with only √ Tregret
Alon Cohen, Tomer Koren, and Yishay Mansour. Learning linear-quadratic regulators efficiently with only √ Tregret. In International Conference on Machine Learning, pages 1300–
-
[8]
Data- enabled predictive control: In the shallows of the DeePC
Jeremy Coulson, John Lygeros, and Florian D¨ orfler. Data- enabled predictive control: In the shallows of the DeePC. In 18th European Control Conference (ECC), pages 307–312, 2019
2019
-
[9]
Online adaptation of data-driven controllers for unknown nonlinear systems.International Journal of Robust and Nonlinear Control, 36(3):1086–1095, 2026
Xiaoyan Dai, Claudio De Persis, and Nima Monshizadeh. Online adaptation of data-driven controllers for unknown nonlinear systems.International Journal of Robust and Nonlinear Control, 36(3):1086–1095, 2026
2026
-
[10]
Formulas for data-driven control: Stabilization, optimality, and robustness.IEEE Transactions on Automatic Control, 65(3):909–924, 2019
Claudio De Persis and Pietro Tesi. Formulas for data-driven control: Stabilization, optimality, and robustness.IEEE Transactions on Automatic Control, 65(3):909–924, 2019
2019
-
[11]
On the sample complexity of the linear quadratic regulator.Foundations of Computational Mathematics, 20(4):633–679, 2020
Sarah Dean, Horia Mania, Nikolai Matni, Benjamin Recht, and Stephen Tu. On the sample complexity of the linear quadratic regulator.Foundations of Computational Mathematics, 20(4):633–679, 2020
2020
-
[12]
Data-driven control: Part two of two: Hot take: Why not go with models?IEEE Control Systems Magazine, 43(6):27–31, 2023
Florian D¨ orfler. Data-driven control: Part two of two: Hot take: Why not go with models?IEEE Control Systems Magazine, 43(6):27–31, 2023
2023
-
[13]
On the certainty-equivalence approach to direct data-driven LQR design.IEEE Transactions on Automatic Control, 68(12):7989–7996, 2023
Florian D¨ orfler, Pietro Tesi, and Claudio De Persis. On the certainty-equivalence approach to direct data-driven LQR design.IEEE Transactions on Automatic Control, 68(12):7989–7996, 2023
2023
-
[14]
Predictive active steering control for autonomous vehicle systems.IEEE Transactions on Control Systems Technology, 15(3):566–580, 2007
Paolo Falcone, Francesco Borrelli, Jahan Asgari, Hongtei Eric Tseng, and Davor Hrovat. Predictive active steering control for autonomous vehicle systems.IEEE Transactions on Control Systems Technology, 15(3):566–580, 2007
2007
-
[15]
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi. Global convergence of policy gradient methods for the linear quadratic regulator. InInternational Conference on Machine Learning, pages 1467–1476, 2018
2018
-
[16]
Toward a theoretical foundation of policy optimization for learning control policies.Annual Review of Control, Robotics, and Autonomous Systems, 6:123–158, 2023
Bin Hu, Kaiqing Zhang, Na Li, Mehran Mesbahi, Maryam Fazel, and Tamer Ba¸ sar. Toward a theoretical foundation of policy optimization for learning control policies.Annual Review of Control, Robotics, and Autonomous Systems, 6:123–158, 2023
2023
-
[17]
A hybrid systems framework for data-based adaptive control of linear time- varying systems.IEEE Transactions on Automatic Control, 2025
Andrea Iannelli and Romain Postoyan. A hybrid systems framework for data-based adaptive control of linear time- varying systems.IEEE Transactions on Automatic Control, 2025
2025
-
[18]
Khalil.Nonlinear Systems
Hassan K. Khalil.Nonlinear Systems. Prentice Hall, 3 edition, 2002
2002
-
[19]
Lakshmikantham, S
V. Lakshmikantham, S. Leela, and A. A. Martynyuk. Practical Stability of Nonlinear Systems. World Scientific, 1990
1990
-
[20]
Adaptive control of unknown linear switched systems via policy gradient methods.European Control Conference (ECC), 2026, accepted
Felix Laurent, Feiran Zhao, Jaap Eising, and Florian D¨ orfler. Adaptive control of unknown linear switched systems via policy gradient methods.European Control Conference (ECC), 2026, accepted
2026
-
[21]
Online data- driven adaptive control for unknown linear time-varying systems
Shenyu Liu, Kaiwen Chen, and Jaap Eising. Online data- driven adaptive control for unknown linear time-varying systems. In2023 62nd IEEE Conference on Decision and Control (CDC), pages 8775–8780. IEEE, 2023
2023
-
[22]
Almost surely √ Tregret for adaptive lqr.IEEE Transactions on Automatic Control, 70(8):5145– 5159, 2025
Yiwen Lu and Yilin Mo. Almost surely √ Tregret for adaptive lqr.IEEE Transactions on Automatic Control, 70(8):5145– 5159, 2025
2025
-
[23]
Certainty equivalence is efficient for linear quadratic control
Horia Mania, Stephen Tu, and Benjamin Recht. Certainty equivalence is efficient for linear quadratic control. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019
2019
-
[24]
Persistency of excitation criteria for linear, multivariable, time-varying systems.Mathematics of Control, Signals and Systems, 1(3):203–226, 1988
Iven MY Mareels and Michel Gevers. Persistency of excitation criteria for linear, multivariable, time-varying systems.Mathematics of Control, Signals and Systems, 1(3):203–226, 1988
1988
-
[25]
Online control of unknown time-varying dynamical systems.Advances in Neural Information Processing Systems, 34:15934–15945, 2021
Edgar Minasyan, Paula Gradu, Max Simchowitz, and Elad Hazan. Online control of unknown time-varying dynamical systems.Advances in Neural Information Processing Systems, 34:15934–15945, 2021
2021
-
[26]
Direct data- driven control of linear time-varying systems.IEEE Transactions on Automatic Control, 68(8):4888–4895, 2023
Benita Nortmann and Thulasi Mylvaganam. Direct data- driven control of linear time-varying systems.IEEE Transactions on Automatic Control, 68(8):4888–4895, 2023
2023
-
[27]
An adaptive data- enabled policy optimization approach for autonomous bicycle control.IEEE Transactions on Control Systems Technology, 2026 (early access)
Niklas Persson, Feiran Zhao, Mojtaba Kaheni, Florian D¨ orfler, and Alessandro V Papadopoulos. An adaptive data- enabled policy optimization approach for autonomous bicycle control.IEEE Transactions on Control Systems Technology, 2026 (early access)
2026
-
[28]
Stable online control of linear time-varying systems
Guannan Qu, Yuanyuan Shi, Sahin Lale, Anima Anandkumar, and Adam Wierman. Stable online control of linear time-varying systems. InLearning for Dynamics and Control, pages 742–753. PMLR, 2021
2021
-
[29]
Online learning of data-driven controllers for unknown switched linear systems.Automatica, 145:110519, 2022
Monica Rotulo, Claudio De Persis, and Pietro Tesi. Online learning of data-driven controllers for unknown switched linear systems.Automatica, 145:110519, 2022
2022
-
[30]
Rugh and Jeff S
Wilson J. Rugh and Jeff S. Shamma. Research on gain scheduling.Automatica, 36(10):1401–1425, 2000
2000
-
[31]
Gain and phase margin for multiloop LQG regulators.IEEE Transactions on Automatic Control, 22(2):173–179, 1977
Michael Safonov and Michael Athans. Gain and phase margin for multiloop LQG regulators.IEEE Transactions on Automatic Control, 22(2):173–179, 1977
1977
-
[32]
Naive exploration is optimal for online LQR
Max Simchowitz and Dylan Foster. Naive exploration is optimal for online LQR. InInternational Conference on Machine Learning, pages 8937–8948. PMLR, 2020
2020
-
[33]
Adaptive model predictive control for linear time varying mimo systems.Automatica, 105:237–245, 2019
Marko Tanaskovic, Lorenzo Fagiano, and Vojislav Gligorovski. Adaptive model predictive control for linear time varying mimo systems.Automatica, 105:237–245, 2019
2019
-
[34]
Teel and Laurent Praly
Andrew R. Teel and Laurent Praly. Tools for semiglobal stabilization by partial state and output feedback.SIAM Journal on Control and Optimization, 33(5):1443–1488, 1995
1995
-
[35]
Statistical learning theory for control: A finite-sample perspective.IEEE Control Systems Magazine, 43(6):67–97, 2023
Anastasios Tsiamis, Ingvar Ziemann, Nikolai Matni, and George J Pappas. Statistical learning theory for control: A finite-sample perspective.IEEE Control Systems Magazine, 43(6):67–97, 2023
2023
-
[36]
Data informativity: a new perspective on data-driven analysis and control.IEEE Transactions on Automatic Control, 65(11):4753–4768, 2020
Henk J Van Waarde, Jaap Eising, Harry L Trentelman, and M Kanat Camlibel. Data informativity: a new perspective on data-driven analysis and control.IEEE Transactions on Automatic Control, 65(11):4753–4768, 2020
2020
-
[37]
Xuerui Wang, Feiran Zhao, Andres J¨ urisson, Florian D¨ orfler, and Roy S. Smith. Unified aeroelastic flutter and loads control via data-enabled policy optimization. IEEE Transactions on Aerospace and Electronic Systems, 61(5):11437–11449, 2025
2025
-
[38]
Xinyi Yi and Ioannis Lestas. Data-driven online control for real-time optimal economic dispatch and temperature regulation in district heating systems.arXiv preprint arXiv:2603.23748, 2026
Pith/arXiv arXiv 2026
-
[39]
Feiran Zhao, Alessandro Chiuso, and Florian D¨ orfler. Policy gradient adaptive control for the LQR: Indirect and direct approaches.arXiv preprint arXiv:2505.03706, 2025
Pith/arXiv arXiv 2025
-
[40]
Data-enabled policy optimization for direct adaptive learning of the LQR.IEEE Transactions on Automatic Control, 70(11):7217–7232, 2025
Feiran Zhao, Florian D¨ orfler, Alessandro Chiuso, and Keyou You. Data-enabled policy optimization for direct adaptive learning of the LQR.IEEE Transactions on Automatic Control, 70(11):7217–7232, 2025
2025
-
[41]
Convergence and sample complexity of policy gradient methods for stabilizing linear systems.IEEE Transactions on Automatic Control, 70(3):1455–1466, 2025
Feiran Zhao, Xingyun Fu, and Keyou You. Convergence and sample complexity of policy gradient methods for stabilizing linear systems.IEEE Transactions on Automatic Control, 70(3):1455–1466, 2025. 17
2025
-
[42]
Direct adaptive control of grid-connected power converters via output-feedback data- enabled policy optimization
Feiran Zhao, Ruohan Leng, Linbin Huang, Huanhai Xin, Keyou You, and Florian D¨ orfler. Direct adaptive control of grid-connected power converters via output-feedback data- enabled policy optimization. In2025 European Control Conference (ECC), pages 2563–2568. IEEE, 2025. 18
2025
This paper was first reviewed by grok-4.5 on July 12, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.