REVIEW 7 cited by
Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches
read the original abstract
Motivated by recent advances of reinforcement learning and direct data-driven control, we propose policy gradient adaptive control (PGAC) for the linear quadratic regulator (LQR), which uses online closed-loop data to improve the control policy while maintaining stability. Our method adaptively updates the policy in feedback by descending the gradient of the LQR cost and is categorized as indirect, when gradients are computed via an estimated model, versus direct, when gradients are derived from data using sample covariance parameterization. Beyond the vanilla gradient, we also showcase the merits of the natural gradient and Gauss-Newton methods for the policy update. Notably, natural gradient descent bridges the indirect and direct PGAC, and the Gauss-Newton method of the indirect PGAC leads to an adaptive version of the celebrated Hewer's algorithm. To account for the uncertainty from noise, we propose a regularization method for both indirect and direct PGAC. For all the considered PGAC approaches, we show closed-loop stability and convergence of the policy to the optimal LQR gain. Simulations validate our theoretical findings and demonstrate the robustness and computational efficiency of PGAC.
Forward citations
Cited by 7 Pith papers
-
A Data-Enabled Primal-Dual Approach for Policy Learning with SDP Formulations
A primal-dual online framework updates policies from closed-loop data for SDP-based control synthesis in linear discrete-time systems, with local linear tracking and global ergodic convergence guarantees under persist...
-
Direct Data-Driven Linear Quadratic Tracking via Policy Optimization
A reference-decoupled reformulation makes direct data-driven LQT equivalent to certainty-equivalence solutions and supports convergent offline and online DeePO algorithms.
-
Adaptive Linear Quadratic Control of Unknown Linear Time-Varying Systems via Policy Gradient Methods
One-step policy-gradient LQR updates with normalized sliding-window least-squares stabilize unknown slowly varying and piecewise-constant linear systems and track frozen-time optima on average.
-
Global Convergence of Policy Gradient Methods for ReLU Controllers in Linear Quadratic Regulation
Model-based policy gradient converges globally to the optimal scalar LQR gain for discounted LQR using overparameterized ReLU networks by reducing the controller to two effective gains on positive and negative half-lines.
-
Sample-Efficient Model-Free Policy Gradient Methods for Stochastic LQR via Robust Linear Regression
Primal-dual robust linear regression enables O(1/epsilon) sample complexity for model-free policy gradient methods on stochastic LQR.
-
Sample-Efficient Model-Free Policy Gradient Methods for Stochastic LQR via Robust Linear Regression
Claims O(1/epsilon) sample complexity for model-free NPG/GNM on stochastic LQR via primal-dual estimation, but the key estimation-error proof equates a max of a sum with a sum of maxes.
-
Stability of Certainty-Equivalent Adaptive LQR for Linear Systems with Unknown Time-Varying Parameters
LMS estimation paired with certainty-equivalent LQR delivers finite-gain ℓ²-stability for linear systems with unknown time-varying parameters and disturbances.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.