Pith. sign in

REVIEW 7 cited by

Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.03706 v2 pith:VZXMOOWE submitted 2025-05-06 math.OC cs.SYeess.SY

Policy Gradient Adaptive Control for the LQR: Indirect and Direct Approaches

classification math.OC cs.SYeess.SY
keywords gradientpgacpolicydirectindirectcontroladaptivemethod
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Motivated by recent advances of reinforcement learning and direct data-driven control, we propose policy gradient adaptive control (PGAC) for the linear quadratic regulator (LQR), which uses online closed-loop data to improve the control policy while maintaining stability. Our method adaptively updates the policy in feedback by descending the gradient of the LQR cost and is categorized as indirect, when gradients are computed via an estimated model, versus direct, when gradients are derived from data using sample covariance parameterization. Beyond the vanilla gradient, we also showcase the merits of the natural gradient and Gauss-Newton methods for the policy update. Notably, natural gradient descent bridges the indirect and direct PGAC, and the Gauss-Newton method of the indirect PGAC leads to an adaptive version of the celebrated Hewer's algorithm. To account for the uncertainty from noise, we propose a regularization method for both indirect and direct PGAC. For all the considered PGAC approaches, we show closed-loop stability and convergence of the policy to the optimal LQR gain. Simulations validate our theoretical findings and demonstrate the robustness and computational efficiency of PGAC.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Data-Enabled Primal-Dual Approach for Policy Learning with SDP Formulations

    eess.SY 2026-07 unverdicted novelty 7.0

    A primal-dual online framework updates policies from closed-loop data for SDP-based control synthesis in linear discrete-time systems, with local linear tracking and global ergodic convergence guarantees under persist...

  2. Direct Data-Driven Linear Quadratic Tracking via Policy Optimization

    eess.SY 2026-05 unverdicted novelty 7.0

    A reference-decoupled reformulation makes direct data-driven LQT equivalent to certainty-equivalence solutions and supports convergent offline and online DeePO algorithms.

  3. Adaptive Linear Quadratic Control of Unknown Linear Time-Varying Systems via Policy Gradient Methods

    math.OC 2026-07 accept novelty 6.0

    One-step policy-gradient LQR updates with normalized sliding-window least-squares stabilize unknown slowly varying and piecewise-constant linear systems and track frozen-time optima on average.

  4. Global Convergence of Policy Gradient Methods for ReLU Controllers in Linear Quadratic Regulation

    math.OC 2026-04 unverdicted novelty 6.0

    Model-based policy gradient converges globally to the optimal scalar LQR gain for discounted LQR using overparameterized ReLU networks by reducing the controller to two effective gains on positive and negative half-lines.

  5. Sample-Efficient Model-Free Policy Gradient Methods for Stochastic LQR via Robust Linear Regression

    eess.SY 2025-12 unverdicted novelty 6.0

    Primal-dual robust linear regression enables O(1/epsilon) sample complexity for model-free policy gradient methods on stochastic LQR.

  6. Sample-Efficient Model-Free Policy Gradient Methods for Stochastic LQR via Robust Linear Regression

    eess.SY 2025-12 reject novelty 5.0

    Claims O(1/epsilon) sample complexity for model-free NPG/GNM on stochastic LQR via primal-dual estimation, but the key estimation-error proof equates a max of a sum with a sum of maxes.

  7. Stability of Certainty-Equivalent Adaptive LQR for Linear Systems with Unknown Time-Varying Parameters

    eess.SY 2025-11 unverdicted novelty 5.0

    LMS estimation paired with certainty-equivalent LQR delivers finite-gain ℓ²-stability for linear systems with unknown time-varying parameters and disturbances.