Projected mini-batch policy gradient attains Õ(1/η) sample complexity for known C^s noise densities and Õ(η^{-(2s+1)/(2s)}) when the density must be estimated, by pairing observations to cancel the cusp-obstruction divergence.
LQR through the lens of first order methods: Discrete-time case
5 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 5representative citing papers
Establishes global linear convergence of entropy-regularized policy gradient in continuous MDPs with log-linear softmax policies under Q-realizability by bounding non-uniform PL constants in two feature regimes.
A scalar-projection federated zeroth-order method for model-free LQR policy learning that reduces per-agent communication from O(d) to O(1) with convergence rate improving in the number of agents.
Riemannian regularization reshapes the policy optimization landscape to enable learning of Kalman gains from data under unknown and rank-deficient covariances with non-asymptotic convergence guarantees.
Relearn LQR combines recursive least squares with policy gradient for on-policy data-driven LQR and proves stability of the full scheme via Lyapunov analysis with averaging and timescale separation.
citing papers explorer
-
Sample Complexity of Policy Gradient for Log-Growth Control
Projected mini-batch policy gradient attains Õ(1/η) sample complexity for known C^s noise densities and Õ(η^{-(2s+1)/(2s)}) when the density must be estimated, by pairing observations to cancel the cusp-obstruction divergence.
-
Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs
Establishes global linear convergence of entropy-regularized policy gradient in continuous MDPs with log-linear softmax policies under Q-realizability by bounding non-uniform PL constants in two feature regimes.
-
Scalar Federated Learning for Linear Quadratic Regulator
A scalar-projection federated zeroth-order method for model-free LQR policy learning that reduces per-agent communication from O(d) to O(1) with convergence rate improving in the number of agents.
-
Learning Kalman Policy for Singular Unknown Covariances via Riemannian Regularization
Riemannian regularization reshapes the policy optimization landscape to enable learning of Kalman gains from data under unknown and rank-deficient covariances with non-asymptotic convergence guarantees.
-
Stability-Certified On-Policy Data-Driven LQR via Recursive Learning and Policy Gradient
Relearn LQR combines recursive least squares with policy gradient for on-policy data-driven LQR and proves stability of the full scheme via Lyapunov analysis with averaging and timescale separation.