REVIEW 4 cited by
On the Sample Complexity of the Linear Quadratic Regulator
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper addresses the optimal control problem known as the Linear Quadratic Regulator in the case when the dynamics are unknown. We propose a multi-stage procedure, called Coarse-ID control, that estimates a model from a few experimental trials, estimates the error in that model with respect to the truth, and then designs a controller using both the model and uncertainty estimate. Our technique uses contemporary tools from random matrix theory to bound the error in the estimation procedure. We also employ a recently developed approach to control synthesis called System Level Synthesis that enables robust control design by solving a convex optimization problem. We provide end-to-end bounds on the relative error in control cost that are nearly optimal in the number of parameters and that highlight salient properties of the system to be controlled such as closed-loop sensitivity and optimal control magnitude. We show experimentally that the Coarse-ID approach enables efficient computation of a stabilizing controller in regimes where simple control schemes that do not take the model uncertainty into account fail to stabilize the true system.
Forward citations
Cited by 4 Pith papers
-
Is Conditional Generative Modeling all you need for Decision-Making?
Return-conditional diffusion models for policies outperform offline RL on benchmarks by circumventing dynamic programming and enable constraint or skill composition.
-
Non-asymptotic Closed-Loop System Identification using Autoregressive Processes and Hankel Model Reduction
For closed-loop data, the REDAR algorithm (VARX fit plus balanced reduction) has one-step-ahead prediction error bounded by the optimal error plus terms that decay with model order p and with sample size T as O(1/√T).
-
Verification of Neural Network Control Policy Under Persistent Adversarial Perturbation
A sufficient-condition algorithm certifies boundedness of a closed-loop neural-network control system under l-infinity-bounded persistent adversarial perturbation, without requiring Lipschitz continuity of the policy.
-
Linear Dynamics: Clustering without identification
The eigenvalues of an unknown linear dynamical system's state-transition matrix can be consistently estimated from output time series by fitting the autoregressive parameters of an ARMA model, at a root-T convergence rate.
Discussion (0). Continue with ORCID to comment.