REVIEW 2 cited by
Learning Linear-Quadratic Regulators Efficiently with only $\sqrt{T}$ Regret
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We present the first computationally-efficient algorithm with $\widetilde O(\sqrt{T})$ regret for learning in Linear Quadratic Control systems with unknown dynamics. By that, we resolve an open question of Abbasi-Yadkori and Szepesv\'ari (2011) and Dean, Mania, Matni, Recht, and Tu (2018).
Forward citations
Cited by 2 Pith papers
-
Regret-Guaranteed Safe Switching: LQR Setting with Unknown Dynamics
Proposes a regret-minimizing algorithm for safe mode switching in unknown-dynamics LQR that achieves expected regret O(|M|^{1/4} n_s^{3/4} + n_m) under infinite mode visits.
-
Multi-agent imitation learning with function approximation: Linear Markov games and beyond
In linear Markov games, behavior cloning's sample complexity hinges on a feature-level concentrability coefficient, and the interactive algorithm LSVI-UCB-ZERO-BC removes concentrability dependence entirely, scaling o...
Discussion (0). Continue with ORCID to comment.