Bounded perturbations of size O(sqrt(epsilon)) still allow derivative-free LQR policy optimization to reach an epsilon-optimal policy with high probability, with sample complexity O(1/epsilon^2 log(1/epsilon)) (one-point) or O(1/epsilon log(1/epsilon)) (two-point).
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
On the Robustness of Derivative-free Methods for Linear Quadratic Regulator
Bounded perturbations of size O(sqrt(epsilon)) still allow derivative-free LQR policy optimization to reach an epsilon-optimal policy with high probability, with sample complexity O(1/epsilon^2 log(1/epsilon)) (one-point) or O(1/epsilon log(1/epsilon)) (two-point).