Bounded perturbations of size O(sqrt(epsilon)) still allow derivative-free LQR policy optimization to reach an epsilon-optimal policy with high probability, with sample complexity O(1/epsilon^2 log(1/epsilon)) (one-point) or O(1/epsilon log(1/epsilon)) (two-point).
Optimal and autonomous control using reinforcement learning: A survey,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
math.OC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
On the Robustness of Derivative-free Methods for Linear Quadratic Regulator
Bounded perturbations of size O(sqrt(epsilon)) still allow derivative-free LQR policy optimization to reach an epsilon-optimal policy with high probability, with sample complexity O(1/epsilon^2 log(1/epsilon)) (one-point) or O(1/epsilon log(1/epsilon)) (two-point).