Single-sample single-timescale actor-critic provably finds the global optimum of linear quadratic regulation on continuous state-action space with O(ε^-2) sample complexity.
Adaptive linear quadratic control using policy iteration
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator
Single-sample single-timescale actor-critic provably finds the global optimum of linear quadratic regulation on continuous state-action space with O(ε^-2) sample complexity.