Adaptive TRPO is shown to be mirror descent with an adaptive proximity term, converging at tilde O(1/sqrt(N)) and at tilde O(1/N) for regularized MDPs.
WhereCω,1 = √ A,Cω,3 = 1 for the euclidean case, and Cω,1 = 1,Cω,3 = logA for the non-euclidean case
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPs
Adaptive TRPO is shown to be mirror descent with an adaptive proximity term, converging at tilde O(1/sqrt(N)) and at tilde O(1/N) for regularized MDPs.