Deepening the m-step Bellman lookahead in PGTS monotonically reduces the set of stationary policies, so the worst local optimum improves with depth.
Beyond the One Step Greedy Approach in Reinforcement Learning
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The famous Policy Iteration algorithm alternates between policy improvement and policy evaluation. Implementations of this algorithm with several variants of the latter evaluation stage, e.g, $n$-step and trace-based returns, have been analyzed in previous works. However, the case of multiple-step lookahead policy improvement, despite the recent increase in empirical evidence of its strength, has to our knowledge not been carefully analyzed yet. In this work, we introduce the first such analysis. Namely, we formulate variants of multiple-step policy improvement, derive new algorithms using these definitions and prove their convergence. Moreover, we show that recent prominent Reinforcement Learning algorithms are, in fact, instances of our framework. We thus shed light on their empirical success and give a recipe for deriving new algorithms for future study.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Policy Gradient with Tree Search: Avoiding Local Optimas through Lookahead
Deepening the m-step Bellman lookahead in PGTS monotonically reduces the set of stationary policies, so the worst local optimum improves with depth.