Learnings Options End-to-End for Continuous Action Tasks

Doina Precup; Jean Harb; Martin Klissarov; Pierre-Luc Bacon

arxiv: 1712.00004 · v1 · pith:FAGNKIXBnew · submitted 2017-11-30 · 💻 cs.LG · cs.AI

Learnings Options End-to-End for Continuous Action Tasks

Martin Klissarov , Pierre-Luc Bacon , Jean Harb , Doina Precup This is my paper

classification 💻 cs.LG cs.AI

keywords optionspolicyresultsaboutwhenaachieveactionactionsarchitecture

0 comments

read the original abstract

We present new results on learning temporally extended actions for continuoustasks, using the options framework (Suttonet al.[1999b], Precup [2000]). In orderto achieve this goal we work with the option-critic architecture (Baconet al.[2017])using a deliberation cost and train it with proximal policy optimization (Schulmanet al.[2017]) instead of vanilla policy gradient. Results on Mujoco domains arepromising, but lead to interesting questions aboutwhena given option should beused, an issue directly connected to the use of initiation sets.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Hierarchical Behaviour Spaces
cs.AI 2026-04 unverdicted novelty 6.0

Hierarchical Behaviour Spaces uses linear combinations of reward functions to induce expressive behavior spaces in hierarchical RL, yielding strong performance on NetHack primarily through better exploration rather th...
Scalable Option Learning in High-Throughput Environments
cs.LG 2025-08 unverdicted novelty 6.0

SOL is a new hierarchical RL algorithm that reaches 35x higher throughput and outperforms flat agents when trained on 30 billion frames in NetHack while showing positive scaling.