Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning

Joshua Achiam; Shankar Sastry

arxiv: 1703.01732 · v1 · pith:673E4YWYnew · submitted 2017-03-06 · 💻 cs.LG

Surprise-Based Intrinsic Motivation for Deep Reinforcement Learning

Joshua Achiam , Shankar Sastry This is my paper

classification 💻 cs.LG

keywords learningexplorationintrinsicmotivationreinforcementrewardstaskscomplex

0 comments

read the original abstract

Exploration in complex domains is a key challenge in reinforcement learning, especially for tasks with very sparse rewards. Recent successes in deep reinforcement learning have been achieved mostly using simple heuristic exploration strategies such as $\epsilon$-greedy action selection or Gaussian control noise, but there are many tasks where these methods are insufficient to make any learning progress. Here, we consider more complex heuristics: efficient and scalable exploration strategies that maximize a notion of an agent's surprise about its experiences via intrinsic motivation. We propose to learn a model of the MDP transition probabilities concurrently with the policy, and to form intrinsic rewards that approximate the KL-divergence of the true transition probabilities from the learned model. One of our approximations results in using surprisal as intrinsic motivation, while the other gives the $k$-step learning progress. We show that our incentives enable agents to succeed in a wide range of environments with high-dimensional state spaces and very sparse rewards, including continuous control tasks and games in the Atari RAM domain, outperforming several other heuristic exploration techniques.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

LLM-Guided Task- and Affordance-Level Exploration in Reinforcement Learning
cs.RO 2025-09 unverdicted novelty 5.0

LLM-TALE steers RL exploration using LLM-generated plans at task and affordance levels with online suboptimality correction, improving sample efficiency and success rates on pick-and-place tasks without human supervision.