Pith. sign in

REVIEW 2 cited by

A Study of Global and Episodic Bonuses for Exploration in Contextual MDPs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03236 v1 pith:ABEQFEI2 submitted 2023-06-05 cs.AI

classification cs.AI
keywords bonusesacrossepisodicglobalsharedstructuredifferenttypes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Exploration in environments which differ across episodes has received increasing attention in recent years. Current methods use some combination of global novelty bonuses, computed using the agent's entire training experience, and \textit{episodic novelty bonuses}, computed using only experience from the current episode. However, the use of these two types of bonuses has been ad-hoc and poorly understood. In this work, we shed light on the behavior of these two types of bonuses through controlled experiments on easily interpretable tasks as well as challenging pixel-based settings. We find that the two types of bonuses succeed in different settings, with episodic bonuses being most effective when there is little shared structure across episodes and global bonuses being effective when more structure is shared. We develop a conceptual framework which makes this notion of shared structure precise by considering the variance of the value function across contexts, and which provides a unifying explanation of our empirical results. We furthermore find that combining the two bonuses can lead to more robust performance across different degrees of shared structure, and investigate different algorithmic choices for defining and combining global and episodic bonuses based on function approximation. This results in an algorithm which sets a new state of the art across 16 tasks from the MiniHack suite used in prior work, and also performs robustly on Habitat and Montezuma's Revenge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The impact of intrinsic rewards on exploration in Reinforcement Learning

    cs.AI 2025-01 conditional novelty 5.0 of 10

    An empirical MiniGrid study shows state-counting is best for low-dimensional observations, maximum entropy is more robust with images, and DIAYN skill learning does not aid exploration.

  2. Deep Reinforcement Learning with Hybrid Intrinsic Reward Model

    cs.LG 2025-01 conditional novelty 4.0 of 10

    A systematic empirical study showing that cycle-based fusion of two or three intrinsic rewards improves exploration and unsupervised RL performance more than single rewards.

Pith tools