Pith. sign in

REVIEW

A Deep Reinforcement Learning Approach to Marginalized Importance Sampling with the Successor Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.06854 v2 pith:HKP7DS2K submitted 2021-06-12 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords approachdeeplearningreinforcementrepresentationsamplingsuccessordensity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Marginalized importance sampling (MIS), which measures the density ratio between the state-action occupancy of a target policy and that of a sampling distribution, is a promising approach for off-policy evaluation. However, current state-of-the-art MIS methods rely on complex optimization tricks and succeed mostly on simple toy problems. We bridge the gap between MIS and deep reinforcement learning by observing that the density ratio can be computed from the successor representation of the target policy. The successor representation can be trained through deep reinforcement learning methodology and decouples the reward optimization from the dynamics of the environment, making the resulting algorithm stable and applicable to high-dimensional domains. We evaluate the empirical performance of our approach on a variety of challenging Atari and MuJoCo environments.

Discussion (0). Sign in to comment.

Pith tools