Pith. sign in

REVIEW

Backplay: "Man muss immer umkehren"

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.06919 v5 pith:CT7FONS2 submitted 2018-07-18 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords backplayimprovetrainingcurriculumdemonstrationefficiencyenvironmentsinitial
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-free reinforcement learning (RL) requires a large number of trials to learn a good policy, especially in environments with sparse rewards. We explore a method to improve the sample efficiency when we have access to demonstrations. Our approach, Backplay, uses a single demonstration to construct a curriculum for a given task. Rather than starting each training episode in the environment's fixed initial state, we start the agent near the end of the demonstration and move the starting point backwards during the course of training until we reach the initial state. Our contributions are that we analytically characterize the types of environments where Backplay can improve training speed, demonstrate the effectiveness of Backplay both in large grid worlds and a complex four player zero-sum game (Pommerman), and show that Backplay compares favorably to other competitive methods known to improve sample efficiency. This includes reward shaping, behavioral cloning, and reverse curriculum generation.

Discussion (0). Continue with ORCID to comment.

Pith tools