Pith. sign in

REVIEW 1 cited by

Augmenting GAIL with BC for sample efficient imitation learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2001.07798 v4 pith:4H5OFESR submitted 2020-01-21 cs.LG stat.ML

classification cs.LGstat.ML
keywords efficientgaillearningsampleexpertimitationmethodspolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning is the problem of recovering an expert policy without access to a reward signal. Behavior cloning and GAIL are two widely used methods for performing imitation learning. Behavior cloning converges in a few iterations but doesn't achieve peak performance due to its inherent iid assumption about the state-action distribution. GAIL addresses the issue by accounting for the temporal dependencies when performing a state distribution matching between the agent and the expert. Although GAIL is sample efficient in the number of expert trajectories required, it is still not very sample efficient in terms of the environment interactions needed for convergence of the policy. Given the complementary benefits of both methods, we present a simple and elegant method to combine both methods to enable stable and sample efficient learning. Our algorithm is very simple to implement and integrates with different policy gradient algorithms. We demonstrate the effectiveness of the algorithm in low dimensional control tasks, gridworlds and in high dimensional image-based tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Creating Hierarchical Dispositions of Needs in an Agent

    cs.LG 2024-11 reject novelty 3.0 of 10

    A hand-designed nested reward formula with a secondary multi-output network is claimed to improve PPO on Pendulum-v1, but the method is underspecified and the evidence is anecdotal.

Pith tools