Pith. sign in

REVIEW 2 cited by

Non-Adversarial Imitation Learning and its Connections to Adversarial Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.03525 v1 pith:LI7RYNUB submitted 2020-08-08 cs.LG cs.ITcs.ROmath.ITstat.ML

classification cs.LGcs.ITcs.ROmath.ITstat.ML
keywords learningimitationadversarialmethodsnon-adversarialpolicyconvergenceformulation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many modern methods for imitation learning and inverse reinforcement learning, such as GAIL or AIRL, are based on an adversarial formulation. These methods apply GANs to match the expert's distribution over states and actions with the implicit state-action distribution induced by the agent's policy. However, by framing imitation learning as a saddle point problem, adversarial methods can suffer from unstable optimization, and convergence can only be shown for small policy updates. We address these problems by proposing a framework for non-adversarial imitation learning. The resulting algorithms are similar to their adversarial counterparts and, thus, provide insights for adversarial imitation learning methods. Most notably, we show that AIRL is an instance of our non-adversarial formulation, which enables us to greatly simplify its derivations and obtain stronger convergence guarantees. We also show that our non-adversarial formulation can be used to derive novel algorithms by presenting a method for offline imitation learning that is inspired by the recent ValueDice algorithm, but does not rely on small policy updates for convergence. In our simulated robot experiments, our offline method for non-adversarial imitation learning seems to perform best when using many updates for policy and discriminator at each iteration and outperforms behavioral cloning and ValueDice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ContraDICE learns policies that imitate expert behavior while repelling undesirable demonstrations via a difference-of-KL objective that is convex when the expert weight dominates.

  2. On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A variational autoencoder over skill sequences produces label-free trajectory embeddings that separate and imitate policies of different ability levels in MuJoCo control tasks.

Pith tools