Pith. sign in

REVIEW 1 cited by

Error Bounds of Imitating Policies and Environments

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.11876 v1 pith:TSHIGJRK submitted 2020-10-22 cs.LG

classification cs.LG
keywords imitationlearningadversarialbehavioralcloningenvironmentgenerativeimitating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning trains a policy by mimicking expert demonstrations. Various imitation methods were proposed and empirically evaluated, meanwhile, their theoretical understanding needs further studies. In this paper, we firstly analyze the value gap between the expert policy and imitated policies by two imitation methods, behavioral cloning and generative adversarial imitation. The results support that generative adversarial imitation can reduce the compounding errors compared to behavioral cloning, and thus has a better sample complexity. Noticed that by considering the environment transition model as a dual agent, imitation learning can also be used to learn the environment model. Therefore, based on the bounds of imitating policies, we further analyze the performance of imitating environments. The results show that environment models can be more effectively imitated by generative adversarial imitation than behavioral cloning, suggesting a novel application of adversarial imitation for model-based reinforcement learning. We hope these results could inspire future advances in imitation learning and model-based reinforcement learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Few-Shot Neuro-Symbolic Imitation Learning for Long-Horizon Planning and Acting

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A neuro-symbolic system learns symbolic task rules and neural control policies from as few as five demonstrations and generalizes to larger unseen task instances.

Pith tools