Pith. sign in

REVIEW 1 cited by

Imitation Learning from Observations under Transition Model Disparity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.11446 v1 pith:SOE6APMF submitted 2022-04-25 cs.LG stat.ML

classification cs.LGstat.ML
keywords expertlearninglearnerdynamicsobservationstransitionalgorithmdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning to perform tasks by leveraging a dataset of expert observations, also known as imitation learning from observations (ILO), is an important paradigm for learning skills without access to the expert reward function or the expert actions. We consider ILO in the setting where the expert and the learner agents operate in different environments, with the source of the discrepancy being the transition dynamics model. Recent methods for scalable ILO utilize adversarial learning to match the state-transition distributions of the expert and the learner, an approach that becomes challenging when the dynamics are dissimilar. In this work, we propose an algorithm that trains an intermediary policy in the learner environment and uses it as a surrogate expert for the learner. The intermediary policy is learned such that the state transitions generated by it are close to the state transitions in the expert dataset. To derive a practical and scalable algorithm, we employ concepts from prior work on estimating the support of a probability distribution. Experiments using MuJoCo locomotion tasks highlight that our method compares favorably to the baselines for ILO with transition dynamics mismatch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations

    cs.CV 2025-02 conditional novelty 5.0 of 10

    DIFF-IL combines per-frame domain-invariant feature extraction with frame-wise time labeling to improve cross-domain imitation learning from images, beating prior methods on 14 tasks.

Pith tools