Pith. sign in

REVIEW 1 cited by

Visually Robust Adversarial Imitation Learning from Videos with Contrastive Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12792 v2 pith:SLCG7HON submitted 2024-06-18 cs.LG cs.CV

classification cs.LGcs.CV
keywords learningimitationc-laiforobustspacevideosadversarialalgorithm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose C-LAIfO, a computationally efficient algorithm designed for imitation learning from videos in the presence of visual mismatch between agent and expert domains. We analyze the problem of imitation from expert videos with visual discrepancies, and introduce a solution for robust latent space estimation using contrastive learning and data augmentation. Provided a visually robust latent space, our algorithm performs imitation entirely within this space using off-policy adversarial imitation learning. We conduct a thorough ablation study to justify our design and test C-LAIfO on high-dimensional continuous robotic tasks. Additionally, we demonstrate how C-LAIfO can be combined with other reward signals to facilitate learning on a set of challenging hand manipulation tasks with sparse rewards. Our experiments show improved performance compared to baseline methods, highlighting the effectiveness of C-LAIfO. To ensure reproducibility, we open source our code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Converting RGB videos into event-like temporal gradients lets an imitation agent ignore appearance differences between expert and learner domains, improving robustness without data augmentation.

Pith tools