Pith. sign in

REVIEW 2 cited by

Confidence-Aware Imitation Learning from Demonstrations with Varying Optimality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.14754 v2 pith:APAY47CJ submitted 2021-10-27 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords demonstrationsoptimalityimitationlearningvaryingcailconfidencelearn
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most existing imitation learning approaches assume the demonstrations are drawn from experts who are optimal, but relaxing this assumption enables us to use a wider range of data. Standard imitation learning may learn a suboptimal policy from demonstrations with varying optimality. Prior works use confidence scores or rankings to capture beneficial information from demonstrations with varying optimality, but they suffer from many limitations, e.g., manually annotated confidence scores or high average optimality of demonstrations. In this paper, we propose a general framework to learn from demonstrations with varying optimality that jointly learns the confidence score and a well-performing policy. Our approach, Confidence-Aware Imitation Learning (CAIL) learns a well-performing policy from confidence-reweighted demonstrations, while using an outer loss to track the performance of our model and to learn the confidence. We provide theoretical guarantees on the convergence of CAIL and evaluate its performance in both simulated and real robot experiments. Our results show that CAIL significantly outperforms other imitation learning methods from demonstrations with varying optimality. We further show that even without access to any optimal demonstrations, CAIL can still learn a successful policy, and outperforms prior work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Latent Diffusion Planning for Imitation Learning

    cs.RO 2025-04 conditional novelty 6.0 of 10

    Latent Diffusion Planning separates latent-state forecasting from action inference, letting imitation learning use action-free and suboptimal data, and outperforms Diffusion Policy in low-demonstration manipulation tasks.

  2. Imitation Learning via Focused Satisficing

    cs.LG 2025-05 conditional novelty 5.0 of 10

    MinSubFI directly minimizes subdominance, a margin-based measure of failing to be acceptable, and empirically reports higher demonstrator acceptability than prior imitation methods.

Pith tools