Pith. sign in

REVIEW 1 cited by

MEGA-DAgger: Imitation Learning with Multiple Imperfect Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.00638 v3 pith:ROSFG4J7 submitted 2023-03-01 cs.LG cs.RO

classification cs.LGcs.RO
keywords expertslearningimitationimperfectinteractivemega-daggermultiplealgorithms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning has been widely applied to various autonomous systems thanks to recent development in interactive algorithms that address covariate shift and compounding errors induced by traditional approaches like behavior cloning. However, existing interactive imitation learning methods assume access to one perfect expert. Whereas in reality, it is more likely to have multiple imperfect experts instead. In this paper, we propose MEGA-DAgger, a new DAgger variant that is suitable for interactive learning with multiple imperfect experts. First, unsafe demonstrations are filtered while aggregating the training data, so the imperfect demonstrations have little influence when training the novice policy. Next, experts are evaluated and compared on scenarios-specific metrics to resolve the conflicted labels among experts. Through experiments in autonomous racing scenarios, we demonstrate that policy learned using MEGA-DAgger can outperform both experts and policies learned using the state-of-the-art interactive imitation learning algorithms such as Human-Gated DAgger. The supplementary video can be found at \url{https://youtu.be/wPCht31MHrw}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dynamic Rank Adjustment in Diffusion Policies for Efficient and Flexible Training

    cs.RO 2025-02 conditional novelty 6.0 of 10

    By dynamically freezing low-rank singular components of diffusion policy weights during training, DRIFT-DAgger cuts training time by roughly 11 to 18 percent while keeping task success near full-rank baselines.

Pith tools