Pith. sign in

REVIEW 3 cited by

Efficient Active Imitation Learning with Random Network Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.01894 v2 pith:SX4DJQMQ submitted 2024-11-04 cs.LG

classification cs.LG
keywords expertimitationlearningactivernd-daggerwhenagentsclear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing agents for complex and underspecified tasks, where no clear objective exists, remains challenging but offers many opportunities. This is especially true in video games, where simulated players (bots) need to play realistically, and there is no clear reward to evaluate them. While imitation learning has shown promise in such domains, these methods often fail when agents encounter out-of-distribution scenarios during deployment. Expanding the training dataset is a common solution, but it becomes impractical or costly when relying on human demonstrations. This article addresses active imitation learning, aiming to trigger expert intervention only when necessary, reducing the need for constant expert input along training. We introduce Random Network Distillation DAgger (RND-DAgger), a new active imitation learning method that limits expert querying by using a learned state-based out-of-distribution measure to trigger interventions. This approach avoids frequent expert-agent action comparisons, thus making the expert intervene only when it is useful. We evaluate RND-DAgger against traditional imitation learning and other active approaches in 3D video games (racing and third-person navigation) and in a robotic locomotion task and show that RND-DAgger surpasses previous methods by reducing expert queries. https://sites.google.com/view/rnd-dagger

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

    cs.AI 2025-09 conditional novelty 6.0 of 10

    EAPO lets a policy model consult a stronger expert during training, anneals that access to zero, and improves independent math reasoning by about 5 points over self-exploratory RL.

  2. Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A learned proxy Q-function that triggers expert help when agent and expert actions diverge reduces human takeover cost and improves imitation learning efficiency in simulated driving and navigation.

  3. TROFI: Trajectory-Ranked Offline Inverse Reinforcement Learning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    TROFI learns a reward model from ranked trajectories, labels an offline dataset with it, and trains a TD3+BC policy, matching ground-truth-reward performance on many D4RL tasks without a hand-coded reward or expert de...

Pith tools