Pith. sign in

REVIEW 6 cited by

Is Behavior Cloning All You Need? Understanding Horizon in Imitation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15007 v2 pith:LPURC2L6 submitted 2024-07-20 cs.LG cs.AImath.STstat.MLstat.TH

classification cs.LGcs.AImath.STstat.MLstat.TH
keywords offlineonlinebehaviorhorizonlearningcloningcomplexitydependence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning (IL) aims to mimic the behavior of an expert in a sequential decision making task by learning from demonstrations, and has been widely applied to robotics, autonomous driving, and autoregressive text generation. The simplest approach to IL, behavior cloning (BC), is thought to incur sample complexity with unfavorable quadratic dependence on the problem horizon, motivating a variety of different online algorithms that attain improved linear horizon dependence under stronger assumptions on the data and the learner's access to the expert. We revisit the apparent gap between offline and online IL from a learning-theoretic perspective, with a focus on the realizable/well-specified setting with general policy classes up to and including deep neural networks. Through a new analysis of behavior cloning with the logarithmic loss, we show that it is possible to achieve horizon-independent sample complexity in offline IL whenever (i) the range of the cumulative payoffs is controlled, and (ii) an appropriate notion of supervised learning complexity for the policy class is controlled. Specializing our results to deterministic, stationary policies, we show that the gap between offline and online IL is smaller than previously thought: (i) it is possible to achieve linear dependence on horizon in offline IL under dense rewards (matching what was previously only known to be achievable in online IL); and (ii) without further assumptions on the policy class, online IL cannot improve over offline IL with the logarithmic loss, even in benign MDPs. We complement our theoretical results with experiments on standard RL tasks and autoregressive language generation to validate the practical relevance of our findings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-agent imitation learning with function approximation: Linear Markov games and beyond

    cs.LG 2026-02 conditional novelty 7.0 of 10

    In linear Markov games, behavior cloning's sample complexity hinges on a feature-level concentrability coefficient, and the interactive algorithm LSVI-UCB-ZERO-BC removes concentrability dependence entirely, scaling o...

  2. Agent-Centric Animal Pose Forecasting

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Agent-centric transformers trained through a composable library reproduce several marginal statistics of courting fly behavior, but discriminators still separate simulated from real flies.

  3. Comparing Behavioural Cloning and Reinforcement Learning for Spacecraft Guidance and Control Networks

    eess.SY 2025-07 conditional novelty 6.0 of 10

    On four spacecraft guidance problems, reinforcement learning trains networks that are more robust to disturbances than behavioural cloning, but behavioural cloning matches optimal control when the expert data is accurate.

  4. Behavioral Exploration: Learning to Explore via In-Context Adaptation

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A coverage-conditioned behavioral cloning policy adapts in-context to its own history, making robots explore new expert-like behaviors online without online reinforcement learning.

  5. Buzz, Choose, Forget: A Meta-Bandit Framework for Bee-Like Decision Making

    cs.LG 2025-10 reject novelty 5.0 of 10

    MAYA reproduces individual bee left/right choices by matching regret trajectories to four bandit policies with a memory window fixed at tau=7, but the tau value and best metric are selected on the same data used for e...

  6. PRISM: Pointcloud Reintegrated Inference via Segmentation and Cross-attention for Manipulation

    cs.RO 2025-07 conditional novelty 4.0 of 10

    PRISM trains a diffusion policy on segmented point-cloud object tokens fused with joint states via cross-attention, reporting 82.0 percent average success across six RoboTwin tasks versus 58.4 percent for DP3 and 22.3...

Pith tools