Pith. sign in

REVIEW 5 cited by

Feedback in Imitation Learning: The Three Regimes of Covariate Shift

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.02872 v2 pith:BR2KXGU2 submitted 2021-02-04 cs.LG cs.ROstat.ML

classification cs.LGcs.ROstat.ML
keywords divergenceshiftbenchmarkscausalcovariateimitationlearningproblems
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning practitioners have often noted that conditioning policies on previous actions leads to a dramatic divergence between "held out" error and performance of the learner in situ. Interactive approaches can provably address this divergence but require repeated querying of a demonstrator. Recent work identifies this divergence as stemming from a "causal confound" in predicting the current action, and seek to ablate causal aspects of current state using tools from causal inference. In this work, we argue instead that this divergence is simply another manifestation of covariate shift, exacerbated particularly by settings of feedback between decisions and input features. The learner often comes to rely on features that are strongly predictive of decisions, but are subject to strong covariate shift. Our work demonstrates a broad class of problems where this shift can be mitigated, both theoretically and practically, by taking advantage of a simulator but without any further querying of expert demonstration. We analyze existing benchmarks used to test imitation learning approaches and find that these benchmarks are realizable and simple and thus insufficient for capturing the harder regimes of error compounding seen in real-world decision making problems. We find, in a surprising contrast with previous literature, but consistent with our theory, that naive behavioral cloning provides excellent results. We detail the need for new standardized benchmarks that capture the phenomena seen in robotics problems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned world-model verification with conformal first-intervention control and latency-aware suffix repair raises RoboCasa365 success 8.5 points over invocation-matched periodic replanning.

  2. Scalable Causal Imitation Learning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Causal SQIL and Causal IQ-Learn combine sliding-window causal adjustment with off-policy soft Q-learning, scaling causal imitation learning to long-horizon continuous control.

  3. The Three Regimes of Offline-to-Online Reinforcement Learning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Offline-to-online fine-tuning works best when the chosen method preserves the stronger of the pretrained policy or the offline dataset; a three-regime taxonomy organizes these choices.

  4. Confounded Causal Imitation Learning with Instrumental Variables

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A causal imitation learning framework that identifies valid instrumental variables from observational data and uses them to learn policies robust to multi-timestep latent confounders.

  5. DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DySS combines state-space feature learning with dynamic query merging and pruning to improve both accuracy and speed for camera-based 3D detection on nuScenes.

Pith tools