Pith. sign in

REVIEW 38 cited by

A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1011.0686 v3 pith:Q7APOC2K submitted 2010-11-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningalgorithmimitationapproachesassumptionsobservationsonlineperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sequential prediction problems such as imitation learning, where future observations depend on previous predictions (actions), violate the common i.i.d. assumptions made in statistical learning. This leads to poor performance in theory and often in practice. Some recent approaches provide stronger guarantees in this setting, but remain somewhat unsatisfactory as they train either non-stationary or stochastic policies and require a large number of iterations. In this paper, we propose a new iterative algorithm, which trains a stationary deterministic policy, that can be seen as a no regret algorithm in an online learning setting. We show that any such no regret algorithm, combined with additional reduction assumptions, must find a policy with good performance under the distribution of observations it induces in such sequential settings. We demonstrate that this new approach outperforms previous approaches on two challenging imitation learning problems and a benchmark sequence labeling problem.

Discussion (0). Sign in to comment.

Forward citations

Cited by 38 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Efficient Uniform Feasible-Set Sampling for Approximate Linear MPC

    eess.SY 2026-04 unverdicted novelty 7.0 of 10

    LMPC-HR sampler cuts computation time by roughly 10x for uniform sampling from polyhedral feasible sets of linear MPC by replacing iterative boundary searches with one convex LP per sample.

  2. RT-H: Action Hierarchies Using Language

    cs.RO 2024-03 conditional novelty 7.0 of 10

    RT-H learns robot policies by first predicting language motions as an intermediate representation and then mapping those plus the high-level task to actions, yielding more robust multi-task performance and the ability...

  3. Solving Rubik's Cube with a Robot Hand

    cs.LG 2019-10 accept novelty 7.0 of 10

    Reinforcement learning models trained only in simulation using automatic domain randomization solve Rubik's cube with a real robot hand.

  4. Concrete Problems in AI Safety

    cs.AI 2016-06 accept novelty 7.0 of 10

    The paper categorizes five concrete AI safety problems arising from flawed objectives, costly evaluation, and learning dynamics.

  5. Static In, Dynamic Out: Counterfactual Action Augmentation for Moving Object Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    SIDO morphs static demonstrations into counterfactual future-pose samples, training a goal-conditioned policy that, paired with a pose predictor, grasps objects whose motion was unseen during training.

  6. CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Action-conditioned world-model verification with conformal first-intervention control and latency-aware suffix repair raises RoboCasa365 success 8.5 points over invocation-matched periodic replanning.

  7. Learned Interventions in Lean 4 grind

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Failure-triggered interventions in Lean's grind tactic add a few solves and a small speedup with zero regressions, while static split-feature prediction fails to beat random.

  8. FORGE-plus: Force-Budgeted Recovery for Contact-Rich Assembly with a Frozen LLM Supervisor

    cs.RO 2026-07 conditional novelty 6.0 of 10

    With a hidden per-episode breaking force, an LLM-set force ceiling plus force-signature recovery achieves 256/256 clean insertions on fragile and robust parts and resolves 40–64% of injected jams in simulation.

  9. Agent-Centric Animal Pose Forecasting

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Agent-centric transformers trained through a composable library reproduce several marginal statistics of courting fly behavior, but discriminators still separate simulated from real flies.

  10. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  11. One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

    cs.RO 2026-07 unverdicted novelty 6.0 of 10

    AutoSERL achieves strong performance on six real-world robot manipulation tasks using RL guided by a single demonstration via sliding-window intervention, safety recovery, and automatic termination.

  12. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Aligning temporal granularity, action subspaces, and train-test conditioning yields SOTA long-horizon mobile and fine-grained manipulation success for a unified world-action model.

  13. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    ABot-M0.5 proposes a unified mobility-and-manipulation world action model using three alignment strategies that achieves state-of-the-art performance on mobile and fine-grained manipulation benchmarks.

  14. SoftSkill: Behavioral Compression for Contextual Adaptation

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    SoftSkill compresses agent skills into length-32 continuous prefixes via next-token training of soft deltas, yielding 5.2-12.5 point gains over SkillOpt on SearchQA and LiveMath while using far fewer tokens.

  15. Training and Evaluating Diffusion Policies with Long Context Lengths

    cs.RO 2026-06 conditional novelty 6.0 of 10

    Naive long-context Diffusion Policies succeed with UNet+Cross-Attention and sufficient data; variable-history training cuts sample complexity in the low-data regime.

  16. AEGIS: A Backup Reflex for Physical AI

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    AEGIS uses activation probes for early-warning detection of high-risk steps in weak policies and selectively escalates to stronger policies, recovering 10.1% of lost trajectories on LIBERO-Spatial while activating the...

  17. Pretraining Recurrent Networks without Recurrence

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    SMT reduces RNN training to supervised learning on memory transitions (m_t, x_{t+1}) to m_{t+1} obtained from a Transformer encoder, enabling time-parallel training with O(1) gradient paths.

  18. Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    LearnWeak specializes small CUAs via weakness detection by a reference agent, targeted task synthesis, and error-aware training, delivering 11+ point gains on OSWorld.

  19. When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    Position-Weighted On-Policy Self-Distillation (PW-OPSD) weights later tokens more heavily after a diagnostic shows position predicts teacher reliability better than entropy, yielding +1.0 and +1.1 Avg@12 gains on AIME...

  20. SigLoMa: Learning Open-World Quadrupedal Loco-Manipulation from Ego-Centric Vision

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    SigLoMa enables dynamic loco-manipulation on quadrupeds from ego-centric 5 Hz vision alone by using Sigma Points for scalable exteroception, an ego-centric Kalman Filter for high-rate state estimation, and an active s...

  21. Behavior-Constrained Reinforcement Learning with Receding-Horizon Credit Assignment for High-Performance Control

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    A behavior-constrained RL framework with receding-horizon credit assignment learns high-performance control policies that stay aligned with expert behavior in race car simulation.

  22. mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs

    cs.RO 2025-12 unverdicted novelty 6.0 of 10

    mimic-video combines internet video pretraining with a flow-matching decoder to achieve state-of-the-art robotic manipulation performance with 10x better sample efficiency than vision-language-action models.

  23. Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning

    cs.RO 2025-10 conditional novelty 6.0 of 10

    An imitation-learning policy with learned contrast-dye control navigates a soft robotic guidewire to unseen aneurysm targets in 83% of simulated fluoroscopy trials, matching clinician teleoperation performance.

  24. Arnold: a generalist muscle transformer policy

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A single transformer policy with a compositional sensorimotor vocabulary achieves expert or super-expert performance on 14 musculoskeletal control tasks spanning four embodiments.

  25. Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories

    cs.LG 2024-11 unverdicted novelty 6.0 of 10

    A Bayesian method uses near-optimality constraints from expert trajectories to estimate transition dynamics in offline model-based reinforcement learning.

  26. Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense

    cs.CR 2026-07 conditional novelty 5.0 of 10

    OSB proposes frozen synthetic-enterprise snapshots with gold posture answers so AI agents can be benchmarked on security investigation via SQL or native vendor APIs.

  27. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 conditional novelty 5.0 of 10

    A VLA policy with auxiliary success/progress heads and AWR+RECAP-style RL finished 1st in the LeHome 2026 simulation round and 2nd on the real robot.

  28. RSC: Decentralized Rigid Formation Flocking for Large-Scale Swarms via Hybrid Predictive Control and Online Reconfiguration

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    RSC achieves 83% success in maintaining rigid formations with 25 UAVs in cluttered environments via hybrid predictive control, APF safety, and stable leader-follower reconfiguration, outperforming baselines below 5%.

  29. SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    SPADE combines sketch-guided path planning with diffusion-augmented imitation learning to achieve better generalization and lower error with fewer parameters than prior methods.

  30. Efficient Uniform Feasible-Set Sampling for Approximate Linear MPC

    eess.SY 2026-04 conditional novelty 5.0 of 10

    LMPC-HR samples the polyhedral feasible set of linear MPC uniformly by replacing iterative boundary search with a single convex LP, cutting data-generation cost by roughly an order of magnitude.

  31. The Cartesian Cut in Agentic AI

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    LLM agents use a Cartesian split between learned prediction and engineered control, enabling modularity but creating sensitivity and bottlenecks unlike integrated biological systems.

  32. RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation

    cs.RO 2025-10 unverdicted novelty 5.0 of 10

    RESample uses exploratory sampling guided by a lightweight Coverage Function to expand VLA training data coverage, yielding 12% performance gains on LIBERO and real-world tasks with 10-20% added samples.

  33. Buzz, Choose, Forget: A Meta-Bandit Framework for Bee-Like Decision Making

    cs.LG 2025-10 reject novelty 5.0 of 10

    MAYA reproduces individual bee left/right choices by matching regret trajectories to four bandit policies with a memory window fixed at tau=7, but the tau value and best metric are selected on the same data used for e...

  34. Beyond Simulation: Benchmarking World Models for Planning and Causality in Autonomous Driving

    cs.RO 2025-08 conditional novelty 5.0 of 10

    Autoregressive traffic world models are overly sensitive to uncontrollable objects, and new delta metrics plus control dropout expose and reduce that sensitivity.

  35. Pretraining Recurrent Networks without Recurrence

    cs.LG 2026-06 conditional novelty 4.0 of 10

    SMT trains nonlinear RNNs by imitating one-step memory-transition labels generated by a Transformer, replacing BPTT's unrolled credit assignment with time-parallel supervised learning.

  36. Imitation Learning Based on Disentangled Representation Learning of Behavioral Characteristics

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A weakly-supervised CVAE with action chunking lets a robot change wiping speed online from instruction labels, but the same mechanism fails to disentangle wiping force and fails on spatial pick-and-place directives.

  37. Vision-Language-Action Models: Experimental Insights from a Real-World UR5 Platform

    cs.RO 2026-06 unverdicted novelty 3.0 of 10

    Real-robot trials with OpenVLA on a UR5e arm show consistent offline-to-closed-loop gaps driven by action semantics, coordinate conventions, temporal alignment, image preprocessing, and dataset quality rather than mod...

  38. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 unverdicted novelty 3.0 of 10

    A competition entry for bimanual garment folding won 1st in simulation and 2nd in reality by making a VLA policy predict its own value quantities to drive advantage estimation, failure detection, and action selection.

Pith tools