Pith. sign in

REVIEW 14 cited by

Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.08558 v3 pith:EQRE24FC submitted 2025-03-11 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords failuredetectionpolicyfailuresimitationroboticdatafail-detect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed impressive robotic manipulation systems driven by advances in imitation learning and generative modeling, such as diffusion- and flow-based approaches. As robot policy performance increases, so does the complexity and time horizon of achievable tasks, inducing unexpected and diverse failure modes that are difficult to predict a priori. To enable trustworthy policy deployment in safety-critical human environments, reliable runtime failure detection becomes important during policy inference. However, most existing failure detection approaches rely on prior knowledge of failure modes and require failure data during training, which imposes a significant challenge in practicality and scalability. In response to these limitations, we present FAIL-Detect, a modular two-stage approach for failure detection in imitation learning-based robotic manipulation. To accurately identify failures from successful training data alone, we frame the problem as sequential out-of-distribution (OOD) detection. We first distill policy inputs and outputs into scalar signals that correlate with policy failures and capture epistemic uncertainty. FAIL-Detect then employs conformal prediction (CP) as a versatile framework for uncertainty quantification with statistical guarantees. Empirically, we thoroughly investigate both learned and post-hoc scalar signal candidates on diverse robotic manipulation tasks. Our experiments show learned signals to be mostly consistently effective, particularly when using our novel flow-based density estimator. Furthermore, our method detects failures more accurately and faster than state-of-the-art (SOTA) failure detection baselines. These results highlight the potential of FAIL-Detect to enhance the safety and reliability of imitation learning-based robotic systems as they progress toward real-world deployment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    ActProbe is an action-space detector that uses temporal consistency error and action chunk magnitude from policy outputs, mapped via LSTM-MLP, to predict failures earlier than baselines across policies and real-robot tasks.

  2. Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

    cs.RO 2026-04 unverdicted novelty 7.0 of 10

    FIDeL detects failures in imitation learning by building compact nominal representations via optimal transport, applying conformal prediction thresholds, and using VLMs for semantic filtering, outperforming baselines ...

  3. The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A single-pass measure of how much a flow-matching action trajectory bends correlates with the model's uncertainty and can flag impending robot failures for free.

  4. RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Adaptive compositional steering of frozen VLAs with a latent-conditioned offline RL flow policy, gated by failure prediction, improves OOD manipulation success by up to +17.3%.

  5. Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Foresight detects failures in long-horizon robotic manipulation using latents from action-conditioned world models trained only on task-level labels and calibrated via functional conformal prediction.

  6. Robot Critics that Sweat the Small Stuff

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Fine-tuning VLMs with pairwise progress supervision from policy rollouts improves fine-grained failure detection and boosts robot manipulation success by 11% real-world and 5.9% in simulation.

  7. Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Tri-Info uses three information theory signals on action diversity, temporal consistency, and state coupling to predict VLA model failures with cross-domain generalization to 83% real-world accuracy.

  8. AEGIS: A Backup Reflex for Physical AI

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    AEGIS uses activation probes for early-warning detection of high-risk steps in weak policies and selectively escalates to stronger policies, recovering 10.1% of lost trajectories on LIBERO-Spatial while activating the...

  9. Flow-based Policy Adaptation without Policy Updates

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    GLOVES learns flow models from limited expert demonstrations to selectively correct actions from non-expert policies or operators toward expert distributions using reverse-flow OOD detection as an intervention gate.

  10. VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

    cs.RO 2026-05 unverdicted novelty 6.0 of 10

    VLAConf is a one-class discriminative method that estimates step-wise task-success confidence for VLA models via anomaly scoring on frozen representations plus step-conditioned modeling, shown to be more efficient tha...

  11. CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding

    cs.RO 2026-01 reject novelty 6.0 of 10

    CycleVLA adds progress-triggered VLM failure checks, subtask backtracking, and MBR consensus decoding to VLAs, raising LIBERO average success from 89.3% to 95.3% and claiming 91% real-robot success.

  12. InFeR: Informed Failure Resilience in Learned Visual Navigation Control

    cs.RO 2025-10 unverdicted novelty 6.0 of 10

    InFeR retrains imitation learning policies with a VIB loss for OOD failure detection and applies Grad-CAM to localize failure sources, enabling heuristic recovery in visual navigation without additional demonstrations.

  13. INSIGHT: INference-time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Token-level uncertainty sequences from a VLA policy, classified by a small transformer, predict when a robot should request human help better than static uncertainty scores.

  14. RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

    cs.RO 2026-07 conditional novelty 5.0 of 10

    RL² improves VLA robot success rates by conditionally composing an offline RL policy's actions with the frozen VLA only when a failure detector flags impending failure.

Pith tools