Pith. sign in

REVIEW 14 cited by

From Imitation to Refinement -- Residual RL for Precise Assembly

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.16677 v4 pith:EXYZIGTH submitted 2024-07-23 cs.RO cs.LG

From Imitation to Refinement -- Residual RL for Precise Assembly

classification cs.RO cs.LG
keywords closed-loopactiondataperformanceresidualchunkscriticaldistribution
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances in Behavior Cloning (BC) have made it easy to teach robots new tasks. However, we find that the ease of teaching comes at the cost of unreliable performance that saturates with increasing data for tasks requiring precision. The performance saturation can be attributed to two critical factors: (a) distribution shift resulting from the use of offline data and (b) the lack of closed-loop corrective control caused by action chucking (predicting a set of future actions executed open-loop) critical for BC performance. Our key insight is that by predicting action chunks, BC policies function more like trajectory "planners" than closed-loop controllers necessary for reliable execution. To address these challenges, we devise a simple yet effective method, ResiP (Residual for Precise Manipulation), that overcomes the reliability problem while retaining BC's ease of teaching and long-horizon capabilities. ResiP augments a frozen, chunked BC model with a fully closed-loop residual policy trained with reinforcement learning (RL) that addresses distribution shifts and introduces closed-loop corrections over open-loop execution of action chunks predicted by the BC trajectory planner. Videos, code, and data: https://residual-assembly.github.io.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VINE: Taming Generative Control Policies for Reinforcement Learning

    cs.RO 2026-07 conditional novelty 7.0

    Reconstructing a fresh noisy interpolation state at every denoising step stabilizes end-to-end value-gradient training of multi-step flow-matching policies and yields state-of-the-art offline and real-robot results.

  2. EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

    cs.RO 2026-06 unverdicted novelty 7.0

    EgoEngine transforms egocentric human videos into high-fidelity robot data enabling zero-shot visuomotor dexterous policy learning without real-robot demonstrations.

  3. Dynamic Execution Horizon Prediction for Chunk-based Robot Policies

    cs.RO 2026-06 unverdicted novelty 7.0

    DEHP adds an online-RL horizon predictor to frozen chunk policies, yielding higher success on precise and long-horizon robot manipulation by adapting chunk length to task stage.

  4. EXPO: Stable Reinforcement Learning with Expressive Policies

    cs.LG 2025-07 conditional novelty 7.0

    EXPO stabilizes online RL for expressive policies by training a base policy with imitation and using a lightweight Gaussian edit policy to select higher-value actions on the fly for sampling and TD backups.

  5. Steering Your Diffusion Policy with Latent Space Reinforcement Learning

    cs.RO 2025-06 unverdicted novelty 7.0

    DSRL steers pretrained diffusion policies for robotics by applying RL to their latent noise inputs, achieving sample-efficient real-world adaptation with only black-box access.

  6. The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    For high-precision manipulation, required demonstration count grows as log(N) ∝ 1/(P−c), where the fitted c varies with sensors, expert, and task complexity.

  7. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  8. RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations

    cs.RO 2026-05 unverdicted novelty 6.0

    RoHIL adapts human-in-the-loop RL policies to new illumination conditions offline by combining world-model image relighting, illumination-retention replay, and anchored Bellman regularisation, improving shifted-light ...

  9. Diffusion Policy Policy Optimization

    cs.RO 2024-09 unverdicted novelty 6.0

    DPPO fine-tunes diffusion policies via policy gradients and outperforms prior RL approaches for diffusion policies and PG-tuned alternatives on robot benchmarks while enabling stable training and hardware deployment.

  10. Dynamic Execution Horizon Prediction for Chunk-based Robot Policies

    cs.RO 2026-06 unverdicted novelty 5.0

    A frozen chunk-based robot policy plus a lightweight RL-trained horizon predictor raises success on high-precision and long-horizon manipulation by adapting open-loop length on the fly.

  11. HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks

    cs.RO 2026-05 unverdicted novelty 5.0

    HDFlow pairs a high-level diffusion planner for subgoals with a low-level rectified flow planner for trajectories, outperforming prior methods on furniture assembly and locomotion-manipulation benchmarks.

  12. HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks

    cs.RO 2026-05 unverdicted novelty 5.0

    HDFlow pairs a high-level diffusion planner for strategic subgoals with a low-level rectified flow planner for efficient trajectories, claiming superior performance on furniture assembly and other long-horizon robotic...

  13. Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

    cs.LG 2026-03 unverdicted novelty 5.0

    Approximation error of constant-depth parallelizable sequence models falls exponentially with depth, via a Lie-algebraic tower of expressivity extensions.

  14. Robot Self-Improvement via Human-Video Dynamics Models

    cs.RO 2026-06 unverdicted novelty 4.0

    Human-video dynamics models enable cross-embodiment robot self-improvement via training-free Dynamics-Guided Action Correction, raising success rates from 40% to 81% on seven real-world tasks.