Pith. sign in

REVIEW 6 cited by

Residual Reinforcement Learning from Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.08050 v1 pith:OD7HCS2E submitted 2021-06-15 cs.LG

Residual Reinforcement Learning from Demonstrations

classification cs.LG
keywords demonstrationsresidualcontrollerlearningtasksinputsreinforcementreward
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Residual reinforcement learning (RL) has been proposed as a way to solve challenging robotic tasks by adapting control actions from a conventional feedback controller to maximize a reward signal. We extend the residual formulation to learn from visual inputs and sparse rewards using demonstrations. Learning from images, proprioceptive inputs and a sparse task-completion reward relaxes the requirement of accessing full state features, such as object and target positions. In addition, replacing the base controller with a policy learned from demonstrations removes the dependency on a hand-engineered controller in favour of a dataset of demonstrations, which can be provided by non-experts. Our experimental evaluation on simulated manipulation tasks on a 6-DoF UR5 arm and a 28-DoF dexterous hand demonstrates that residual RL from demonstrations is able to generalize to unseen environment conditions more flexibly than either behavioral cloning or RL fine-tuning, and is capable of solving high-dimensional, sparse-reward tasks out of reach for RL from scratch.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand

    cs.RO 2026-06 unverdicted novelty 7.0

    DexCompose achieves 77.4% average success on 16 composite dexterous tasks by using role-aware residual composition with explicit finger ownership to combine pretrained policies without destructive interference.

  2. AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance

    cs.RO 2026-06 unverdicted novelty 6.0

    AnyBody distills a privileged teacher tracker into a latent unit-sphere representation and uses a masked transformer to drive humanoid control from arbitrary keypoint subsets.

  3. Diffusion Policy Policy Optimization

    cs.RO 2024-09 unverdicted novelty 6.0

    DPPO fine-tunes diffusion policies via policy gradients and outperforms prior RL approaches for diffusion policies and PG-tuned alternatives on robot benchmarks while enabling stable training and hardware deployment.

  4. Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning

    cs.AI 2026-07 conditional novelty 5.5

    Offline residual SAC correction on CNN–LSTM kinematic outputs improves continuous 3D EEG motor-imagery decoding by ~21–42% in correlation and RMSE across 2D and VR feedback without feeding EEG to the RL agent.

  5. TacCoRL: Integrating Tactile Feedback into VLA via Simulation

    cs.RO 2026-06 unverdicted novelty 5.0

    TacCoRL integrates tactile feedback into VLA policies via real-aligned simulation co-training and RL, raising average success from 50% to 72.5% on four bimanual contact-rich tasks with direct real-robot transfer.

  6. An Agency-Transferring Model-Free Policy Enhancement Technique

    cs.LG 2026-06 unverdicted novelty 5.0

    A model-free RL method arbitrates between a functional baseline policy and a learning policy, transferring agency over time to yield a standalone policy with high goal-reaching rates and competitive returns on continu...