Pith. sign in

REVIEW 22 cited by

Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21845 v3 pith:Q7K3K5ZM submitted 2024-10-29 cs.RO cs.AI

Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning

classification cs.RO cs.AI
keywords manipulationlearningpoliciesroboticapproachcomplexdexteroushuman-in-the-loop
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Reinforcement learning (RL) holds great promise for enabling autonomous acquisition of complex robotic manipulation skills, but realizing this potential in real-world settings has been challenging. We present a human-in-the-loop vision-based RL system that demonstrates impressive performance on a diverse set of dexterous manipulation tasks, including dynamic manipulation, precision assembly, and dual-arm coordination. Our approach integrates demonstrations and human corrections, efficient RL algorithms, and other system-level design choices to learn policies that achieve near-perfect success rates and fast cycle times within just 1 to 2.5 hours of training. We show that our method significantly outperforms imitation learning baselines and prior RL approaches, with an average 2x improvement in success rate and 1.8x faster execution. Through extensive experiments and analysis, we provide insights into the effectiveness of our approach, demonstrating how it learns robust, adaptive policies for both reactive and predictive control strategies. Our results suggest that RL can indeed learn a wide range of complex vision-based manipulation policies directly in the real world within practical training times. We hope this work will inspire a new generation of learned robotic manipulation techniques, benefiting both industrial applications and research advancements. Videos and code are available at our project website https://hil-serl.github.io/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Improving Robotic Generalist Policies via Flow Reversal Steering

    cs.RO 2026-06 unverdicted novelty 7.0

    Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.

  2. Steering Your Diffusion Policy with Latent Space Reinforcement Learning

    cs.RO 2025-06 unverdicted novelty 7.0

    DSRL steers pretrained diffusion policies for robotics by applying RL to their latent noise inputs, achieving sample-efficient real-world adaptation with only black-box access.

  3. The Curse of Precision: A Data Scaling Law for High-Precision Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0

    For high-precision manipulation, required demonstration count grows as log(N) ∝ 1/(P−c), where the fitted c varies with sensors, expert, and task complexity.

  4. FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

    cs.RO 2026-07 conditional novelty 6.0

    Human corrective actions can be inverted into noise-space targets that train a lightweight latent policy to steer frozen flow/diffusion robot models from a handful of interventions while preserving pretrained skills.

  5. OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies

    cs.RO 2026-07 conditional novelty 6.0

    A policy-agnostic two-stage real-world RL method learns tactile residual corrections on frozen visual policies, lifting contact-rich task success from 5–40% to 85–100% in under 80 minutes.

  6. RhinoVLA Technical Report

    cs.RO 2026-06 conditional novelty 6.0

    RhinoVLA matches π0.5-scale VLA task performance while reaching 11.69 Hz end-to-end inference on the Huixi R1 edge SoC via token-efficient Qwen3-VL, a 72D unified action interface, and hardware co-optimization.

  7. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 unverdicted novelty 6.0

    M²-VLA shows that generalized VLMs can serve as direct backbones for robotic manipulation by selectively extracting task-critical features via Mixture of Layers and adding Meta Skill Modules for efficient trajectory learning.

  8. $M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills

    cs.RO 2026-04 conditional novelty 6.0

    Freezing a VLM backbone and routing its layers through Mixture-of-Layers plus a Meta-Skill memory yields higher success and stronger zero-shot generalization than fine-tuned VLAs on LIBERO and real robots.

  9. RL Token: Bootstrapping Online RL with Vision-Language-Action Models

    cs.LG 2026-04 unverdicted novelty 6.0

    RL Token enables sample-efficient online RL fine-tuning of large VLAs, delivering up to 3x speed gains and higher success rates on real-robot manipulation tasks within minutes to hours.

  10. QDTraj: Exploration of Diverse Trajectory Primitives for Articulated Objects Robotic Manipulation

    cs.RO 2026-04 unverdicted novelty 6.0

    QDTraj uses Quality-Diversity algorithms with sparse rewards to produce at least five times more diverse high-performing trajectories for articulated object manipulation than compared methods, validated across 30 obje...

  11. VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

    cs.RO 2025-05 conditional novelty 6.0

    VLA-RL applies online RL to pretrained VLAs, yielding a 4.5% gain over strong baselines on 40 LIBERO manipulation tasks and matching commercial models like π₀-FAST.

  12. AhaRobot: A Low-Cost Open-Source Bimanual Mobile Manipulator for Embodied AI

    cs.RO 2025-03 conditional novelty 6.0

    AhaRobot delivers 0.7 mm repeatability on a $1000 bimanual platform using dual-motor compensation and a novel 26-faced marker handle that cuts tracking error 80% versus a 6-faced baseline.

  13. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 conditional novelty 5.0

    A VLA policy with auxiliary success/progress heads and AWR+RECAP-style RL finished 1st in the LeHome 2026 simulation round and 2nd on the real robot.

  14. Scaling by Diversified Experience for Vision-Language-Action Models

    cs.CV 2026-06 unverdicted novelty 5.0

    SyVLA uses Intention Decoupling and similar-sample guided RL on diversified experiences to improve VLA model task success and out-of-distribution generalization while keeping vision-language abilities.

  15. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 5.0

    RhinoVLA uses a token-efficient Qwen3-VL backbone, continuous Action Expert, and unified cross-robot interface to match π0.5 performance while hitting 11.69 Hz on Huixi R1 edge SoC.

  16. RhinoVLA Technical Report

    cs.RO 2026-06 conditional novelty 5.0

    RhinoVLA runs a ~2.5B-parameter robot policy end-to-end at 11.69 Hz on the Huixi R1 edge SoC with LIBERO accuracy within a few points of π0.5 and mixed real-robot results.

  17. REAP: Reinforcement-Learning End-to-End Autonomous Parking with Gaussian Splatting Simulator for Real2Sim2Real Transfer

    cs.RO 2026-05 unverdicted novelty 5.0

    REAP trains an end-to-end SAC policy with behavior cloning and collision penalties inside a 3DGS Real2Sim simulator and transfers it to physical vehicles, succeeding in narrow mechanical parking slots.

  18. RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction

    cs.RO 2025-09 conditional novelty 5.0

    Robot policies trained on human interventions that rewind to a familiar state and then correct the mistake achieve higher long-horizon success and better data efficiency than imitation on full demonstrations alone.

  19. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  20. EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

    cs.RO 2026-05 unverdicted novelty 4.0

    EXPO-FT enables pretrained VLA policies to reach 30/30 success on complex manipulation tasks using an average of 19.1 minutes of online robot data while outperforming prior RL approaches.

  21. Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

    cs.RO 2026-06 unverdicted novelty 3.0

    A competition entry for bimanual garment folding won 1st in simulation and 2nd in reality by making a VLA policy predict its own value quantities to drive advantage estimation, failure detection, and action selection.

  22. RhinoVLA Technical Report

    cs.RO 2026-06 unverdicted novelty 3.0

    RhinoVLA cuts VLM tokens with a Qwen3-VL backbone and continuous action expert, adds a unified cross-robot interface, and reaches real-time 11.69 Hz on Huixi R1 while matching π0.5 downstream performance.