Pith. sign in

REVIEW 8 cited by

Sim-to-Real Reinforcement Learning for Vision-Based Dexterous Manipulation on Humanoids

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.20396 v2 pith:C32IHJXZ submitted 2025-02-27 cs.RO cs.AIcs.CVcs.LGcs.SYeess.SY

classification cs.ROcs.AIcs.CVcs.LGcs.SYeess.SY
keywords manipulationlearningsim-to-realdexteroustasksvision-basedbimanualhumanoid
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning generalizable robot manipulation policies, especially for complex multi-fingered humanoids, remains a significant challenge. Existing approaches primarily rely on extensive data collection and imitation learning, which are expensive, labor-intensive, and difficult to scale. Sim-to-real reinforcement learning (RL) offers a promising alternative, but has mostly succeeded in simpler state-based or single-hand setups. How to effectively extend this to vision-based, contact-rich bimanual manipulation tasks remains an open question. In this paper, we introduce a practical sim-to-real RL recipe that trains a humanoid robot to perform three challenging dexterous manipulation tasks: grasp-and-reach, box lift and bimanual handover. Our method features an automated real-to-sim tuning module, a generalized reward formulation based on contact and object goals, a divide-and-conquer policy distillation framework, and a hybrid object representation strategy with modality-specific augmentation. We demonstrate high success rates on unseen objects and robust, adaptive policy behaviors -- highlighting that vision-based dexterous manipulation via sim-to-real RL is not only viable, but also scalable and broadly applicable to real-world humanoid manipulation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routing

    cs.RO 2026-07 conditional novelty 7.0 of 10

    SILO enables the first reported zero-shot sim-to-real RL transfer for multi-stage cable routing by approximating cables as articulated rigid links and executing policy actions inside a synchronized digital twin.

  2. Closing the Reality Gap: Zero-Shot Sim-to-Real Deployment for Dexterous Force-Based Grasping and Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Zero-shot sim-to-real RL policies on a five-finger hand achieve commandable grasp-force tracking and in-hand reorientation using dense tactile simulation, current-to-torque calibration, and actuator randomization.

  3. ClutterDexGrasp: A Sim-to-Real System for General Dexterous Grasping in Cluttered Scenes

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A simulation-trained teacher-student policy achieves zero-shot sim-to-real closed-loop target-oriented dexterous grasping in cluttered scenes, with 83.9 percent real-world success.

  4. CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 5.5 of 10

    Residual waypoint RL around a frozen VLA prior, scaled via heterogeneous CARLA/H100 infrastructure, raises closed-loop driving score and success rate on longest6 v2 and Bench2Drive.

  5. Where to Touch, How to Contact: A Hierarchical RL-MPC Framework for Geometry-Aware Sim-to-Real Manipulation

    cs.RO 2026-01 conditional novelty 5.0 of 10

    A hierarchical RL-MPC framework with a 'contact intention' interface achieves data-efficient, robust non-prehensile manipulation that transfers zero-shot to a real robot.

  6. HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    HERMES converts a single human motion demonstration into a deployable mobile bimanual dexterous manipulation policy, using RL, depth-image distillation, and closed-loop PnP pose refinement.

  7. EmbodieDreamer: Advancing Real2Sim2Real Transfer for Policy Training via Embodied World Modeling

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A Real2Sim2Real framework that aligns simulator dynamics via differentiable parameter fitting and renders photorealistic policy-training videos with a diffusion model, improving real-world manipulation success.

  8. Learning Dexterous Object Handover

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A dual quaternion-based reward function yields more robust learned object handover with a four-finger hand than Euler or matrix rotation rewards in simulation.

Pith tools