Pith. sign in

REVIEW 12 cited by

Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04436 v1 pith:AW75AUUM submitted 2024-03-07 cs.RO cs.AIcs.LGcs.SYeess.SY

classification cs.ROcs.AIcs.LGcs.SYeess.SY
keywords humanoidreal-timeteleoperationwhole-bodymotionmotionsachievehuman
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of human movements for humanoid robots, we propose a scalable "sim-to-data" process to filter and pick feasible motions using a privileged motion imitator. Afterwards, we train a robust real-time humanoid motion imitator in simulation using these refined motions and transfer it to the real humanoid robot in a zero-shot manner. We successfully achieve teleoperation of dynamic whole-body motions in real-world scenarios, including walking, back jumping, kicking, turning, waving, pushing, boxing, etc. To the best of our knowledge, this is the first demonstration to achieve learning-based real-time whole-body humanoid teleoperation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding a panoramic camera feed to a vision-language-action policy raises end-to-end success on four real-world mobile two-arm tasks from 30% to 73%.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Online co-training of a text-to-motion generator and a humanoid tracker on simulated G1 improves generator executability and zero-shot tracker coverage beyond static replay or one-way filtering.

  4. SafeMimic: Towards Safe and Autonomous Human-to-Robot Imitation for Mobile Manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SafeMimic enables a mobile robot to safely and autonomously adapt a single third-person human video into a successful multi-step manipulation strategy.

  5. GMT: General Motion Tracking for Humanoid Whole-Body Control

    cs.RO 2025-06 conditional novelty 6.0 of 10

    GMT trains a single unified humanoid policy using adaptive sampling and mixture-of-experts, achieving lower tracking errors than a re-implemented ExBody2 across diverse whole-body motions.

  6. From Experts to a Generalist: Toward General Whole-Body Control for Humanoid Robots

    cs.RO 2025-06 conditional novelty 6.0 of 10

    BumbleBee, an expert-to-generalist pipeline using autoencoder-based motion clustering and per-cluster delta action models, reports state-of-the-art whole-body control on a Unitree G1 humanoid, with success rates of 89...

  7. SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SkillBlender pretrains reusable goal-conditioned skills and blends them with softmax per-joint weights to solve simulated humanoid loco-manipulation tasks with one or two reward terms.

  8. Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.

  9. HuBE: Cross-Embodiment Human-like Behavior Execution for Humanoid Robots

    cs.RO 2025-08 reject novelty 5.0 of 10

    HuBE is a closed-loop pose-generation framework that produces context-appropriate, human-like upper-body motions for multiple humanoid robots, trained on an LLM-annotated dataset with bone-scaling augmentation.

  10. EMP: Executable Motion Prior for Humanoid Robot Standing Upper-body Motion Imitation

    cs.RO 2025-07 conditional novelty 5.0 of 10

    A state-conditioned executable motion prior network modifies upper-body motion targets so a humanoid can imitate human gestures while maintaining balance.

  11. Realizing Text-Driven Motion Generation on NAO Robot: A Reinforcement Learning-Optimized Control Pipeline

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A text-to-motion diffusion model paired with a norm-position/rotation mapping network and a sim-to-real reinforcement learning controller enables a NAO robot to perform queried gestures such as waving and boxing.

  12. Immersive Social Interaction with VR and LLM-Assisted Humanoids

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Novice operators achieved 80% success on object manipulation and 70% on social cube-passing using a VR-and-LLM-assisted humanoid teleoperation framework on a Unitree H1.

Pith tools