Pith. sign in

REVIEW 40 cited by

OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08858 v1 pith:LGUKWDML submitted 2024-06-13 cs.RO cs.CVcs.LGcs.SYeess.SY

classification cs.ROcs.CVcs.LGcs.SYeess.SY
keywords omnih2owhole-bodyhumanoidlearningteleoperationautonomycontroldatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present OmniH2O (Omni Human-to-Humanoid), a learning-based system for whole-body humanoid teleoperation and autonomy. Using kinematic pose as a universal control interface, OmniH2O enables various ways for a human to control a full-sized humanoid with dexterous hands, including using real-time teleoperation through VR headset, verbal instruction, and RGB camera. OmniH2O also enables full autonomy by learning from teleoperated demonstrations or integrating with frontier models such as GPT-4. OmniH2O demonstrates versatility and dexterity in various real-world whole-body tasks through teleoperation or autonomy, such as playing multiple sports, moving and manipulating objects, and interacting with humans. We develop an RL-based sim-to-real pipeline, which involves large-scale retargeting and augmentation of human motion datasets, learning a real-world deployable policy with sparse sensor input by imitating a privileged teacher policy, and reward designs to enhance robustness and stability. We release the first humanoid whole-body control dataset, OmniH2O-6, containing six everyday tasks, and demonstrate humanoid whole-body skill learning from teleoperated datasets.

Discussion (0). Sign in to comment.

Forward citations

Cited by 40 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Adding a panoramic camera feed to a vision-language-action policy raises end-to-end success on four real-world mobile two-arm tasks from 30% to 73%.

  2. Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.

  3. GenTrack: Physical Alignment for Robot-Native Motion Generation and Zero-Shot Humanoid Tracking

    cs.RO 2026-08 conditional novelty 6.0 of 10

    Online co-training of a text-to-motion generator and a humanoid tracker on simulated G1 improves generator executability and zero-shot tracker coverage beyond static replay or one-way filtering.

  4. First Deployable Dynamic-CoM: A Unified Policy and Method-Agnostic Benchmark for Humanoid Single-Leg Balance

    cs.RO 2026-08 conditional novelty 6.0 of 10

    A support-relative dynamic capture-point observation, reconstructible without base linear velocity, lets a humanoid policy hold clean single-leg balance at 86/90 in simulation and deploy on a Unitree G1 without distillation.

  5. DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A hybrid kinesthetic-arm-plus-webcam-hand teleoperation interface achieved 17x/3x higher demonstration throughput than vision baselines and trained a 90%-success pick-and-place policy in a ten-person study.

  6. What Matters in Humanoid General Motion Tracking? An Empirical Study

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A controlled ablation of humanoid motion-tracking pipelines shows that explicit reference joint velocities and a short observation history improve tracking, while residual actions and teacher-student training yield on...

  7. Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A multi-source 16,074-clip quadruped motion library plus a flow-matching generalist tracker shows empirical data scaling and zero-shot unseen tracking, integrated with all-terrain locomotion and real-robot deployment.

  8. WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    WristMimic achieves comparable or superior object manipulation retargeting by supervising wrist kinematics while letting finger behavior emerge from object and contact dynamics.

  9. ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A force-aware humanoid benchmark pairs synchronized human motion-force data with simulation-based force replay to evaluate whole-body control policies under realistic physical disturbances.

  10. Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization

    cs.RO 2026-03 conditional novelty 6.0 of 10

    A physics-aware motion-retargeting pipeline that uses ground-reaction-force-derived heel-toe contacts produces dynamically feasible humanoid references and improves downstream imitation learning.

  11. Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A decoupled whole-body RL policy with a force-based lean reward enables a Unitree G1 humanoid to pull with up to 167.7 N, beating prior controllers by 69–75%.

  12. PHUMA: Physically Reliable Humanoid Locomotion Dataset

    cs.RO 2025-10 conditional novelty 6.0 of 10

    PHUMA is a curated 73-hour humanoid locomotion corpus whose physical-reliability metrics are partly defined by the same losses used to optimize it, and whose imitation success claims are confounded by in-distribution ...

  13. Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A 10,300-demonstration, 260-task multimodal humanoid manipulation dataset with baseline policy evaluations and a cloud evaluation platform.

  14. UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Embodiment-Aware Diffusion Policy steers a UMI-trained diffusion policy with controller tracking-cost gradients at inference time, improving aerial manipulation success in simulation and real flights.

  15. A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A neural retargeting pipeline maps human motion to humanoid robot motion at 5000+ frames per second using a shared latent space and physics-based fine-tuning, filtering noise and producing physically feasible trajectories.

  16. Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Claimed first large-scale egocentric and multi-view dataset of human-object-human assistance (11.4 hours, 1.2M frames) with three benchmarks; only the abstract was assessable because the submitted body text is a diffe...

  17. TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A reinforcement-learned time optimization policy that adaptively slows upper-body motion clips improves stability and precision of humanoid standing manipulation at a modest time cost.

  18. A3FR: Agile 3D Gaussian Splatting with Incremental Gaze Tracked Foveated Rendering in Virtual Reality

    cs.GR 2025-07 conditional novelty 6.0 of 10

    A3FR parallelizes CPU gaze tracking with GPU 3D Gaussian Splatting rendering using incremental early-exit gaze predictions, cutting end-to-end foveated rendering latency by up to 2x without measured quality loss.

  19. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  20. TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A type-guided teleoperation system that selects predefined dexterous hand poses with a language model outperforms retargeting-based teleoperation on nine real-world tasks and improves imitation learning success.

  21. Learning Motion Skills with Adaptive Assistive Curriculum Force in Humanoid Robots

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A2CF uses an adaptive assistive-force agent to guide humanoid robots through training, yielding faster convergence and robust policies that work without the external force.

  22. GMT: General Motion Tracking for Humanoid Whole-Body Control

    cs.RO 2025-06 conditional novelty 6.0 of 10

    GMT trains a single unified humanoid policy using adaptive sampling and mixture-of-experts, achieving lower tracking errors than a re-implemented ExBody2 across diverse whole-body motions.

  23. SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending

    cs.RO 2025-06 conditional novelty 6.0 of 10

    SkillBlender pretrains reusable goal-conditioned skills and blends them with softmax per-joint weights to solve simulated humanoid loco-manipulation tasks with one or two reward terms.

  24. MoRE: Mixture of Residual Experts for Humanoid Lifelike Gaits Learning on Complex Terrains

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A two-stage reinforcement learning pipeline with a mixture of latent residual experts gives a Unitree G1 humanoid multiple commanded human-like gaits over complex terrains.

  25. Understanding and Mitigating Network Latency Effect on Teleoperated-Robot with Extended Reality

    cs.RO 2025-06 reject novelty 6.0 of 10

    TeleXR decouples robot control and XR visualization from network delays by having each side reconstruct the other's state from local sensor data.

  26. Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.

  27. Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CogRobot uses optical flow as an intermediate variable to fine-tune a text-to-video model for predicting bimanual robot trajectories, then maps those predictions to actions with a goal-conditioned diffusion policy.

  28. H2-COMPACT: Human-Humanoid Co-Manipulation via Adaptive Contact Trajectory Policies

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A hierarchical framework that maps wrist force/torque into velocity commands and then into stable leg motions lets a humanoid robot carry loads cooperatively with a human using only haptic cues.

  29. ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments

    cs.CV 2026-07 unverdicted novelty 5.5 of 10

    Dual-expert RL plus difficulty-aware multi-teacher distillation improves physics-based human–scene interaction imitation under complex 3D geometry versus prior single-policy baselines.

  30. ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

    cs.RO 2026-03 conditional novelty 5.0 of 10

    An open-loop generation-then-tracking system maps one egocentric image plus language into Unitree G1 whole-body interactions using only human egocentric motion data.

  31. RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking

    cs.RO 2025-09 conditional novelty 5.0 of 10

    Residual-action reinforcement learning, with selective corrections on hip and knee pitch joints, enables zero-shot long-horizon dance tracking on real humanoid robots.

  32. GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    GBC unifies MoCap retargeting and imitation learning into one framework that trains whole-body humanoid policies across multiple robot morphologies in simulation.

  33. Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.

  34. Hierarchical Vision-Language Planning for Multi-Step Humanoid Manipulation

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A three-layer hierarchical system using a VLM planner and VLM skill monitor with imitation-learned skills and an RL tracking policy achieved 73% success on a real humanoid pick-and-place task.

  35. Whole-Body Conditioned Egocentric Video Prediction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    An autoregressive conditional diffusion transformer predicts future egocentric video from whole-body 3D pose sequences, trained on Nymeria, with atomic action and long-horizon evaluations.

  36. Multi-Loco: Unifying Multi-Embodiment Legged Locomotion via Reinforcement Learning Augmented Diffusion

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A single diffusion-plus-residual-RL policy, trained on four robot morphologies using zero-padded observations and actions, outperforms per-robot PPO baselines in simulation and transfers to real robots.

  37. Multi-Embodiment Robotic Retargeting via Guided Diffusion Model

    cs.RO 2025-05 reject novelty 5.0 of 10

    A graph-conditioned diffusion model retargets motions across heterogeneous robot embodiments without needing target-robot motion data, yet lacks baseline comparisons and error bars in its validation.

  38. From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

    cs.RO 2025-05 reject novelty 4.0 of 10

    A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.

  39. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

  40. Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why

    cs.LG 2025-07 conditional novelty 3.0 of 10

    Feature-based and GAN-based imitation learning should be selected by task priorities (fidelity, diversity, interpretability, adaptability), not by paradigm loyalty.

Pith tools