Pith. sign in

REVIEW 17 cited by

Visual Imitation Enables Contextual Humanoid Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.03729 v5 pith:CP34PCEB submitted 2025-05-06 cs.RO cs.CV

classification cs.ROcs.CV
keywords controlenvironmenthumanoidhumanoidschairscontextualpipelinerobots
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How can we teach humanoids to climb staircases and sit on chairs using the surrounding environment context? Arguably, the simplest way is to just show them-casually capture a human motion video and feed it to humanoids. We introduce VIDEOMIMIC, a real-to-sim-to-real pipeline that mines everyday videos, jointly reconstructs the humans and the environment, and produces whole-body control policies for humanoid robots that perform the corresponding skills. We demonstrate the results of our pipeline on real humanoid robots, showing robust, repeatable contextual control such as staircase ascents and descents, sitting and standing from chairs and benches, as well as other dynamic whole-body skills-all from a single policy, conditioned on the environment and global root commands. VIDEOMIMIC offers a scalable path towards teaching humanoids to operate in diverse real-world environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow Matching Policy Gradients

    cs.LG 2025-07 conditional novelty 7.0 of 10

    FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.

  2. PhiZero: A World Model Built Around Physical Language

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A self-supervised discrete physical-language bottleneck plus a VLM reasoner lets a world model predict state transitions before rendering video, improving physical coherence and enabling zero-shot motion transfer.

  3. Extreme-RGMT: Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A two-stage continual-learning framework lets a generalist humanoid tracking policy acquire highly dynamic acrobatic skills while preserving its general-purpose motion capabilities.

  4. ContactMimic: Humanoid Object Interaction via Contact Control

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A humanoid tracking policy is trained with contact-following rewards and trajectory augmentation to decouple physical contact from keypoint geometry, enabling runtime contact control.

  5. Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A multi-source 16,074-clip quadruped motion library plus a flow-matching generalist tracker shows empirical data scaling and zero-shot unseen tracking, integrated with all-terrain locomotion and real-robot deployment.

  6. World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

    cs.RO 2026-02 conditional novelty 6.0 of 10

    World-VLA-Loop alternately fine-tunes a video world model and a VLA policy, using RL inside the simulator to boost real-world success rates by up to 36.7 percentage points over two iterations.

  7. Thor: Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A decoupled whole-body RL policy with a force-based lean reward enables a Unitree G1 humanoid to pull with up to 167.7 N, beating prior controllers by 69–75%.

  8. PHUMA: Physically Reliable Humanoid Locomotion Dataset

    cs.RO 2025-10 conditional novelty 6.0 of 10

    PHUMA is a curated 73-hour humanoid locomotion corpus whose physical-reliability metrics are partly defined by the same losses used to optimize it, and whose imitation success claims are confounded by in-distribution ...

  9. CBF-RL: Safety Filtering Reinforcement Learning in Training with Control Barrier Functions

    cs.RO 2025-10 conditional novelty 6.0 of 10

    Training RL policies with a closed-form CBF safety filter plus CBF reward lets a Unitree G1 humanoid avoid obstacles and climb stairs without a runtime safety filter.

  10. A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A neural retargeting pipeline maps human motion to humanoid robot motion at 5000+ frames per second using a shared latent space and physics-based fine-tuning, filtering noise and producing physically feasible trajectories.

  11. Robot Drummer: Learning Rhythmic Skills for Humanoid Drumming

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A simulated Unitree G1 humanoid learns to drum dozens of popular songs from MIDI with high F1 scores using a Rhythmic Contact Chain and temporal decomposition.

  12. Crowd4D: Scene-Aware Monocular 4D Crowd Reconstruction

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Crowd4D introduces HSIP scene-anchored proxies and structural coherence regularization to reconstruct scene-consistent 4D crowds from monocular video, outperforming DyCrowd on VirtualCrowd.

  13. ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

    cs.RO 2026-03 conditional novelty 5.0 of 10

    An open-loop generation-then-tracking system maps one egocentric image plus language into Unitree G1 whole-body interactions using only human egocentric motion data.

  14. Consensus-based optimization (CBO): Towards Global Optimality in Robotics

    cs.RO 2026-02 conditional novelty 5.0 of 10

    On long-horizon, underactuated, and high-dimensional simulated robot planning tasks, consensus-based optimization finds lower-cost trajectories than MPPI, CEM, and CMA-ES.

  15. DPL: Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction

    cs.RO 2025-10 conditional novelty 5.0 of 10

    Combining a blind-backbone policy, cross-attention terrain reconstruction from depth plus proprioception, and realistic synthetic depth with noise enables depth-only full-sized humanoid locomotion over stairs, slopes,...

  16. HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    HERMES converts a single human motion demonstration into a deployable mobile bimanual dexterous manipulation policy, using RL, depth-image distillation, and closed-loop PnP pose refinement.

  17. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

Pith tools