HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.
hub
Karen Liu
40 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.
Rhythm transfers interactive whole-body behaviors from simulation to real dual Unitree G1 humanoids via interaction-aware retargeting and graph-reward RL.
A force-aware humanoid benchmark pairs synchronized human motion-force data with simulation-based force replay to evaluate whole-body control policies under realistic physical disturbances.
An auto-regressive diffusion planner trained with scheduled prefix sampling, coupled asynchronously to a pretrained universal tracker, enables closed-loop humanoid whole-body control with replanning under disturbances and zero-shot moving-target reaching.
X-Morph retargets human motions to kinematically plausible references for multiple legged morphologies, trains privileged RL trackers, and distills them into deployable policies that generalize and enable teleoperation and text-conditioned generation.
SceneBot conditions a humanoid tracking policy on motion references and contact labels, using reconstructed scene-interaction data to unify free-space locomotion with contact-rich manipulation and terrain tasks.
PressMimic fuses RGB and pressure for pose estimation via FRAPPE++ and uses pressure signals in RL policy PSP, backed by the MotionPRO dataset, to achieve physically consistent humanoid motion imitation.
TaskNPoint lets humanoid robots learn dynamic skills such as tennis backhands from single short human video demonstrations plus under one hour of single-GPU simulation training, achieving zero-shot generalization to new goal locations without per-task reward tuning.
OpenHLM is an empirical recipe yielding a whole-body humanoid VLA model that outperforms GR00T N1.6 and Ψ0 baselines on long-horizon tasks using less than half the demonstration time.
PhysDrift generates executable humanoid co-speech motions directly from speech via robot-native data curated by IK-EER, claiming better alignment and plausibility than human-centric retargeting.
A humanoid robot can carry out diverse manipulation tasks in a zero-shot way by imitating one AI-generated video, using contact-aware trajectory optimization instead of task-specific policy training.
A three-stage motion-guided curriculum RL framework trains humanoid robots for soccer shooting, achieving 48.6% lower shot error and 2.96x higher velocity than baselines in simulation and sub-meter accuracy with 13.10 m/s ball speed on a real Unitree G1.
MPC-based retargeting framework enables cross-morphology whole-body teleoperation from a single XR device via dynamic feasibility optimization, state synchronization, and SLAM feedback, with reported gains in simulation and real-world tests.
DiscoForcing introduces a causal diffusion-forcing model with a hybrid temporal schedule for stable real-time audio-to-motion generation under abrupt audio changes.
BifrostUMI enables robot-free human demonstration capture via VR and wrist cameras to train visuomotor policies that predict keypoint trajectories for transfer to humanoid whole-body control through retargeting.
ExoActor uses exocentric video generation to implicitly model robot-environment-object interactions and converts the resulting videos into task-conditioned humanoid control sequences.
X2-N is a transformable wheel-legged humanoid robot with a reinforcement learning whole-body controller that enables dual-mode locomotion and manipulation across varied terrains.
A weightlessness mechanism enables humanoid robots to dynamically relax joints for stable, contact-rich motions across diverse environments without task-specific tuning.
CLAW composes motion primitives from a kinematic planner, tracks them with a low-level controller in MuJoCo to produce physically grounded trajectories, and generates segment- and trajectory-level language annotations via templates for scalable motion-language data collection on the Unitree G1.
RoSHI is a hybrid wearable that combines sparse IMUs and egocentric SLAM to capture accurate full-body 3D pose and shape data in natural environments for robot learning.
NMR uses VAE-based clustered expert physics refinement and a CNN-Transformer to learn dynamics-aware retargeting, eliminating joint jumps and self-collisions on Unitree G1 while accelerating downstream control policies.
cuRoboV2 unifies B-spline optimization, GPU-native dense signed distance fields, and scalable whole-body kinematics and dynamics to achieve 99.7% success on payloaded manipulators and 99.6% collision-free IK on 48-DoF humanoids.
HAIC enables robust humanoid interactions with underactuated objects by predicting their dynamics from proprioceptive history and using a world model for adaptive control.
citing papers explorer
-
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
HumanoidArena is a new benchmark of 7 leg-critical HOI/HSI tasks that evaluates egocentric hierarchical whole-body policies in humanoids and finds performance is strongly conditioned on the low-level GMT used.
-
ReActor: Reinforcement Learning for Physics-Aware Motion Retargeting
ReActor jointly optimizes motion retargeting and RL policy training with an approximate gradient to generate physically consistent robot motions from human references using only sparse body correspondences.
-
Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids
Rhythm transfers interactive whole-body behaviors from simulation to real dual Unitree G1 humanoids via interaction-aware retargeting and graph-reward RL.
-
ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations
A force-aware humanoid benchmark pairs synchronized human motion-force data with simulation-based force replay to evaluate whole-body control policies under realistic physical disturbances.
-
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
An auto-regressive diffusion planner trained with scheduled prefix sampling, coupled asynchronously to a pretrained universal tracker, enables closed-loop humanoid whole-body control with replanning under disturbances and zero-shot moving-target reaching.
-
X-Morph: Human Motion Priors for Scalable Robot Learning Across Morphologies
X-Morph retargets human motions to kinematically plausible references for multiple legged morphologies, trains privileged RL trackers, and distills them into deployable policies that generalize and enable teleoperation and text-conditioned generation.
-
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
SceneBot conditions a humanoid tracking policy on motion references and contact labels, using reconstructed scene-interaction data to unify free-space locomotion with contact-rich manipulation and terrain tasks.
-
PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation
PressMimic fuses RGB and pressure for pose estimation via FRAPPE++ and uses pressure signals in RL policy PSP, backed by the MotionPRO dataset, to achieve physically consistent humanoid motion imitation.
-
TaskNPoint: How to Teach Your Humanoid to Hit a Backhand in Minutes
TaskNPoint lets humanoid robots learn dynamic skills such as tennis backhands from single short human video demonstrations plus under one hour of single-GPU simulation training, achieving zero-shot generalization to new goal locations without per-task reward tuning.
-
OpenHLM: An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation
OpenHLM is an empirical recipe yielding a whole-body humanoid VLA model that outperforms GR00T N1.6 and Ψ0 baselines on long-horizon tasks using less than half the demonstration time.
-
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
PhysDrift generates executable humanoid co-speech motions directly from speech via robot-native data curated by IK-EER, claiming better alignment and plausibility than human-centric retargeting.
-
GenHOI: Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training
A humanoid robot can carry out diverse manipulation tasks in a zero-shot way by imitating one AI-generated video, using contact-aware trajectory optimization instead of task-specific policy training.
-
RoboNaldo: Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning
A three-stage motion-guided curriculum RL framework trains humanoid robots for soccer shooting, achieving 48.6% lower shot error and 2.96x higher velocity than baselines in simulation and sub-meter accuracy with 13.10 m/s ball speed on a real Unitree G1.
-
X-OP: Cross-Morphology Whole-Body Teleoperation via MPC Retargeting
MPC-based retargeting framework enables cross-morphology whole-body teleoperation from a single XR device via dynamic feasibility optimization, state synchronization, and SLAM feedback, with reported gains in simulation and real-world tests.
-
DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing
DiscoForcing introduces a causal diffusion-forcing model with a hybrid temporal schedule for stable real-time audio-to-motion generation under abrupt audio changes.
-
BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation
BifrostUMI enables robot-free human demonstration capture via VR and wrist cameras to train visuomotor policies that predict keypoint trajectories for transfer to humanoid whole-body control through retargeting.
-
ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control
ExoActor uses exocentric video generation to implicitly model robot-environment-object interactions and converts the resulting videos into task-conditioned humanoid control sequences.
-
X2-N: A Transformable Wheel-legged Humanoid Robot with Dual-mode Locomotion and Manipulation
X2-N is a transformable wheel-legged humanoid robot with a reinforcement learning whole-body controller that enables dual-mode locomotion and manipulation across varied terrains.
-
Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot
A weightlessness mechanism enables humanoid robots to dynamically relax joints for stable, contact-rich motions across diverse environments without task-specific tuning.
-
CLAW: Composable Language-Annotated Whole-body Motion Generation
CLAW composes motion primitives from a kinematic planner, tracks them with a low-level controller in MuJoCo to produce physically grounded trajectories, and generates segment- and trajectory-level language annotations via templates for scalable motion-language data collection on the Unitree G1.
-
RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild
RoSHI is a hybrid wearable that combines sparse IMUs and egocentric SLAM to capture accurate full-body 3D pose and shape data in natural environments for robot learning.
-
Make Tracking Easy: Neural Motion Retargeting for Humanoid Whole-body Control
NMR uses VAE-based clustered expert physics refinement and a CNN-Transformer to learn dynamics-aware retargeting, eliminating joint jumps and self-collisions on Unitree G1 while accelerating downstream control policies.
-
cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots
cuRoboV2 unifies B-spline optimization, GPU-native dense signed distance fields, and scalable whole-body kinematics and dynamics to achieve 99.7% success on payloaded manipulators and 99.6% collision-free IK on 48-DoF humanoids.
-
HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model
HAIC enables robust humanoid interactions with underactuated objects by predicting their dynamics from proprioceptive history and using a world model for adaptive control.
-
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
Humanoid-LLA converts unconstrained natural language commands into stable whole-body motions for humanoid robots using a unified motion vocabulary and two-stage supervised-plus-reinforcement fine-tuning.
-
Duet: Dual-Robot Understanding via Efficient Teaching
DUET pretrains collaborative policies on human-human VR demonstrations then fine-tunes on minimal robot teleoperation data, achieving equal or better performance than robot-only baselines with 5.4x faster collection across four tasks.
-
Proprioceptive-visual correspondence enables self-other distinction in humanoid robots
Proprioceptive-visual correspondence lets a humanoid robot acquire self-other distinction and a 3D self-model without labels or kinematics.
-
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
-
VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands
VAIC distills a teacher policy into a vision-and-proprioception student policy using recurrent adaptation and decoupled commands, enabling diverse real-robot tasks like box carrying and skateboarding that outperform baselines.
-
OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
OASIS generates scalable simulation data for humanoid loco-manipulation via 3D generative asset reconstruction and domain randomization, yielding a policy with higher zero-shot real-world success than real-robot teleoperation data.
-
T-GMP: Terrain-conditioned Generative Motion Priors for Versatile and Natural Humanoid Locomotion
T-GMP learns a terrain-conditioned latent motion manifold via CVAE from demonstrations and integrates it into an adversarial pipeline with a foothold penalty for versatile, natural humanoid locomotion.
-
HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers
HANDOFF is a distilled mixture-of-experts humanoid whole-body controller that follows a compact task-space interface, matches SOTA velocity tracking, provides large manipulation workspace on Unitree G1, and supports VLM-driven agentic planning with no task-specific data.
-
Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Humanoid-GPT is a causal Transformer pre-trained on a unified billion-scale motion dataset that tracks dynamic behaviors with zero-shot generalization to unseen motions and tasks.
-
Human2Humanoid: Physics-Aware Cross-Morphology Motion Retargeting for Humanoid Robots
Human2Humanoid is an unsupervised motion retargeting framework using CycleGAN, skeleton-aware GCN, end-effector consistency loss, and physics-aware constraints to transfer human motions to humanoid robots without paired data.
-
Constrained Whole-Body Tracking for Humanoid Robots
ConstrainedMimic integrates operational space control and control barrier functions into RL tracking policies to enforce arbitrary runtime constraints on humanoid kinematics and dynamics while preserving contact modes and tracking goals.
-
SPRINT: Efficient Spectral Priors for Humanoid Athletic Sprints
SPRINT generates sprint trajectories for humanoids via spectral priors from five human motion sequences, achieving 6 m/s peak velocity with zero-shot sim-to-real transfer on Unitree G1.
-
Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
Any2Any transfers humanoid whole-body tracking models across embodiments via kinematic alignment followed by targeted PEFT, matching full-training performance with 1% of the data and compute on tested platforms.
-
Booster Lab: A Data-Centric Pipeline for Learning Deployable Humanoid Locomotion Policies
Describes an integrated pipeline for curating motion data, adapting real-to-sim models, applying AMP-based RL, and deploying locomotion policies on Booster T1 and K1 humanoid robots.
- WARP: Whole-Body Retargeting for Learning from Offline Human Demonstrations
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control