Rhythm transfers interactive whole-body behaviors from simulation to real dual Unitree G1 humanoids via interaction-aware retargeting and graph-reward RL.
hub
Trackanymotionsunderanydisturbances.arXivpreprintarXiv:2509.13833
27 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Switch-JustDance turns Nintendo Switch Just Dance into a reproducible benchmark for evaluating humanoid whole-body controllers on real hardware using the game's scoring system.
A multi-source 16,074-clip quadruped motion library plus a flow-matching generalist tracker shows empirical data scaling and zero-shot unseen tracking, integrated with all-terrain locomotion and real-robot deployment.
Wrist-guided whole-body RL retargets human–object interactions without finger pose supervision, matching supervised methods and generalizing across hand morphologies in simulation.
FADA is a three-stage Planner-IDM method that achieves few-shot domain adaptation for humanoid control by distilling an oracle policy then finetuning only the IDM on short target-domain rollouts via supervised learning.
CWI decouples MoCap data for upper-body manipulation and lower-body locomotion, using dual discriminators and multi-critic training plus distillation to produce a policy that works from hand poses and velocity commands alone.
SceneBot conditions a humanoid tracking policy on motion references and contact labels, using reconstructed scene-interaction data to unify free-space locomotion with contact-rich manipulation and terrain tasks.
PressMimic fuses RGB and pressure for pose estimation via FRAPPE++ and uses pressure signals in RL policy PSP, backed by the MotionPRO dataset, to achieve physically consistent humanoid motion imitation.
Stubborn introduces a unified RL framework with yaw-aligned representation, Bernoulli probabilistic termination, and adaptive sampling for robust humanoid motion tracking and fall recovery.
A data-centric approach shows that less than 3% of AMASS motion data, filtered by physics feasibility, diversity, and complexity, yields better humanoid tracking policies than the full dataset.
A multi-condition latent diffusion model transfers human motion styles to diverse humanoid robot contents with physics regularizations, achieving 96% success in real-robot trials on Unitree G1.
Diff-CAST replaces GAN discriminators with diffusion-based priors and adds symmetric command conditioning plus constrained RL to enable versatile, drift-free, and hardware-safe quadruped locomotion.
SixthSense infers whole-body contact events and wrenches in humanoids from proprioception and IMU data alone by tokenizing histories and estimating a sparse contact-event flow with conditional flow matching.
AssistMimic is the first multi-agent RL method that successfully tracks assistive human-human interaction motions in simulation by using partner-aware policies, single-agent initialization, dynamic reference retargeting, and contact-promoting rewards.
HAIC enables robust humanoid interactions with underactuated objects by predicting their dynamics from proprioceptive history and using a world model for adaptive control.
TeleGate achieves high-precision real-time whole-body teleoperation of humanoid robots by dynamically gating between expert policies and using a VAE motion prior to infer future intent from history, outperforming distillation baselines on dynamic motions with only 2.5 hours of mocap data.
DynaWM adds a world model as dynamics regularizer and momentum targets to teacher-student distillation, yielding better terrain encoding and smoother stair traversal for bipedal-wheeled robots.
HoloAgent-0 is a unified embodied agent framework with Embodied AgentOS, 3D spatial memory, and embodied skills, deployed and evaluated on real robot hardware for navigation and manipulation tasks.
CTS-MoE combines a dense MoE actor with perception-based gating and a multi-critic architecture to enable adaptive perceptive locomotion on discontinuous terrain in a single-stage teacher-student training setup.
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
TAM is a policy-agnostic torque adaptation module trained in randomized simulation that improves zero-shot real-robot performance on dynamic manipulation tasks compared to system identification and RMA baselines.
M3imic unifies heterogeneous motion modalities via encoders into a shared latent space for a single RL-trained whole-body controller achieving high sim success and sim-to-real transfer on Unitree G1.
Humanoid-GPT is a causal Transformer pre-trained on a unified billion-scale motion dataset that tracks dynamic behaviors with zero-shot generalization to unseen motions and tasks.
HoloMotion-1 trains a MoE Transformer policy on hybrid video and MoCap motion data to achieve robust zero-shot tracking that transfers directly to real humanoid robots.
citing papers explorer
-
Rhythm: Learning Interactive Whole-Body Control for Dual Humanoids
Rhythm transfers interactive whole-body behaviors from simulation to real dual Unitree G1 humanoids via interaction-aware retargeting and graph-reward RL.
-
Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game
Switch-JustDance turns Nintendo Switch Just Dance into a reproducible benchmark for evaluating humanoid whole-body controllers on real hardware using the game's scoring system.
-
Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report
A multi-source 16,074-clip quadruped motion library plus a flow-matching generalist tracker shows empirical data scaling and zero-shot unseen tracking, integrated with all-terrain locomotion and real-robot deployment.
-
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
Wrist-guided whole-body RL retargets human–object interactions without finger pose supervision, matching supervised methods and generalizing across hand morphologies in simulation.
-
FADA: Few-Shot Domain Adaptation via Dynamics Alignment for Humanoid Control
FADA is a three-stage Planner-IDM method that achieves few-shot domain adaptation for humanoid control by distilling an oracle policy then finetuning only the IDM on short target-domain rollouts via supervised learning.
-
CWI: Composite Humanoid Whole-Body Imitation System for Loco-manipulation
CWI decouples MoCap data for upper-body manipulation and lower-body locomotion, using dual discriminators and multi-critic training plus distillation to produce a policy that works from hand poses and velocity commands alone.
-
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
SceneBot conditions a humanoid tracking policy on motion references and contact labels, using reconstructed scene-interaction data to unify free-space locomotion with contact-rich manipulation and terrain tasks.
-
PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation
PressMimic fuses RGB and pressure for pose estimation via FRAPPE++ and uses pressure signals in RL policy PSP, backed by the MotionPRO dataset, to achieve physically consistent humanoid motion imitation.
-
Stubborn: A Streamlined and Unified Reinforcement Learning Framework for Robust Motion Tracking and Fall Recovery for Humanoids
Stubborn introduces a unified RL framework with yaw-aligned representation, Bernoulli probabilistic termination, and adaptive sampling for robust humanoid motion tracking and fall recovery.
-
LIMMT: Less is More for Motion Tracking
A data-centric approach shows that less than 3% of AMASS motion data, filtered by physics feasibility, diversity, and complexity, yields better humanoid tracking policies than the full dataset.
-
Bionic Human-Motion Style Transfer for Physically Executable Whole-Body Control of Humanoid Robots
A multi-condition latent diffusion model transfers human motion styles to diverse humanoid robot contents with physics regularizations, achieving 96% success in real-robot trials on Unitree G1.
-
Constraint-Aware Diffusion Priors for High-Fidelity and Versatile Quadruped Locomotion
Diff-CAST replaces GAN discriminators with diffusion-based priors and adds symmetric command conditioning plus constrained RL to enable versatile, drift-free, and hardware-safe quadruped locomotion.
-
SixthSense: Task-Agnostic Proprioception-Only Whole-Body Wrench Estimation for Humanoids
SixthSense infers whole-body contact events and wrenches in humanoids from proprioception and IMU data alone by tokenizing histories and estimating a sparse contact-event flow with conditional flow matching.
-
Learning to Assist: Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning
AssistMimic is the first multi-agent RL method that successfully tracks assistive human-human interaction motions in simulation by using partner-aware policies, single-agent initialization, dynamic reference retargeting, and contact-promoting rewards.
-
HAIC: Humanoid Agile Object Interaction Control via Dynamics-Aware World Model
HAIC enables robust humanoid interactions with underactuated objects by predicting their dynamics from proprioceptive history and using a world model for adaptive control.
-
TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior
TeleGate achieves high-precision real-time whole-body teleoperation of humanoid robots by dynamically gating between expert policies and using a VAE motion prior to infer future intent from history, outperforming distillation baselines on dynamic motions with only 2.5 hours of mocap data.
-
DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs
DynaWM adds a world model as dynamics regularizer and momentum targets to teacher-student distillation, yielding better terrain encoding and smoother stair traversal for bipedal-wheeled robots.
-
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory
HoloAgent-0 is a unified embodied agent framework with Embodied AgentOS, 3D spatial memory, and embodied skills, deployed and evaluated on real robot hardware for navigation and manipulation tasks.
-
CTS-MoE: Implicit Terrain Adaptation via Mixture-of-Experts for Perceptive Locomotion
CTS-MoE combines a dense MoE actor with perception-based gating and a multi-critic architecture to enable adaptive perceptive locomotion on discontinuous terrain in a single-stage teacher-student training setup.
-
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
-
TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation
TAM is a policy-agnostic torque adaptation module trained in randomized simulation that improves zero-shot real-robot performance on dynamic manipulation tasks compared to system identification and RMA baselines.
-
M3imic: Learning a Versatile Whole-Body Controller for Multimodal Motion Mimicking
M3imic unifies heterogeneous motion modalities via encoders into a shared latent space for a single RL-trained whole-body controller achieving high sim success and sim-to-real transfer on Unitree G1.
-
Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Humanoid-GPT is a causal Transformer pre-trained on a unified billion-scale motion dataset that tracks dynamic behaviors with zero-shot generalization to unseen motions and tasks.
-
HoloMotion-1 Technical Report
HoloMotion-1 trains a MoE Transformer policy on hybrid video and MoCap motion data to achieve robust zero-shot tracking that transfers directly to real humanoid robots.
-
Tree Learning: A Multi-Skill Continual Learning Framework for Humanoid Robots
Tree Learning uses root-branch parameter inheritance and multi-modal adaptation to enable continual multi-skill learning in humanoid robots, achieving higher rewards and 100% retention versus joint training in Unity simulations.
-
SplitAdapter: Load-Aware Humanoid Loco-Manipulation via Factorized Adaptation
SplitAdapter factorizes adaptation into load-aware and dynamics-aware encoders using split world-model objectives, GRL regularization, and hierarchical FiLM, reporting higher full-task success than baselines across 2-6 kg masses and 0-60 cm heights.
- SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control