REVIEW 20 cited by
Universal Humanoid Motion Representations for Physics-Based Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a universal motion representation that encompasses a comprehensive range of motor skills for physics-based humanoid control. Due to the high dimensionality of humanoids and the inherent difficulties in reinforcement learning, prior methods have focused on learning skill embeddings for a narrow range of movement styles (e.g. locomotion, game characters) from specialized motion datasets. This limited scope hampers their applicability in complex tasks. We close this gap by significantly increasing the coverage of our motion representation space. To achieve this, we first learn a motion imitator that can imitate all of human motion from a large, unstructured motion dataset. We then create our motion representation by distilling skills directly from the imitator. This is achieved by using an encoder-decoder structure with a variational information bottleneck. Additionally, we jointly learn a prior conditioned on proprioception (humanoid's own pose and velocities) to improve model expressiveness and sampling efficiency for downstream tasks. By sampling from the prior, we can generate long, stable, and diverse human motions. Using this latent space for hierarchical RL, we show that our policies solve tasks using human-like behavior. We demonstrate the effectiveness of our motion representation by solving generative tasks (e.g. strike, terrain traversal) and motion tracking using VR controllers.
Forward citations
Cited by 20 Pith papers
-
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
An autoregressive diffusion model with a hybrid explicit-root/latent-body representation generates real-time, controllable 3D human motion from text and spatial constraints.
-
Flow Matching Policy Gradients
FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.
-
WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation
WristMimic achieves comparable or superior object manipulation retargeting by supervising wrist kinematics while letting finger behavior emerge from object and contact dynamics.
-
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
At compression ratios above 8x, grouping tokens via a Krylov-projected LSH of an implicit attention kernel preserves LLM output quality far better than block averaging, at a large preprocessing latency cost.
-
Distinguishing Imitation Error from Intrinsic Motion Learning Difficulty
A physics-based score (MDS) predicts how hard a motion is for a humanoid to imitate by measuring how much joint torques must change under small pose perturbations.
-
SimGenHOI: Physically Realistic Whole-Body Humanoid-Object Interaction via Generative Modeling and Reinforcement Learning
SimGenHOI generates physically plausible humanoid-object interaction sequences by predicting sparse key actions with a diffusion transformer and tracking them with a contact-aware reinforcement learning policy in simulation.
-
SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
SkillBlender pretrains reusable goal-conditioned skills and blends them with softmax per-joint weights to solve simulated humanoid loco-manipulation tasks with one or two reward terms.
-
PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation
A text-to-motion pipeline that projects motions through physics-based imitation for training and post-processing, plus new consistency and marker-interaction losses.
-
KinTwin: Imitation Learning with Torque and Muscle Driven Biomechanical Models Enables Precise Replication of Able-Bodied and Impaired Movement from Markerless Motion Capture
A single imitation-learning policy tracks the movements of 467 people, including people with gait impairments, with torque- and muscle-driven biomechanical models and estimates joint torques, ground reaction forces, a...
-
Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning
PDC trains a single egocentric-vision policy that lets a simulated humanoid search for, grasp, and place objects and open drawers without privileged state information.
-
DexTrack: Towards Generalizable Neural Tracking Control for Dexterous Manipulation from Human References
A neural controller combining RL and imitation learning on iteratively mined demonstrations tracks human kinematic references for dexterous manipulation, yielding over 10% higher success rates than prior baselines.
-
Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs
A single reinforcement-learned controller uses masked motion demonstrations to catch up, combine, and complete humanoid motions from sparse multi-modal directives.
-
Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking
Mimicking-Bench provides six humanoid-scene interaction tasks with 23K human motion references and a retarget-track-imitate pipeline that beats data-free RL on average success.
-
A Plug-and-Play Physical Motion Restoration Approach for In-the-Wild High-Difficulty Motions
A mask-guided motion correction module plus test-time adapted physics imitation restores physically plausible human motion for high-difficulty in-the-wild videos.
-
RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
Residual-action reinforcement learning, with selective corrections on hip and knee pitch joints, enables zero-shot long-horizon dance tracking on real humanoid robots.
-
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.
-
Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control
A decoupled humanoid controller combines IK-based arm control with an RL locomotion policy conditioned on a CVAE motion prior, improving manipulation precision while maintaining walking stability.
-
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.
-
Feature-Based vs. GAN-Based Learning from Demonstrations: When and Why
Feature-based and GAN-based imitation learning should be selected by task priorities (fidelity, diversity, interpretability, adaptability), not by paradigm loyalty.
-
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.
Discussion (0). Continue with ORCID to comment.