An autoregressive diffusion model with a hybrid explicit-root/latent-body representation generates real-time, controllable 3D human motion from text and spatial constraints.
hub
arXiv preprint arXiv:2603.15546 , year=
19 Pith papers cite this work. Polarity classification is still indexing.
hub tools
years
2026 19representative citing papers
Releases body-pose dataset for communicative intent recognition and introduces consistency-based reliability measure with a proof bounding its accuracy probability for real-time robot applications.
DanceCrafter generates high-fidelity, text-controlled dance sequences using a new Choreographic Syntax framework and a large fine-grained motion dataset.
Generates 48,000 synthetic VLK trajectories in 3D-reconstructed scenes to train a policy for egocentric perception-based humanoid navigation and object transport, shown on physical Unitree G1 robot.
An auto-regressive diffusion planner trained with scheduled prefix sampling, coupled asynchronously to a pretrained universal tracker, enables closed-loop humanoid whole-body control with replanning under disturbances and zero-shot moving-target reaching.
X-Morph retargets human motions to kinematically plausible references for multiple legged morphologies, trains privileged RL trackers, and distills them into deployable policies that generalize and enable teleoperation and text-conditioned generation.
T2Mo generates controllable dynamic 3D shapes by conditioning on both text semantics and 3D trajectories with a shape-grounded embedding for arbitrary inputs.
Text2BFM aligns language with a frozen BFM via a text-aligned variational behavioral bottleneck to generate long motions by decoding latents into policy actions.
AnyMo is a masked-modeling framework for any-modality human motion generation trained on the new OmniHuMo dataset of 5,000+ hours of multimodal motion sequences.
MIND introduces a multi-scale intent diffusion framework that bridges text commands and low-level actions via holistic and immediate intent predictors for physics-based humanoid control.
SCRIPT presents a scalable diffusion policy with JAST-DiT architecture, nonlinear history conditioning, and RLHR post-training that claims to outperform prior methods on text alignment, motion quality, and physical realism while scaling on a 1200-hour dataset.
AnchorRoute couples anchor-conditioned generation via AnchorKV on a frozen text-to-motion diffusion prior with residual-routed refinement through RouteSolver on piecewise-affine interval bases.
CLAW composes motion primitives from a kinematic planner, tracks them with a low-level controller in MuJoCo to produce physically grounded trajectories, and generates segment- and trajectory-level language annotations via templates for scalable motion-language data collection on the Unitree G1.
Humanoid-DART bootstraps goal-conditioned humanoid loco-manipulation policies from sparse demonstrations via diffusion trajectory generation plus RL tracking.
TEXEDO uses test-time sampling and a combined feasibility-semantic reward model to select executable, text-aligned motions for humanoid robots from a pretrained generator.
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
VAIC distills a teacher policy into a vision-and-proprioception student policy using recurrent adaptation and decoupled commands, enabling diverse real-robot tasks like box carrying and skateboarding that outperform baselines.
UMo presents a sparse MoE-based unified model for real-time co-speech avatar animation that claims superior quality under latency constraints via keyframe-centric design and multi-stage audio-augmented training.
citing papers explorer
-
ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation
An autoregressive diffusion model with a hybrid explicit-root/latent-body representation generates real-time, controllable 3D human motion from text and spatial constraints.
-
Real-time body pose non-verbal communication with a consistency-based reliability measure
Releases body-pose dataset for communicative intent recognition and introduces consistency-based reliability measure with a proof bounding its accuracy probability for real-time robot applications.
-
DanceCrafter: Fine-Grained Text-Driven Controllable Dance Generation via Choreographic Syntax
DanceCrafter generates high-fidelity, text-controlled dance sequences using a new Choreographic Syntax framework and a large fine-grained motion dataset.
-
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Generates 48,000 synthetic VLK trajectories in 3D-reconstructed scenes to train a policy for egocentric perception-based humanoid navigation and object transport, shown on physical Unitree G1 robot.
-
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
An auto-regressive diffusion planner trained with scheduled prefix sampling, coupled asynchronously to a pretrained universal tracker, enables closed-loop humanoid whole-body control with replanning under disturbances and zero-shot moving-target reaching.
-
X-Morph: Human Motion Priors for Scalable Robot Learning Across Morphologies
X-Morph retargets human motions to kinematically plausible references for multiple legged morphologies, trains privileged RL trackers, and distills them into deployable policies that generalize and enable teleoperation and text-conditioned generation.
-
Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text
T2Mo generates controllable dynamic 3D shapes by conditioning on both text semantics and 3D trajectories with a shape-grounded embedding for arbitrary inputs.
-
Plan, Don't Pose: Long Composite Motion Generation with Text-Aligned BFM
Text2BFM aligns language with a frozen BFM via a text-aligned variational behavioral bottleneck to generate long motions by decoding latents into policy actions.
-
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
AnyMo is a masked-modeling framework for any-modality human motion generation trained on the new OmniHuMo dataset of 5,000+ hours of multimodal motion sequences.
-
MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
MIND introduces a multi-scale intent diffusion framework that bridges text commands and low-level actions via holistic and immediate intent predictors for physics-based humanoid control.
-
SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control
SCRIPT presents a scalable diffusion policy with JAST-DiT architecture, nonlinear history conditioning, and RLHR post-training that claims to outperform prior methods on text alignment, motion quality, and physical realism while scaling on a 1200-hour dataset.
-
AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro
AnchorRoute couples anchor-conditioned generation via AnchorKV on a frozen text-to-motion diffusion prior with residual-routed refinement through RouteSolver on piecewise-affine interval bases.
-
CLAW: Composable Language-Annotated Whole-body Motion Generation
CLAW composes motion primitives from a kinematic planner, tracks them with a low-level controller in MuJoCo to produce physically grounded trajectories, and generates segment- and trajectory-level language annotations via templates for scalable motion-language data collection on the Unitree G1.
-
Humanoid-DART: Humanoid Loco-Manipulation using Diffusion-guided Augmentation through Relabeling and Tracking
Humanoid-DART bootstraps goal-conditioned humanoid loco-manipulation policies from sparse demonstrations via diffusion trajectory generation plus RL tracking.
-
TEXEDO : Test Time Scaling for Controller-aware Language-conditioned Humanoid Motion Generation
TEXEDO uses test-time sampling and a combined feasibility-semantic reward model to select executable, text-aligned motions for humanoid robots from a pretrained generator.
-
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
-
VAIC: Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands
VAIC distills a teacher policy into a vision-and-proprioception student policy using recurrent adaptation and decoupled commands, enabling diverse real-robot tasks like box carrying and skateboarding that outperform baselines.
-
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
UMo presents a sparse MoE-based unified model for real-time co-speech avatar animation that claims superior quality under latency constraints via keyframe-centric design and multi-stage audio-augmented training.
- AnyAct: Towards Human Reenactment of Character Motion From Video