SAMoR encodes motions of arbitrary skeletons into a fixed set of 8 part tokens via graph-transformer encoding, cross-attention pooling, and residual vector quantization, enabling cross-topology reconstruction, transfer, and text-conditioned generation.
hub Mixed citations
Human Motion Diffusion Model
Mixed citation behavior. Most common role is background (50%).
abstract
Natural and expressive human motion generation is the holy grail of computer animation. It is a challenging task, due to the diversity of possible motion, human perceptual sensitivity to it, and the difficulty of accurately describing it. Therefore, current generative solutions are either low-quality or limited in expressiveness. Diffusion models, which have already shown remarkable generative capabilities in other domains, are promising candidates for human motion due to their many-to-many nature, but they tend to be resource hungry and hard to control. In this paper, we introduce Motion Diffusion Model (MDM), a carefully adapted classifier-free diffusion-based generative model for the human motion domain. MDM is transformer-based, combining insights from motion generation literature. A notable design-choice is the prediction of the sample, rather than the noise, in each diffusion step. This facilitates the use of established geometric losses on the locations and velocities of the motion, such as the foot contact loss. As we demonstrate, MDM is a generic approach, enabling different modes of conditioning, and different generation tasks. We show that our model is trained with lightweight resources and yet achieves state-of-the-art results on leading benchmarks for text-to-motion and action-to-motion. https://guytevet.github.io/mdm-page/ .
hub tools
citation-role summary
citation-polarity summary
representative citing papers
PatternGSL introduces a learnable specification language for sewing patterns that lets vision-language models reconstruct explicit, simulation-ready 3D garments from single images, backed by a new 300K paired dataset.
GARD is a context-aware autoregressive diffusion model for gloss-wise sign language production using inter-gloss transition guidance and global motion harmonizer, claiming superior linguistic accuracy and motion similarity on Phoenix-T and CSL-Daily datasets.
NextMotionQA benchmark reveals VLMs have critical gaps in fine-grained human motion understanding and align with experts on coarse judgment (κ=0.70) but not fine-grained (κ=0.10).
DrawMotion is a diffusion-based framework that fuses text and hand-drawn stickman conditions via a Multi-Condition Module and training-free guidance to generate 3D human motions.
A hypernetwork maps style motion embeddings to LoRA updates that stylize text-driven motion diffusion models with improved generalization to unseen styles via contrastive structuring of the style space.
ScaleMoGen applies next-scale autoregressive prediction to human motion generation with multi-scale skeletal-temporal bitwise token maps, reporting SOTA FID on HumanML3D.
MaMi-HOI counters geometric forgetting in diffusion models via a Geometry-Aware Proximity Adapter for precise contacts and a Kinematic Harmony Adapter for natural whole-body postures in human-object interactions.
MotionGRPO models diffusion sampling as a Markov decision process optimized with Group Relative Policy Optimization, using hybrid rewards and noise injection to boost sample diversity and local joint precision in egocentric motion recovery.
DanceCrafter generates high-fidelity, text-controlled dance sequences using a new Choreographic Syntax framework and a large fine-grained motion dataset.
TeMuDance enables text-based semantic control over music-conditioned dance generation by using motion as a bridge to align existing unpaired datasets and training a lightweight text branch on a frozen diffusion backbone with noise-filtered supervision.
CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.
FrameCache uses a Screen-Cache-Match strategy and Trajectory-Aware Autoregressive Generation to convert past frames into causal guidance for temporally coherent human animation videos.
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
Ego3DLM jointly predicts past and future 3D body pose and motion descriptions in a single autoregressive pass, conditioned on egocentric video, 3D scene features, and three-point tracking, achieving state-of-the-art on the Nymeria benchmark.
DeSeG decouples semantic intent from geometric constraints in human-scene interaction synthesis using a residual CVAE planner and a physics-regularized diffusion executor, reducing scene penetration by 47% and improving semantic alignment by 29% over SOTA baselines on the Lingo dataset.
A single causal diffusion model with an anchor–relational motion representation generates streaming solo and two-person motion and smooth solo–social transitions from incremental text.
A training-free 3D scene-adaptive human animation framework that controls human motion and camera trajectories via ground-adaptive retargeting and visibility-masked point-cloud fusion in a diffusion backbone.
ICMPG combines LLM-based candidate generation with MPC-style physical simulation and semantic scoring to produce text-driven human motions that are both plausible and faithful.
Proposes a feed-forward keyframe-conditioned in-betweening method for arbitrary 4D meshes using a topology-agnostic VAE and MMDiT-based rectified flow model.
THREAD is a diffusion model for generating feasible backbone trajectories of hybrid rigid-soft manipulators conditioned on environment geometry, reporting 92.4% success and 5x fewer collisions in simulation plus real-world transfer.
MOCHI enhances noisy collaborative human-object interaction captures via grasp optimization followed by diffusion-based full-body refinement that incorporates interaction information into single-person motion priors.
Multimodal framework uses LLM conditioning and a pathology-aware tokeniser to generate synthetic gait sequences that raise GRU classifier accuracy to 92.77% when mixed with real data under leave-one-subject-out evaluation.
A multi-condition latent diffusion model transfers human motion styles to diverse humanoid robot contents with physics regularizations, achieving 96% success in real-robot trials on Unitree G1.
citing papers explorer
-
SAMoR: Motion Modelling for Articulated Objects of Any Skeleton and Topology
SAMoR encodes motions of arbitrary skeletons into a fixed set of 8 part tokens via graph-transformer encoding, cross-attention pooling, and residual vector quantization, enabling cross-topology reconstruction, transfer, and text-conditioned generation.
-
PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments
PatternGSL introduces a learnable specification language for sewing patterns that lets vision-language models reconstruct explicit, simulation-ready 3D garments from single images, backed by a new 300K paired dataset.
-
Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production
GARD is a context-aware autoregressive diffusion model for gloss-wise sign language production using inter-gloss transition guidance and global motion harmonizer, claiming superior linguistic accuracy and motion similarity on Phoenix-T and CSL-Daily datasets.
-
NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models
NextMotionQA benchmark reveals VLMs have critical gaps in fine-grained human motion understanding and align with experts on coarse judgment (κ=0.70) but not fine-grained (κ=0.10).
-
DrawMotion: Generating 3D Human Motions by Freehand Drawing
DrawMotion is a diffusion-based framework that fuses text and hand-drawn stickman conditions via a Multi-Condition Module and training-free guidance to generate 3D human motions.
-
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
A hypernetwork maps style motion embeddings to LoRA updates that stylize text-driven motion diffusion models with improved generalization to unseen styles via contrastive structuring of the style space.
-
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation
ScaleMoGen applies next-scale autoregressive prediction to human motion generation with multi-scale skeletal-temporal bitwise token maps, reporting SOTA FID on HumanML3D.
-
MaMi-HOI: Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation
MaMi-HOI counters geometric forgetting in diffusion models via a Geometry-Aware Proximity Adapter for precise contacts and a Kinematic Harmony Adapter for natural whole-body postures in human-object interactions.
-
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
MotionGRPO models diffusion sampling as a Markov decision process optimized with Group Relative Policy Optimization, using hybrid rewards and noise injection to boost sample diversity and local joint precision in egocentric motion recovery.
-
DanceCrafter: Fine-Grained Text-Driven Controllable Dance Generation via Choreographic Syntax
DanceCrafter generates high-fidelity, text-controlled dance sequences using a new Choreographic Syntax framework and a large fine-grained motion dataset.
-
TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation
TeMuDance enables text-based semantic control over music-conditioned dance generation by using motion as a bridge to align existing unpaired datasets and training a lightweight text branch on a frozen diffusion backbone with noise-filtered supervision.
-
CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos
CoMoVi co-generates 3D human motions and 2D videos synchronously in a single diffusion denoising loop using 3D-to-2D projection and dual-branch diffusion with 3D-2D cross attentions.
-
Screen, Cache, and Match: A Training-Free Causality-Consistent Reference Frame Framework for Human Animation
FrameCache uses a Screen-Cache-Match strategy and Trajectory-Aware Autoregressive Generation to convert past frames into causal guidance for temporally coherent human animation videos.
-
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
-
Ego-Human Motion Prediction with 3D-Aware LLM
Ego3DLM jointly predicts past and future 3D body pose and motion descriptions in a single autoregressive pass, conditioned on egocentric video, 3D scene features, and three-point tracking, achieving state-of-the-art on the Nymeria benchmark.
-
DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction
DeSeG decouples semantic intent from geometric constraints in human-scene interaction synthesis using a residual CVAE planner and a physics-regularized diffusion executor, reducing scene penetration by 47% and improving semantic alignment by 29% over SOTA baselines on the Lingo dataset.
-
ARMS: Anchor-Relational Motion Streaming for Seamless Solo-Social Motion Transitions
A single causal diffusion model with an anchor–relational motion representation generates streaming solo and two-person motion and smooth solo–social transitions from incremental text.
-
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
A training-free 3D scene-adaptive human animation framework that controls human motion and camera trajectories via ground-adaptive retargeting and visibility-masked point-cloud fusion in a diffusion backbone.
-
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
ICMPG combines LLM-based candidate generation with MPC-style physical simulation and semantic scoring to produce text-driven human motions that are both plausible and faithful.
-
Feed-forward Motion In-betweening for Any 4D
Proposes a feed-forward keyframe-conditioned in-betweening method for arbitrary 4D meshes using a topology-agnostic VAE and MMDiT-based rectified flow model.
-
THREAD: Trajectory Planning for Hybrid Rigid-Soft Manipulators with Environment-Aware Diffusion
THREAD is a diffusion model for generating feasible backbone trajectories of hybrid rigid-soft manipulators conditioned on environment geometry, reporting 92.4% success and 5x fewer collisions in simulation plus real-world transfer.
-
MOCHI: Motion Enhancement of Collaborative Human-object Interactions
MOCHI enhances noisy collaborative human-object interaction captures via grasp optimization followed by diffusion-based full-body refinement that incorporates interaction information into single-person motion priors.
-
LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations
Multimodal framework uses LLM conditioning and a pathology-aware tokeniser to generate synthetic gait sequences that raise GRU classifier accuracy to 92.77% when mixed with real data under leave-one-subject-out evaluation.
-
Bionic Human-Motion Style Transfer for Physically Executable Whole-Body Control of Humanoid Robots
A multi-condition latent diffusion model transfers human motion styles to diverse humanoid robot contents with physics regularizations, achieving 96% success in real-robot trials on Unitree G1.
-
Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking From Sparse Inertial Sensors and Ranging-Based Between-Sensor Distances
Ultra Diffusion Poser improves sparse inertial human pose estimation by reconstructing 3D sensor layouts from UWB ranging measurements and using UWB-Diffusion Guidance in a diffusion model, claiming up to 22% lower joint position error than prior work.
-
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
PhyGenHOI couples a motion diffusion model for humans with material point method simulation for objects on 3D Gaussians, using attraction loss, contact re-simulation, and masked video-SDS to produce physically consistent dynamic interactions from text.
-
SCRIPT: Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-based Humanoid Control
SCRIPT presents a scalable diffusion policy with JAST-DiT architecture, nonlinear history conditioning, and RLHR post-training that claims to outperform prior methods on text alignment, motion quality, and physical realism while scaling on a 1200-hour dataset.
-
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing
Recasts sampling-based nonconvex optimization as smoothed gradient descent to obtain non-asymptotic convergence guarantees and introduces the DIDA annealed algorithm that converges to the global optimum.
-
Towards Continuous Sign Language Conversation from Isolated Signs
Constructs continuous sign conversation data from isolated signs using retrieval and diffusion models to train a direct sign-to-sign conversational AI.
-
R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow
R-DMesh proposes a VAE-based disentanglement of base mesh, motion trajectories, and rectification offset plus Triflow Attention and rectified-flow diffusion to produce 4D meshes aligned to video despite initial pose mismatch.
-
PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation
PhysiGen reduces interpenetration in text-driven 3D human interaction generation by simplifying meshes to geometric primitives for fast collision detection and guiding optimization with collision regions.
-
Exploring the Potential of Probabilistic Transformer for Time Series Modeling: A Report on the ST-PT Framework
ST-PT turns transformers into explicit factor graphs for time series, enabling structural injection of symbolic priors, per-sample conditional generation, and principled latent autoregressive forecasting via MFVI iterations.
-
IAM: Identity-Aware Human Motion and Shape Joint Generation
IAM jointly synthesizes motion sequences and body shape parameters conditioned on multimodal identity signals to achieve more realistic and identity-consistent human motions.
-
Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
A flow-matching model derives manipulation strategies from object affordance, adds an adversarial interaction prior, and uses stability simulation to generate natural, effective human-human co-manipulation motions.
-
Visually-grounded Humanoid Agents
A coupled world-agent framework uses 3D Gaussian reconstruction and first-person RGB-D perception with iterative planning to enable goal-directed, collision-avoiding humanoid behavior in novel reconstructed scenes.
-
Next-Scale Autoregressive Models for Text-to-Motion Generation
Next-scale autoregressive modeling with cross-scale and in-scale refinements produces SOTA text-to-motion generation by enforcing coarse-to-fine causal hierarchy.
-
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
LLaMo scales pretrained LLMs for unified motion-language tasks by encoding motion into continuous causal latents and adding a flow-matching head for real-time autoregressive generation and captioning.
-
MARRS: Masked Autoregressive Unit-based Reaction Synthesis
MARRS synthesizes fine-grained reaction motions via unit-distinguished VAE, masked action-conditioned fusion, mutual unit modulation, and compact MLP diffusion predictors.
-
ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments
Dual-expert RL plus difficulty-aware multi-teacher distillation improves physics-based human–scene interaction imitation under complex 3D geometry versus prior single-policy baselines.
-
Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling
Scaling both synthetic motion data volume and VQ-VAE codebook size yields richer motion vocabularies and measurable gains on standard generation benchmarks.
-
Social Structure Matters in 3D Human-Human Interaction Generation
Introduces a Solo-to-Social planner-executor framework where LLMs decompose HHI into phases and roles, then a LoRA-adapted solo motion model grounds them into partner-aware 3D motion.
-
TEXEDO : Test Time Scaling for Controller-aware Language-conditioned Humanoid Motion Generation
TEXEDO uses test-time sampling and a combined feasibility-semantic reward model to select executable, text-aligned motions for humanoid robots from a pretrained generator.
-
OMG: Omni-Modal Motion Generation for Generalist Humanoid Control
OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.
-
Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation
Two-stage pipeline with region-aware attention and Mamba-enhanced diffusion achieves SOTA accuracy, naturalness and temporal coherence on audio-driven portrait animation benchmarks using a new 380-hour dataset.
-
Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization
Sketch2Motion is a diffusion-guided skeleton optimization framework that generates text-driven 3D animations from 2D sketches for biped, quadruped, and other articulated characters.
-
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
Introduces dual pose-image representation, cross-modal alignment, and iterative construction to improve prompt alignment and diversity in multi-person text-to-image generation.
-
Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control
DAJI is a hierarchical framework using distillation and autoregressive generation to learn future-aware joint intents for language-conditioned humanoid robot control.
-
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
KANMultiSign generates sign language poses from notation via coarse-to-fine multi-scale supervision and compact KAN-Transformer modules, achieving lower DTW joint error with fewer parameters than baselines on several language corpora.
-
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction
Uni-HOI learns the joint distribution of text, human motion, and object motion using LLMs and VQ-VAEs in a two-stage training process for multiple HOI tasks.
-
EgoMotion: Hierarchical Reasoning and Diffusion for Egocentric Vision-Language Motion Generation
EgoMotion decouples reasoning from motion synthesis in egocentric vision-language tasks by mapping inputs to motion primitives via VLM then using diffusion to produce grounded and coherent 3D trajectories.