Pith. sign in

REVIEW 22 cited by

MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.15001 v1 pith:BPQLSSYO submitted 2022-08-31 cs.CV

MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model

classification cs.CV
keywords motionmotiondiffusegenerationhumanmethodstext-drivendemonstratesdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Human motion modeling is important for many modern graphics applications, which typically require professional skills. In order to remove the skill barriers for laymen, recent motion generation methods can directly generate human motions conditioned on natural languages. However, it remains challenging to achieve diverse and fine-grained motion generation with various text inputs. To address this problem, we propose MotionDiffuse, the first diffusion model-based text-driven motion generation framework, which demonstrates several desired properties over existing methods. 1) Probabilistic Mapping. Instead of a deterministic language-motion mapping, MotionDiffuse generates motions through a series of denoising steps in which variations are injected. 2) Realistic Synthesis. MotionDiffuse excels at modeling complicated data distribution and generating vivid motion sequences. 3) Multi-Level Manipulation. MotionDiffuse responds to fine-grained instructions on body parts, and arbitrary-length motion synthesis with time-varied text prompts. Our experiments show MotionDiffuse outperforms existing SoTA methods by convincing margins on text-driven motion generation and action-conditioned motion generation. A qualitative analysis further demonstrates MotionDiffuse's controllability for comprehensive motion generation. Homepage: https://mingyuan-zhang.github.io/projects/MotionDiffuse.html

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

    cs.CV 2026-04 unverdicted novelty 8.0

    ViPS distills a compact, controllable distribution of valid joint configurations for any auto-rigged mesh from video diffusion priors, matching 4D-trained methods in plausibility while generalizing zero-shot to unseen...

  2. DrawMotion: Generating 3D Human Motions by Freehand Drawing

    cs.CV 2026-05 unverdicted novelty 7.0

    DrawMotion is a diffusion-based framework that fuses text and hand-drawn stickman conditions via a Multi-Condition Module and training-free guidance to generate 3D human motions.

  3. ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    ScaleMoGen introduces a scale-wise autoregressive framework that quantizes motions into hierarchical discrete tokens and predicts next-scale maps to achieve SOTA FID 0.030 on HumanML3D and text-guided editing.

  4. Dynamic Full-body Motion Agent with Object Interaction via Blending Pre-trained Modular Controllers

    cs.CV 2026-05 unverdicted novelty 7.0

    A two-stage framework augments HOI data with dynamic priors and blends pre-trained dynamic motion and static interaction agents via a composer network to enable long-term dynamic human-object interactions with higher ...

  5. DanceCrafter: Fine-Grained Text-Driven Controllable Dance Generation via Choreographic Syntax

    cs.CV 2026-04 unverdicted novelty 7.0

    DanceCrafter generates high-fidelity, text-controlled dance sequences using a new Choreographic Syntax framework and a large fine-grained motion dataset.

  6. ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes

    cs.CV 2026-04 unverdicted novelty 7.0

    ViPS learns a universal, controllable pose space for auto-rigged meshes by transferring motion priors from video diffusion models, matching SOTA performance on plausibility and diversity while enabling zero-shot gener...

  7. ExpertEdit: Learning Skill-Aware Motion Editing from Expert Videos

    cs.CV 2026-04 unverdicted novelty 7.0

    ExpertEdit edits novice motions to expert skill levels by learning a motion prior from unpaired videos and infilling masked skill-critical spans.

  8. Human Motion Diffusion Model

    cs.CV 2022-09 unverdicted novelty 7.0

    MDM is a classifier-free diffusion model that generates expressive human motions by predicting clean samples rather than noise, supporting text and action conditioning and outperforming prior methods on standard benchmarks.

  9. GIRAF: Towards Generalizable Human Interactions with Articulated Objects

    cs.CV 2026-07 conditional novelty 6.0

    A text-conditioned diffusion model using dynamic object-centric BPS, mixed-domain training, and contact augmentation produces generalizable full-body locomotion-to-articulated-object interaction sequences that beat ad...

  10. In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics

    cs.RO 2026-06 unverdicted novelty 6.0

    ICMPG combines LLM-based candidate generation with MPC-style physical simulation and semantic scoring to produce text-driven human motions that are both plausible and faithful.

  11. Feed-forward Motion In-betweening for Any 4D

    cs.CV 2026-06 unverdicted novelty 6.0

    Proposes a feed-forward keyframe-conditioned in-betweening method for arbitrary 4D meshes using a topology-agnostic VAE and MMDiT-based rectified flow model.

  12. EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

    cs.RO 2026-06 unverdicted novelty 6.0

    EgoPriMo learns a unified egocentric motion prior with a Triple-stream DiT model that supports reconstruction, generation, and forecasting of SMPL motions from egocentric views and text, outperforming prior methods an...

  13. Bionic Human-Motion Style Transfer for Physically Executable Whole-Body Control of Humanoid Robots

    cs.RO 2026-06 unverdicted novelty 6.0

    A multi-condition latent diffusion model transfers human motion styles to diverse humanoid robot contents with physics regularizations, achieving 96% success in real-robot trials on Unitree G1.

  14. IAM: Identity-Aware Human Motion and Shape Joint Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    IAM jointly synthesizes motion sequences and body shape parameters conditioned on multimodal identity signals to achieve more realistic and identity-consistent human motions.

  15. HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

    cs.CV 2026-04 unverdicted novelty 6.0

    HO-Flow synthesizes realistic hand-object motions from text and canonical 3D objects via an interaction-aware VAE and masked flow matching, reporting SOTA physical plausibility and diversity on GRAB, OakInk, and DexYCB.

  16. InterCMDM: Block-Causal Diffusion for Autoregressive Human Interaction Generation

    cs.CV 2026-07 unverdicted novelty 5.0

    InterCMDM proposes a block-causal latent diffusion framework with dual-stream causal transformers and multi-task attention masks for autoregressive text-conditioned two-person interaction generation and reports SOTA r...

  17. Zero-Gated Language-conditioned Human Motion Prediction

    cs.CV 2026-06 unverdicted novelty 5.0

    ZGL injects frozen CLIP text embeddings of VLM-generated motion captions into a DCT Transformer via zero-gated adapters and reports lower MPJPE than pose-only baselines on Human3.6M with transfer to CMUMocap.

  18. GPC: Large-Scale Generative Pretraining for Transferable Motor Control

    cs.CV 2026-06 unverdicted novelty 5.0

    GPC learns a motion vocabulary via Finite Scalar Quantization and end-to-end RL, then trains an autoregressive transformer for next-token control generation, achieving 99.98% motion reproduction success with emergent ...

  19. OMG: Omni-Modal Motion Generation for Generalist Humanoid Control

    cs.RO 2026-06 unverdicted novelty 5.0

    OMG is a diffusion model for omni-modal whole-body humanoid motion generation that uses language, audio, and reference motions after large-scale data curation to achieve state-of-the-art performance and adaptation.

  20. Coordinate-Based Dual-Constrained Autoregressive Motion Generation

    cs.CV 2026-04 unverdicted novelty 5.0

    CDAMD is a new autoregressive text-to-motion framework operating on continuous motion coordinates with dual constraints and diffusion-inspired components, establishing new benchmarks and claiming SOTA fidelity plus se...

  21. Evaluating Idle Animation Believability: a User Perspective

    cs.HC 2025-09 unverdicted novelty 4.0

    Users cannot distinguish genuine from acted idle animations but perceive handmade and recorded ones differently; ReActIdle dataset released to simplify future recording.

  22. Agent AI: Surveying the Horizons of Multimodal Interaction

    cs.AI 2024-01 unverdicted novelty 4.0

    The paper defines Agent AI as interactive multimodal systems that perceive grounded data and generate embodied actions, arguing this approach can mitigate hallucinations in foundation models.