Pith. sign in

REVIEW 3 cited by

Text-driven Motion Generation: Overview, Challenges and Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.09379 v1 pith:GD2I4CI4 submitted 2025-05-14 cs.CV

classification cs.CV
keywords motiongenerationchallengesdirectionsfuturehumanmethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling animated characters. This makes it especially useful in areas like virtual reality, gaming, human-computer interaction, and robotics. In this review, we first revisit the traditional perspective on motion synthesis, where models focused on predicting future poses from observed initial sequences, often conditioned on action labels. We then provide a comprehensive and structured survey of modern text-to-motion generation approaches, categorizing them from two complementary perspectives: (i) architectural, dividing methods into VAE-based, diffusion-based, and hybrid models; and (ii) motion representation, distinguishing between discrete and continuous motion generation strategies. In addition, we explore the most widely used datasets, evaluation methods, and recent benchmarks that have shaped progress in this area. With this survey, we aim to capture where the field currently stands, bring attention to its key challenges and limitations, and highlight promising directions for future exploration. We hope this work offers a valuable starting point for researchers and practitioners working to push the boundaries of language-driven human motion synthesis.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ARMS: Anchor-Relational Motion Streaming for Seamless Solo-Social Motion Transitions

    cs.CV 2026-07 accept novelty 6.0 of 10

    A single causal diffusion model with an anchor–relational motion representation generates streaming solo and two-person motion and smooth solo–social transitions from incremental text.

  2. Social Structure Matters in 3D Human-Human Interaction Generation

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Introduces a Solo-to-Social planner-executor framework where LLMs decompose HHI into phases and roles, then a LoRA-adapted solo motion model grounds them into partner-aware 3D motion.

  3. Coordinate-Based Dual-Constrained Autoregressive Motion Generation

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    CDAMD is a new autoregressive text-to-motion framework operating on continuous motion coordinates with dual constraints and diffusion-inspired components, establishing new benchmarks and claiming SOTA fidelity plus se...

Pith tools