Pith. sign in

REVIEW 1 cited by

M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.14502 v1 pith:NXTWIGK2 submitted 2024-07-19 cs.CV

classification cs.CV
keywords m2d2mdiffusiondiscretegenerationmotionmodelsmulti-motionactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Multi-Motion Discrete Diffusion Models (M2D2M), a novel approach for human motion generation from textual descriptions of multiple actions, utilizing the strengths of discrete diffusion models. This approach adeptly addresses the challenge of generating multi-motion sequences, ensuring seamless transitions of motions and coherence across a series of actions. The strength of M2D2M lies in its dynamic transition probability within the discrete diffusion model, which adapts transition probabilities based on the proximity between motion tokens, encouraging mixing between different modes. Complemented by a two-phase sampling strategy that includes independent and joint denoising steps, M2D2M effectively generates long-term, smooth, and contextually coherent human motion sequences, utilizing a model trained for single-motion generation. Extensive experiments demonstrate that M2D2M surpasses current state-of-the-art benchmarks for motion generation from text descriptions, showcasing its efficacy in interpreting language semantics and generating dynamic, realistic motions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single language model takes speech, text, or motion tokens as input and generates expressive 3D body motion, text, or emotion labels, using pretraining on unpaired audio and motion data.

Pith tools