Pith. sign in

REVIEW 6 cited by

MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.19786 v1 pith:5IJBCQQO submitted 2024-11-29 cs.CV cs.CLcs.LG

MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks

classification cs.CV cs.CLcs.LG
keywords motionmodelgenerationmotediffusionhandletaskscaptioning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recently, human motion analysis has experienced great improvement due to inspiring generative models such as the denoising diffusion model and large language model. While the existing approaches mainly focus on generating motions with textual descriptions and overlook the reciprocal task. In this paper, we present~\textbf{MoTe}, a unified multi-modal model that could handle diverse tasks by learning the marginal, conditional, and joint distributions of motion and text simultaneously. MoTe enables us to handle the paired text-motion generation, motion captioning, and text-driven motion generation by simply modifying the input context. Specifically, MoTe is composed of three components: Motion Encoder-Decoder (MED), Text Encoder-Decoder (TED), and Moti-on-Text Diffusion Model (MTDM). In particular, MED and TED are trained for extracting latent embeddings, and subsequently reconstructing the motion sequences and textual descriptions from the extracted embeddings, respectively. MTDM, on the other hand, performs an iterative denoising process on the input context to handle diverse tasks. Experimental results on the benchmark datasets demonstrate the superior performance of our proposed method on text-to-motion generation and competitive performance on motion captioning.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Encoder-Free Human Motion Understanding via Structured Motion Descriptions

    cs.CV 2026-04 accept novelty 8.0

    At the critical temperature of the 2D Potts model (q>4), a disordered layer emerges between two ordered phases, and its boundaries converge to a Brownian watermelon under diffusive scaling.

  2. Encoder-Free Human Motion Understanding via Structured Motion Descriptions

    cs.CV 2026-04 unverdicted novelty 7.0

    SMD converts human motion data into structured text descriptions, enabling LLMs to reach new state-of-the-art results on motion question answering and captioning without learned encoders.

  3. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

    cs.CV 2025-12 unverdicted novelty 7.0

    ViBES introduces a speech-language-behavior model using modality-specific transformer experts that jointly generates dialogue and 3D body actions, showing gains over separate co-speech and text-to-motion baselines on ...

  4. MUGEN: A Unified Framework for Efficient Motion Understanding and Generation

    cs.LG 2026-07 conditional novelty 6.0

    MUGEN lets a language model generate and read motion through a few continuous latent slots, achieving competitive retrieval and captioning with one draw and K language-model steps.

  5. LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

    cs.CV 2026-02 unverdicted novelty 6.0

    LLaMo scales pretrained LLMs for unified motion-language tasks by encoding motion into continuous causal latents and adding a flow-matching head for real-time autoregressive generation and captioning.

  6. IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

    cs.CV 2025-12 conditional novelty 6.0

    Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.