REVIEW 6 cited by
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
read the original abstract
Recently, human motion analysis has experienced great improvement due to inspiring generative models such as the denoising diffusion model and large language model. While the existing approaches mainly focus on generating motions with textual descriptions and overlook the reciprocal task. In this paper, we present~\textbf{MoTe}, a unified multi-modal model that could handle diverse tasks by learning the marginal, conditional, and joint distributions of motion and text simultaneously. MoTe enables us to handle the paired text-motion generation, motion captioning, and text-driven motion generation by simply modifying the input context. Specifically, MoTe is composed of three components: Motion Encoder-Decoder (MED), Text Encoder-Decoder (TED), and Moti-on-Text Diffusion Model (MTDM). In particular, MED and TED are trained for extracting latent embeddings, and subsequently reconstructing the motion sequences and textual descriptions from the extracted embeddings, respectively. MTDM, on the other hand, performs an iterative denoising process on the input context to handle diverse tasks. Experimental results on the benchmark datasets demonstrate the superior performance of our proposed method on text-to-motion generation and competitive performance on motion captioning.
Forward citations
Cited by 6 Pith papers
-
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
At the critical temperature of the 2D Potts model (q>4), a disordered layer emerges between two ordered phases, and its boundaries converge to a Brownian watermelon under diffusive scaling.
-
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
SMD converts human motion data into structured text descriptions, enabling LLMs to reach new state-of-the-art results on motion question answering and captioning without learned encoders.
-
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
ViBES introduces a speech-language-behavior model using modality-specific transformer experts that jointly generates dialogue and 3D body actions, showing gains over separate co-speech and text-to-motion baselines on ...
-
MUGEN: A Unified Framework for Efficient Motion Understanding and Generation
MUGEN lets a language model generate and read motion through a few continuous latent slots, achieving competitive retrieval and captioning with one draw and K language-model steps.
-
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
LLaMo scales pretrained LLMs for unified motion-language tasks by encoding motion into continuous causal latents and adding a flow-matching head for real-time autoregressive generation and captioning.
-
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.