Pith. sign in

REVIEW 1 cited by

M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.08293 v2 pith:EO6SR6QN submitted 2025-05-13 cs.GR cs.AIcs.CVcs.SDeess.AS

M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis

classification cs.GR cs.AIcs.CVcs.SDeess.AS
keywords gesturehumanmulti-granulargesturesmotiontokensaudiofull-body
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and predicting the tokens of each frame from the input audio. However, one observation is that the number of frames required for a complete expressive human gesture, defined as granularity, varies among different human gesture patterns. Existing systems fail to model these gesture patterns due to the fixed granularity of their gesture tokens. To solve this problem, we propose a novel framework named Multi-Granular Gesture Generator (M3G) for audio-driven holistic gesture generation. In M3G, we propose a novel Multi-Granular VQ-VAE (MGVQ-VAE) to tokenize motion patterns and reconstruct motion sequences from different temporal granularities. Subsequently, we proposed a multi-granular token predictor that extracts multi-granular information from audio and predicts the corresponding motion tokens. Then M3G reconstructs the human gestures from the predicted tokens using the MGVQ-VAE. Both objective and subjective experiments demonstrate that our proposed M3G framework outperforms the state-of-the-art methods in terms of generating natural and expressive full-body human gestures.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SentiAvatar: Towards Expressive and Interactive Digital Humans

    cs.CV 2026-04 unverdicted novelty 7.0

    SentiAvatar generates expressive interactive 3D avatars in real time by combining a 37-hour mocap dialogue dataset with a pre-trained motion foundation model and an audio-aware plan-then-infill architecture that separ...