REVIEW 5 cited by
CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deriving co-speech 3D gestures has seen tremendous progress in virtual avatar animation. Yet, the existing methods often produce stiff and unreasonable gestures with unseen human speech inputs due to the limited 3D speech-gesture data. In this paper, we propose CoCoGesture, a novel framework enabling vivid and diverse gesture synthesis from unseen human speech prompts. Our key insight is built upon the custom-designed pretrain-fintune training paradigm. At the pretraining stage, we aim to formulate a large generalizable gesture diffusion model by learning the abundant postures manifold. Therefore, to alleviate the scarcity of 3D data, we first construct a large-scale co-speech 3D gesture dataset containing more than 40M meshed posture instances across 4.3K speakers, dubbed GES-X. Then, we scale up the large unconditional diffusion model to 1B parameters and pre-train it to be our gesture experts. At the finetune stage, we present the audio ControlNet that incorporates the human voice as condition prompts to guide the gesture generation. Here, we construct the audio ControlNet through a trainable copy of our pre-trained diffusion model. Moreover, we design a novel Mixture-of-Gesture-Experts (MoGE) block to adaptively fuse the audio embedding from the human speech and the gesture features from the pre-trained gesture experts with a routing mechanism. Such an effective manner ensures audio embedding is temporal coordinated with motion features while preserving the vivid and diverse gesture generation. Extensive experiments demonstrate that our proposed CoCoGesture outperforms the state-of-the-art methods on the zero-shot speech-to-gesture generation. The dataset will be publicly available at: https://mattie-e.github.io/GES-X/
Forward citations
Cited by 5 Pith papers
-
PhysiInter: Integrating Physical Mapping for High-Fidelity Human Interaction Generation
A text-to-motion pipeline that projects motions through physics-based imitation for training and post-processing, plus new consistency and marker-interaction losses.
-
Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
Co3Gesture generates coherent concurrent co-speech gestures for two speakers using bilateral diffusion branches with temporal interaction and mutual attention, evaluated on the new GES-Inter dataset.
-
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
VividListener uses a Diffusion Transformer with a Responsive Interaction Module and Emotional Intensity Tags to generate controllable listener head dynamics over long multi-turn dialogues, trained on the new ListenerX...
-
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
EMO2 generates talking-head videos by first predicting hand poses from audio and then using those hand signals to drive a video diffusion model that synthesizes face and upper-body motion.
-
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
A comprehensive survey that unifies generative AI techniques for character animation across facial, gesture, motion, and 3D asset generation, with a shared taxonomy and resource list.
Discussion (0). Continue with ORCID to comment.