A two-stage VQ-VAE and crossmodal transformer with coherence and relevance losses produces semantically aware co-speech gestures, beating four baselines on BEAT and TED Expressive for FGD, diversity, and SRGR.
Gesturedif- fuclip: Gesture diffusion model with clip latents
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning
A two-stage VQ-VAE and crossmodal transformer with coherence and relevance losses produces semantically aware co-speech gestures, beating four baselines on BEAT and TED Expressive for FGD, diversity, and SRGR.