A multi-granular VQ-VAE that encodes and reconstructs motion at several time scales, paired with an audio-to-token predictor, lowers FGD and beats EMAGE in perceptual A/B tests on BEAT2 full-body gesture generation.
Robot behavior toolkit: generating effective social behaviors for robots
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.GR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
A multi-granular VQ-VAE that encodes and reconstructs motion at several time scales, paired with an audio-to-token predictor, lowers FGD and beats EMAGE in perceptual A/B tests on BEAT2 full-body gesture generation.