Co-speech gesture video is generated by first predicting 2D skeleton motion from audio via feature-concatenated diffusion, then rendering with an off-the-shelf video model, supported by a new 405-hour public dataset.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Democratizing High-Fidelity Co-Speech Gesture Video Generation
Co-speech gesture video is generated by first predicting 2D skeleton motion from audio via feature-concatenated diffusion, then rendering with an off-the-shelf video model, supported by a new 405-hour public dataset.