EMO2 generates talking-head videos by first predicting hand poses from audio and then using those hand signals to drive a video diffusion model that synthesizes face and upper-body motion.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
EMO2 generates talking-head videos by first predicting hand poses from audio and then using those hand signals to drive a video diffusion model that synthesizes face and upper-body motion.