A two-stage key-value memory model generates personalized 3D facial animation from audio alone, topping prior methods on FVE, LVE, FID, LDTW, and lip-max on VOCASET and BIWI.
Facetalk: Audio-driven motion diffusion for neural parametric head models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
A two-stage key-value memory model generates personalized 3D facial animation from audio alone, topping prior methods on FVE, LVE, FID, LDTW, and lip-max on VOCASET and BIWI.