A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.
S.; Nagrani, A.; and Zisserman, A
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MuteSwap: Visual-informed Silent Video Identity Conversion
A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.