A movie dubbing system that uses a vision-language model to extract scene type and speaker attributes from silent video, then feeds those attributes as extra conditions into a diffusion-based speech generator, supported by a new 7.2-hour annotated movie dataset.
More than words: In-the-wild visually-driven prosody for text-to-speech,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
A movie dubbing system that uses a vision-language model to extract scene type and speaker attributes from silent video, then feeds those attributes as extra conditions into a diffusion-based speech generator, supported by a new 7.2-hour annotated movie dataset.