A movie dubbing system that uses a vision-language model to extract scene type and speaker attributes from silent video, then feeds those attributes as extra conditions into a diffusion-based speech generator, supported by a new 7.2-hour annotated movie dataset.
Datasets Emilia is a comprehensive multilingual speech generation dataset containing a total of 101,654 hours of speech data across six languages [21]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
A movie dubbing system that uses a vision-language model to extract scene type and speaker attributes from silent video, then feeds those attributes as extra conditions into a diffusion-based speech generator, supported by a new 7.2-hour annotated movie dataset.