M2CI-Dubber improves dubbing prosody by extracting global sentence-level and local phoneme-level features from multimodal context and fusing them with the current text through attention and graph interaction.
V2c: Visual voice cloning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
M2CI-Dubber improves dubbing prosody by extracting global sentence-level and local phoneme-level features from multimodal context and fusing them with the current text through attention and graph interaction.