A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audio models on a new benchmark.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
A single multimodal language model can generate intelligible speech and synchronized background audio jointly from a silent video, transcript, and reference voice, outperforming separately concatenated speech and audio models on a new benchmark.