A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.