A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.
Datasets LRS2-BBC[25] is an English audio-visual dataset from BBC programs, comprising over 220 hours of video
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling
A lip-to-speech model that predicts prosody from an audio prompt and content from lip-reading, then fuses them to synthesize speech, achieving strong benchmark results on LRS2 and LRS3.