VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.
The conversation: Deep audio-visual speech enhance- ment
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
baseline 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
baseline 1polarities
baseline 1representative citing papers
citing papers explorer
-
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.