VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.
Looking into your speech: Learning cross-modal affinity for audio-visual speech sepa- ration
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.