VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.
Real time speech enhancement in the waveform domain
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
VoiceFormer fuses text, video, and audio in a transformer to separate a target speaker, and stays robust when audio and video are misaligned by up to 200 ms.