Concatenating speech-recognition and active-speaker visual embeddings improves audio-visual speech enhancement in low-SNR multi-speaker settings, and the CPU real-time system is released open source.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
Concatenating speech-recognition and active-speaker visual embeddings improves audio-visual speech enhancement in low-SNR multi-speaker settings, and the CPU real-time system is released open source.