A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.
Knowing who to listen to in speech recognition: Visually guided beamforming,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Cocktail-Party Audio-Visual Speech Recognition
A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.