A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.
We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cocktail-Party Audio-Visual Speech Recognition
A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.