A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.
The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cocktail-Party Audio-Visual Speech Recognition
A model fine-tuned on a new silent-face-aware audio-visual dataset achieves 39.2% WER in extreme cocktail-party noise, a 67% relative reduction over the prior baseline's 119%.