CoGenAV learns audio-visual speech representations that achieve 1.27% WER on LRS2 AVSR and 20.5% WER on LRS2 VSR using 223 hours of labeled data.
Xlavs-r: Cross-lingual audio-visual speech representation learning for noise-robust speech perception, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
CoGenAV learns audio-visual speech representations that achieve 1.27% WER on LRS2 AVSR and 20.5% WER on LRS2 VSR using 223 hours of labeled data.