A 269-hour Vietnamese AVSR dataset is collected automatically from YouTube, and an AV-HuBERT baseline degrades far less than audio-only ASR under babble noise.
The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
A 269-hour Vietnamese AVSR dataset is collected automatically from YouTube, and an AV-HuBERT baseline degrades far less than audio-only ASR under babble noise.