A 269-hour Vietnamese AVSR dataset is collected automatically from YouTube, and an AV-HuBERT baseline degrades far less than audio-only ASR under babble noise.
ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
A 269-hour Vietnamese AVSR dataset is collected automatically from YouTube, and an AV-HuBERT baseline degrades far less than audio-only ASR under babble noise.