DQ-Data2vec adds online K-means quantizers with cluster counts matched to language and phoneme counts to decouple these features during multilingual speech pre-training, improving phoneme and word error rates on CommonVoice.
Self- supervised learning with random-projection quantizer for speech recognition,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
DQ-Data2vec adds online K-means quantizers with cluster counts matched to language and phoneme counts to decouple these features during multilingual speech pre-training, improving phoneme and word error rates on CommonVoice.