Adding multiple codebooks, a KL-divergence regularizer, and cluster-specific codebooks to BEST-RQ improves LibriSpeech ASR word error rates by up to 30.6% relative with faster convergence.
wav2vec 2.0: a framework for self-supervised learning of speech representations,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimized Self-supervised Training with BEST-RQ for Speech Recognition
Adding multiple codebooks, a KL-divergence regularizer, and cluster-specific codebooks to BEST-RQ improves LibriSpeech ASR word error rates by up to 30.6% relative with faster convergence.