Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

· 2026 · cs.SD · arXiv 2606.30671

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates pseudolabels using a fixed online quantization scheme, which simplifies training but provides weaker supervision than HuBERT-style models that iteratively refine pseudo-labels. In this work, we improve online pseudo-label generation while preserving simplicity. We propose three modifications: replacing the quantizer's linear projection with Principal Component Analysis (PCA), updating the codebook via iterative codebook refinement, and introducing an additional codebook updated via codebook distillation. We pre-train on the LibriSpeech 960-hour dataset and fine-tune using 100 hours of supervised LibriSpeech data. With all three modifications enabled, we achieve a 12% relative reduction in word error rate (WER) on the LibriSpeech test-other set, improving from 10.1% to 8.8%.

representative citing papers

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

cs.SD · 2026-06-24 · unverdicted · novelty 4.0

Three modifications to BEST-RQ quantization (PCA projection, iterative codebook refinement, codebook distillation) reduce WER from 10.1% to 8.8% on LibriSpeech test-other.

citing papers explorer

Showing 1 of 1 citing paper.

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition cs.SD · 2026-06-24 · unverdicted · none · ref 2 · internal anchor
Three modifications to BEST-RQ quantization (PCA projection, iterative codebook refinement, codebook distillation) reduce WER from 10.1% to 8.8% on LibriSpeech test-other.

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

fields

years

verdicts

representative citing papers

citing papers explorer