Applying differentiable k-means to token-based ASR lets the tokenizer, the speech encoder, and the recognizer be optimized together, improving WER and making tokens more phoneme-like.
Experimental Setup Our experiments were conducted using ESPnet [24], follow- ing the basic configurations described in [11, 12]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Differentiable K-means for Fully-optimized Discrete Token-based ASR
Applying differentiable k-means to token-based ASR lets the tokenizer, the speech encoder, and the recognizer be optimized together, improving WER and making tokens more phoneme-like.