Using 80 ms speech segments and 16,384 sound tokens improves zero-shot spoken language understanding and cuts training cost by up to 70%.
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, their effect on the performance of SLMs remains unclear. This paper investigates two key aspects of speech tokenization: the segmentation width and the cluster size of discrete units. First, we segment speech signals into fixed/variable widths and pooled representations. We then train K-means models in multiple cluster sizes. Through the evaluation on zero-shot spoken language understanding benchmarks, we find the positive effect of moderately coarse segmentation and bigger cluster size. Notably, among the best-performing models, the most efficient one achieves a 50% reduction in training data and a 70% decrease in training runtime. Our analysis highlights the importance of combining multiple tokens to enhance fine-grained spoken language understanding.
citation-role summary
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Using 80 ms speech segments and 16,384 sound tokens improves zero-shot spoken language understanding and cuts training cost by up to 70%.