Using 80 ms speech segments and 16,384 sound tokens improves zero-shot spoken language understanding and cuts training cost by up to 70%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Using 80 ms speech segments and 16,384 sound tokens improves zero-shot spoken language understanding and cuts training cost by up to 70%.