Integrating a VAE with an autoregressive prior into token-based spoken language modeling learns continuous variational features that improve naturalness without hand-engineered pitch, with human raters preferring the generated speech.
High fidelity neural audio compression
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
Integrating a VAE with an autoregressive prior into token-based spoken language modeling learns continuous variational features that improve naturalness without hand-engineered pitch, with human raters preferring the generated speech.