WAND converts full-attention AR-TTS models to global-plus-sliding-window attention with curriculum fine-tuning and teacher distillation, claiming quality preservation with up to 66.2% KV-cache savings and near-constant latency.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
WAND converts full-attention AR-TTS models to global-plus-sliding-window attention with curriculum fine-tuning and teacher distillation, claiming quality preservation with up to 66.2% KV-cache savings and near-constant latency.