CuteTTS combines a semantically aligned causal VAE, patch-level autoregression, and guidance-step distillation to deliver efficient zero-shot voice cloning in a 0.2B-parameter streaming system.
Advances in Neural Information Processing Systems , volume =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
CuteTTS combines a semantically aligned causal VAE, patch-level autoregression, and guidance-step distillation to deliver efficient zero-shot voice cloning in a 0.2B-parameter streaming system.