SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.
Soulx-podcast: Towards realis- tic long-form podcasts with dialectal and paralinguistic diversity,
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 1polarities
background 1representative citing papers
FlashTTS delivers a streaming TTS system using multi-track input processing and X-pred mean flow matching to reach 325 ms latency in two function evaluations while retaining zero-shot voice cloning.
The paper announces the ISCSLP 2026 CoT-TTS Challenge with text- and audio-context tracks, large-scale bilingual datasets, and a Qwen3-based baseline requiring both reasoning output and speech generation.
citing papers explorer
-
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
SwanVoice is a zero-shot TTS system for 1-4 speakers that reports higher richness and hierarchy scores than open-source baselines on monologue and dialogue tasks via mixed training and DiffusionNFT post-training.
-
FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation
FlashTTS delivers a streaming TTS system using multi-track input processing and X-pred mean flow matching to reach 325 ms latency in two function evaluations while retaining zero-shot voice cloning.
-
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech
The paper announces the ISCSLP 2026 CoT-TTS Challenge with text- and audio-context tracks, large-scale bilingual datasets, and a Qwen3-based baseline requiring both reasoning output and speech generation.
- MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech