WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.
Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
eess.AS 2representative citing papers
Introduces the E-VOC corpus and shows that five ITTS systems, including gpt-4o-mini-tts as the best, still default to adult voices and struggle with fine-grained expressive control.
citing papers explorer
-
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.
-
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
Introduces the E-VOC corpus and shows that five ITTS systems, including gpt-4o-mini-tts as the best, still default to adult voices and struggle with fine-grained expressive control.