WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.
Emosphere++: Emotion-controllable zero- shot text-to-speech via emotion-adaptive spherical vector
2 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
An emotion TTS system adjusts Classifier-Free Guidance strength according to text-style semantic mismatch; it shows small emotion-accuracy gains, but headline baselines and subjective results are absent from the main text.
citing papers explorer
-
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.
-
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
An emotion TTS system adjusts Classifier-Free Guidance strength according to text-style semantic mismatch; it shows small emotion-accuracy gains, but headline baselines and subjective results are absent from the main text.