WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.
Deep dubbing: End-to-end auto-audiobook system with text-to-timbre and context-aware instruct-tts,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
WordVoice achieves explicit, decoupled word-level control over five acoustic dimensions in LLM-based TTS via a bound-token acoustic planning mechanism and fine-grained style modulation, supported by a new 4.7k-hour bilingual dataset.