Pith. sign in

REVIEW 5 cited by

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.06602 v3 pith:5YBVB564 submitted 2024-12-09 cs.CL cs.AIcs.LGcs.MMcs.SDeess.AS

classification cs.CLcs.AIcs.LGcs.MMcs.SDeess.AS
keywords controllablecontrollanguagesurveycomprehensivelargemodelsspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area. This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts. We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS. This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field. One can visit https://github.com/imxtx/awesome-controllabe-speech-synthesis for a comprehensive paper list and updates.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation

    cs.SD 2026-04 unverdicted novelty 7.0 of 10

    CapTalk unifies single-utterance and dialogue voice design via utterance- and speaker-level captions plus a hierarchical variational module for stable timbre with adaptive expression.

  2. TokenChain: A Discrete Speech Chain via Semantic Token Modeling

    eess.AS 2025-10 unverdicted novelty 7.0 of 10

    TokenChain demonstrates that a discrete semantic-token interface can sustain effective chain learning between ASR and TTS, yielding faster convergence and lower error rates on LibriSpeech and TED-LIUM.

  3. FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    FC-TTS presents a zero-shot TTS framework that integrates disentangled speech representations with architectural choices, training framework, and auxiliary objectives to enable independent style and timbre control fro...

  4. Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models

    cs.SD 2025-09 conditional novelty 6.0 of 10

    Speech DF Arena standardizes audio deepfake detection benchmarking across 14 datasets and 15 systems, showing that most open-source detectors have high error rates on out-of-domain attacks.

  5. Position: Towards Responsible Evaluation for Text-to-Speech

    eess.AS 2025-10 conditional novelty 5.0 of 10

    A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.

Pith tools