Pith. sign in

REVIEW 15 cited by

PromptTTS 2: Describing and Generating Voices with Text Prompt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.02285 v2 pith:4PZ3W27V submitted 2023-09-05 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords textpromptsspeechpromptvoicevariabilitygenerationinformation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice variability, using text prompts (descriptions) is more user-friendly since speech prompts can be hard to find or may not exist at all. TTS approaches based on the text prompt face two main challenges: 1) the one-to-many problem, where not all details about voice variability can be described in the text prompt, and 2) the limited availability of text prompt datasets, where vendors and large cost of data labeling are required to write text prompts for speech. In this work, we introduce PromptTTS 2 to address these challenges with a variation network to provide variability information of voice not captured by text prompts, and a prompt generation pipeline to utilize the large language models (LLM) to compose high quality text prompts. Specifically, the variation network predicts the representation extracted from the reference speech (which contains full information about voice variability) based on the text prompt representation. For the prompt generation pipeline, it generates text prompts for speech with a speech language understanding model to recognize voice attributes (e.g., gender, speed) from speech and a large language model to formulate text prompts based on the recognition results. Experiments on a large-scale (44K hours) speech dataset demonstrate that compared to the previous works, PromptTTS 2 generates voices more consistent with text prompts and supports the sampling of diverse voice variability, thereby offering users more choices on voice generation. Additionally, the prompt generation pipeline produces high-quality text prompts, eliminating the large labeling cost. The demo page of PromptTTS 2 is available online.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Direction-vector style interpolation plus KV-cache swap and sliding-window masking unlock continuous inter- and intra-utterance style control in prompt-based autoregressive TTS without training.

  2. Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.

  3. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  4. Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.

  5. InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems

    cs.CL 2025-06 conditional novelty 6.0 of 10

    InstructTTSEval introduces three tasks and 6,000 English and Chinese test cases for judging how well TTS systems follow natural-language style instructions, using Gemini as the evaluator.

  6. Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

    cs.MM 2025-06 conditional novelty 6.0 of 10

    A two-stage language-model TTS system uses quantized masked-autoencoder style tokens plus discrete attribute labels to achieve fine-grained style control with stable content.

  7. VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.

  8. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

    cs.SD 2026-07 unverdicted novelty 5.0 of 10

    AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.

  9. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  10. MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A multi-modal emotion prompt encoder and prosody predictor let MPE-TTS control emotion from speech, text, or image while preserving speaker timbre.

  11. MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A zero-shot pipeline that creates character voices from AI-generated faces and LLM-written prosody instructions can produce expressive audiobooks without extra training or manual annotation, though human quality score...

  12. Gender Bias in Instruction-Guided Speech Synthesis Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Parler-TTS speech models produce gender-stereotyped voices for occupation prompts, and prompt-based mitigation is inconsistent and can reverse the bias.

  13. CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

    eess.AS 2026-08 conditional novelty 4.0 of 10

    CtrlSpeech adds phone-level pitch, loudness, and duration controls to a diffusion-transformer TTS backbone, improving local prosody control while preserving zero-shot speaker similarity.

  14. Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

    cs.SD 2024-12 conditional novelty 4.0 of 10

    A speech-to-speech dialogue model that predicts mel-spectrograms with flow matching, jointly trained with discrete text tokens, achieves lower WER than a discrete speech token baseline.

  15. WavChat: A Survey of Spoken Dialogue Models

    eess.AS 2024-11 conditional novelty 4.0 of 10

    WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.

Pith tools