Pith. sign in

REVIEW 3 cited by

PromptTTS: Controllable Text-to-Speech with Text Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12171 v1 pith:G5N4D7NA submitted 2022-11-22 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords speechstylepromptttstextcontentdescriptionscorrespondingdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently. Beyond text and image generation, in this work, we explore the possibility of utilizing text descriptions to guide speech synthesis. Thus, we develop a text-to-speech (TTS) system (dubbed as PromptTTS) that takes a prompt with both style and content descriptions as input to synthesize the corresponding speech. Specifically, PromptTTS consists of a style encoder and a content encoder to extract the corresponding representations from the prompt, and a speech decoder to synthesize speech according to the extracted style and content representations. Compared with previous works in controllable TTS that require users to have acoustic knowledge to understand style factors such as prosody and pitch, PromptTTS is more user-friendly since text descriptions are a more natural way to express speech style (e.g., ''A lady whispers to her friend slowly''). Given that there is no TTS dataset with prompts, to benchmark the task of PromptTTS, we construct and release a dataset containing prompts with style and content information and the corresponding speech. Experiments show that PromptTTS can generate speech with precise style control and high speech quality. Audio samples and our dataset are publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

    eess.AS 2026-03 conditional novelty 6.0 of 10

    Dual-encoder speech-text models trained on rich intrinsic and situational style captions outperform prior CLAP-style baselines on retrieval, classification, and inference-time TTS style guidance.

  2. RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    PCA of repeated synthesized utterances with fixed inputs can reveal controllable prosodic features that can be enrolled as new prompts via fine-tuning.

  3. EmoNews: A Spoken Dialogue System for Expressive News Conversations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    An emotional spoken dialogue system that uses a sentiment analyzer to pick an emotion tag and PromptTTS to synthesize matching speech outperforms a neutral baseline on perceived emotional appropriateness, but not sign...

Pith tools