REVIEW 15 cited by
PromptTTS 2: Describing and Generating Voices with Text Prompt
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Speech conveys more information than text, as the same word can be uttered in various voices to convey diverse information. Compared to traditional text-to-speech (TTS) methods relying on speech prompts (reference speech) for voice variability, using text prompts (descriptions) is more user-friendly since speech prompts can be hard to find or may not exist at all. TTS approaches based on the text prompt face two main challenges: 1) the one-to-many problem, where not all details about voice variability can be described in the text prompt, and 2) the limited availability of text prompt datasets, where vendors and large cost of data labeling are required to write text prompts for speech. In this work, we introduce PromptTTS 2 to address these challenges with a variation network to provide variability information of voice not captured by text prompts, and a prompt generation pipeline to utilize the large language models (LLM) to compose high quality text prompts. Specifically, the variation network predicts the representation extracted from the reference speech (which contains full information about voice variability) based on the text prompt representation. For the prompt generation pipeline, it generates text prompts for speech with a speech language understanding model to recognize voice attributes (e.g., gender, speed) from speech and a large language model to formulate text prompts based on the recognition results. Experiments on a large-scale (44K hours) speech dataset demonstrate that compared to the previous works, PromptTTS 2 generates voices more consistent with text prompts and supports the sampling of diverse voice variability, thereby offering users more choices on voice generation. Additionally, the prompt generation pipeline produces high-quality text prompts, eliminating the large labeling cost. The demo page of PromptTTS 2 is available online.
Forward citations
Cited by 15 Pith papers
-
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
Direction-vector style interpolation plus KV-cache swap and sliding-window masking unlock continuous inter- and intra-utterance style control in prompt-based autoregressive TTS without training.
-
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.
-
Next Tokens Denoising for Speech Synthesis
Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.
-
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
Vo-Ve, a 44-dimension vector of voice-attribute probabilities, offers attribute-level explanations for speaker similarity, but its discrimination accuracy and listener-above-chance validation are modest.
-
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
InstructTTSEval introduces three tasks and 6,000 English and Chinese test cases for judging how well TTS systems follow natural-language style instructions, using Gemini as the evaluator.
-
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
A two-stage language-model TTS system uses quantized masked-autoencoder style tokens plus discrete attribute labels to achieve fine-grained style control with stable content.
-
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.
-
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
AutoSIFT disentangles text-describable style categories from residual speech styles and selectively infills only the categories the user specifies.
-
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.
-
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt
A multi-modal emotion prompt encoder and prosody predictor let MPE-TTS control emotion from speech, text, or image while preserving speaker timbre.
-
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
A zero-shot pipeline that creates character voices from AI-generated faces and LLM-written prosody instructions can produce expressive audiobooks without extra training or manual annotation, though human quality score...
-
Gender Bias in Instruction-Guided Speech Synthesis Models
Parler-TTS speech models produce gender-stereotyped voices for occupation prompts, and prompt-based mitigation is inconsistent and can reverse the bias.
-
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
CtrlSpeech adds phone-level pitch, loudness, and duration controls to a diffusion-transformer TTS backbone, improving local prosody control while preserving zero-shot speaker similarity.
-
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
A speech-to-speech dialogue model that predicts mel-spectrograms with flow matching, jointly trained with discrete text tokens, achieves lower WER than a discrete speech token baseline.
-
WavChat: A Survey of Spoken Dialogue Models
WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.
Discussion (0). Continue with ORCID to comment.