REVIEW 18 cited by
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference speech recordings, limiting creative applications. Alternatively, natural language prompting of speaker identity and style has demonstrated promising results and provides an intuitive method of control. However, reliance on human-labeled descriptions prevents scaling to large datasets. Our work bridges the gap between these two approaches. We propose a scalable method for labeling various aspects of speaker identity, style, and recording conditions. We then apply this method to a 45k hour dataset, which we use to train a speech language model. Furthermore, we propose simple methods for increasing audio fidelity, significantly outperforming recent work despite relying entirely on found data. Our results demonstrate high-fidelity speech generation in a diverse range of accents, prosodic styles, channel conditions, and acoustic conditions, all accomplished with a single model and intuitive natural language conditioning. Audio samples can be heard at https://text-description-to-speech.com/.
Forward citations
Cited by 18 Pith papers
-
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
LRLspoof corpus and threshold-transfer evaluation demonstrate that spoof detection performance varies markedly across languages, identifying language as an independent domain shift factor.
-
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.
-
Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer
Depth-only pruning of a flow-matching Hindi TTS teacher, followed by staged re-fine-tuning, produces 131–190M students with ASR-WER close to the teacher and real-time laptop inference.
-
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
Direction-vector style interpolation plus KV-cache swap and sliding-window masking unlock continuous inter- and intra-utterance style control in prompt-based autoregressive TTS without training.
-
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.
-
Multi-interaction TTS toward professional recording reproduction
Multi-turn textual directions can iteratively refine the speaking style of synthesized speech through a learned embedding refiner, with modest but measurable alignment to the directions.
-
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
VoiceStar uses a progress-based rotary position embedding and mixed prompt training to give zero-shot voice cloning precise duration control and much longer output than training clips.
-
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
A new 13,000-hour dataset with 24 million LLM-generated text descriptions enables the first open-source text-driven TTS for 24 Indian languages, with reported high speaker, emotion, and cross-lingual control.
-
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
A unified multi-trait speech benchmark, Vox-Profile, evaluates foundation models on age, sex, accent, emotion, voice quality, fluency, and expressiveness, and demonstrates downstream uses in ASR analysis and speech ge...
-
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
No single Urdu TTS system wins across all metrics, and the system listeners liked most, Google Gemini, was the furthest from reference audio on objective acoustic scores.
-
Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
FSQ-based audio codecs outperform RVQ-based codecs in simulated bit-error transmission, preserving intelligibility at bit-flip rates up to 10%.
-
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts
MoE-TTS adds frozen text-expert MoE modules to a Qwen3-based TTS system and reports better out-of-domain description alignment than ElevenLabs and MiniMax on a small hand-built test set.
-
Unlocking Speech Instruction Data Potential with Query Rewriting
A multi-LLM rewriting and multi-agent validation pipeline makes text-to-speech synthesized speech instruction data far more usable and improves downstream speech instruction following.
-
SecureSpeech: Prompt-based Speaker and Content Protection
A dual anonymization pipeline that uses an LLM to replace sensitive entities and a prompt-driven TTS to generate speech with a new, unrelated voice.
-
MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications
MATE is an open-source multi-agent system for accessibility that uses a fine-tuned BERT model, trained on a new AI-generated dataset, to recognize and execute modality conversion tasks.
-
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
RV-TTS synthesizes speech whose voice matches a face image, including artistic portraits, with natural-language control over pace, noise, distance, tone, and recording place.
-
BoSS: Beyond-Semantic Speech
Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.
-
Optimizing Multilingual Text-To-Speech with Accents & Emotions
A TTS system built on Parler-TTS is claimed to improve accent accuracy and emotional expressiveness for Hindi and Indian English, but the paper lacks detailed architecture and baseline evidence.
Discussion (0). Sign in to comment.