REVIEW 37 cited by
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high/medium resource languages, limiting the applications of these models in most of the low/medium resource languages. In this paper, we aim to alleviate this issue by proposing and making publicly available the XTTS system. Our method builds upon the Tortoise model and adds several novel modifications to enable multilingual training, improve voice cloning, and enable faster training and inference. XTTS was trained in 16 languages and achieved state-of-the-art (SOTA) results in most of them.
Forward citations
Cited by 37 Pith papers
-
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
SemBridge supervises autoregressive states with discrete semantic tokens during training, improving content fidelity of continuous-latent speech generation without changing inference.
-
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
A 260-hour emotional deepfake benchmark spanning 21 attack systems shows state-of-the-art speech deepfake detectors degrade badly on emotionally expressive and LALM-based spoofing.
-
Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech
Experience-Calibrated Contrastive Decoding, a training-free decoding method that strengthens text alignment signals, reduces speech hallucination errors across four LM-based TTS models and nine languages.
-
Tell me Habibi, is it Real or Fake?
ArEnAV, the first large-scale Arabic-English code-switched audio-visual deepfake dataset, makes current state-of-the-art detectors fail much more than on monolingual data.
-
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
ShiftySpeech is a large benchmark demonstrating that audio deepfake detectors degrade under distribution shifts such as noise, emotion, language, and newly released vocoders.
-
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
MovieBench is the first public movie-level benchmark for long video generation, with hierarchical annotations, a character bank with audio, and new character-consistency metrics.
-
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization
A jointly trained speech tokenizer and flow-matching decoder achieve strong zero-shot TTS and voice conversion on English and Mandarin, with WER below ground truth in several tests.
-
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
A four-backbone SSL ensemble achieves near-perfect deepfake detection in the ImageCLEF 2026 track, while the same team's generated audio ranks first in the generation track, revealing an OR-versus-AND asymmetry betwee...
-
Large Audio Language Models for Spoofing-Aware Speaker Verification
Adapted LALMs can reach competitive spoofing-aware speaker verification (89.3% accuracy, 0.19 min a-DCF on an ASVspoof5 subset), though zero-shot performance is near chance.
-
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
Underrepresented gender in training suffers higher deepfake-detection error; WavLM gaps stay large under balance, and all post-hoc calibrations leave the EER gap fixed at 1.317 pp.
-
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis
A compact tokenizer-free non-autoregressive flow-matching DiT synthesizes Turkish speech in frozen AudioVAE2 latents at WER 8.0% / CER 3.0%, beating larger open cloners while running at RTF 0.11 on consumer GPUs.
-
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
A Taiwan-specific tokenizer, language model, and bridge to a reused acoustic stack cut code-switching TTS CER from 11.45% to 4.81%, with 65.6% listener preference.
-
Conversational Human Audio-visual Talking Dialogue Generation
CHAT generates mutually responsive dyadic audio-visual dialogue clips from a single text prompt and yields a 50k synthetic pre-training set that improves facial reaction models on REACT 2024.
-
ParsVoice: A Large-Scale Multi-Speaker Persian Speech Corpus for Text-to-Speech Synthesis
ParsVoice is an open ~1,800–2,200-hour Persian audiobook-derived speech-text corpus, much larger than prior open Persian TTS datasets.
-
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
Evaluating voice anonymization at low false-positive rates reveals much stronger membership inference and attribute leakage than Equal Error Rate reports.
-
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
XEmoRAG synthesizes Thai speech with emotions cloned from Chinese reference audio by retrieving matching Thai prompts and aligning prosody with flow matching, outperforming baseline TTS in emotion similarity and intel...
-
AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
AV-Deepfake1M++ is a 2.05-million-clip audio-visual deepfake benchmark spanning three source corpora, nine generators, and 36 perturbations, with challenge results showing large generalization gaps.
-
ClaritySpeech: Dementia Obfuscation in Speech
An ASR, text-obfuscation, and zero-shot TTS pipeline lowers automatic dementia detection in speech by 10 to 16 percent F1 while improving intelligibility, with only moderate speaker similarity.
-
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
A two-stage language-model TTS system uses quantized masked-autoencoder style tokens plus discrete attribute labels to achieve fine-grained style control with stable content.
-
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.
-
LoRP-TTS: Low-Rank Personalized Text-To-Speech
LoRP-TTS shows that per-prompt LoRA fine-tuning with one short recording improves speaker similarity in Voicebox-based zero-shot TTS, at some cost in inference time.
-
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis
Continuous autoregressive text-to-speech with a Gaussian-mixture codec matches or beats a discrete-codec VALL-E baseline with a fraction of the language model parameters.
-
SOLAMI: Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters
SOLAMI is an end-to-end social vision-language-action model that takes a user's speech and body motion as input and generates a 3D character's spoken and gestural responses in one pass.
-
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
QTTS models speech as sequences from a multi-codebook RVQ audio codec whose first codebook is trained with ASR supervision, aiming for higher-fidelity TTS than single-codebook systems.
-
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.
-
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges
A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.
-
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation
DS-TTS adds a second MFCC-based style encoder and a length-adaptive variance adapter to a StyleSpeech-style TTS model, reporting higher speaker similarity but not lower WER than two strong baselines.
-
Speaking images. A novel framework for the automated self-description of artworks
A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.
-
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
A universal adversarial perturbation framework claiming to protect speech against zero-shot voice cloning by degrading cloned outputs while preserving input naturalness.
-
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
Fine-tuning TTS models on tens of hours of real audio enables generation of 500,000 hours of synthetic speech that reduces ASR error rates by over 30% on Whisper-large-v3.
-
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
MiniMax-Speech reports state-of-the-art zero-shot voice cloning quality using a learnable speaker encoder and Flow-VAE, without requiring reference transcripts.
-
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
IndexTTS combines character and pinyin modeling to make Chinese polyphone pronunciation controllable, and reports improved zero-shot voice cloning and naturalness over open-source TTS baselines.
-
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
A masked generative model pre-trained on unlabeled speech then fine-tuned per task matches or beats task-specific systems across TTS, voice conversion, speaker extraction, enhancement, and lip-to-speech.
-
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
CrossSpeech++ disentangles language and speaker information in speech generation and reports improved cross-lingual TTS naturalness over prior systems on four languages.
-
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder
Confucius4-TTS performs transcript-free, cross-lingual zero-shot voice cloning in 14 languages with competitive intelligibility and speaker similarity.
-
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
A challenge entry combining Wav2Vec-AASIST audio scores with lightweight handcrafted-feature video scores via calibration and maxout reports 92.78% AUC on AV-Deepfake1M++ testA.
-
WavChat: A Survey of Spoken Dialogue Models
WavChat categorizes spoken dialogue models into cascaded and end-to-end paradigms and surveys speech representations, training strategies, streaming, duplex interaction, datasets, and evaluation benchmarks.
Discussion (0). Continue with ORCID to comment.