Pith. sign in

REVIEW 23 cited by

Better speech synthesis through scaling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07243 v2 pith:TSMOJEVX submitted 2023-05-12 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords imagebeengenerationmodelspeechsynthesisadvancesamounts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the field of image generation has been revolutionized by the application of autoregressive transformers and DDPMs. These approaches model the process of image generation as a step-wise probabilistic processes and leverage large amounts of compute and data to learn the image distribution. This methodology of improving performance need not be confined to images. This paper describes a way to apply advances in the image generative domain to speech synthesis. The result is TorToise -- an expressive, multi-voice text-to-speech system. All model code and trained weights have been open-sourced at https://github.com/neonbjb/tortoise-tts.

Discussion (0). Sign in to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    FlexiSLM is the first spoken language model supporting dynamic and controllable frame rates on speech input and output, outperforming fixed-rate 7B models at high quality and enabling faster inference at lower rates l...

  2. Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    Sarashina2.2-TTS achieves SOTA kanji reading accuracy via data scaling and Joyo-kanji-targeted synthesis, introduces the Joyo Kanji Yomi Benchmark and Kana-CER metric, and shows stable cross-lingual performance.

  3. An Evaluation Framework for Text-to-Speech Voice Reconstruction

    eess.AS 2026-06 unverdicted novelty 6.0 of 10

    The paper introduces a subjective-objective evaluation framework using Best Worst Scaling and a novel dual-reference distributional measure to better assess intelligibility versus speaker identity trade-offs in TTS vo...

  4. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    cs.SD 2026-05 unverdicted novelty 6.0 of 10

    X-Voice achieves zero-shot cross-lingual voice cloning across 30 languages by using IPA as a unified phonetic representation and a two-stage training process that first generates its own audio prompts then fine-tunes ...

  5. Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Jointly training the watermark embedder/detector with the source separator enables ~1% bit-error-rate recovery of per-stem watermarks after mixing and separation, where independent training yields 15–35%.

  6. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  7. SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods

    cs.SD 2025-07 conditional novelty 6.0 of 10

    SpeechFake is a large-scale multilingual deepfake speech dataset with baseline experiments showing improved generalization to unseen generation methods.

  8. TTS-1 Technical Report

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new TTS system family combines pre-training, supervised fine-tuning, and GRPO reinforcement learning with a 48 kHz codec to produce multilingual speech with in-context voice cloning.

  9. Step-Audio 2 Technical Report

    cs.CL 2025-07 unverdicted novelty 6.0 of 10

    Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and c...

  10. DeePen: Penetration Testing for Audio Deepfake Detection

    cs.CR 2025-02 unverdicted novelty 6.0 of 10

    DeePen demonstrates that both production and academic audio deepfake detectors can be reliably deceived by simple signal processing attacks such as time-stretching or echo addition, with some attacks resistible via re...

  11. Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

    eess.AS 2024-06 unverdicted novelty 6.0 of 10

    Seed-TTS models produce speech matching human naturalness and speaker similarity, with added controllability via self-distillation and reinforcement learning.

  12. MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

    cs.SD 2024-01 unverdicted novelty 6.0 of 10

    MLAAD provides a large-scale multi-language synthetic audio dataset for training and evaluating audio anti-spoofing models, showing better training performance than InTheWild and FakeOrReal and alternating superiority...

  13. FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation

    eess.AS 2026-06 unverdicted novelty 5.0 of 10

    FlashTTS delivers a streaming TTS system using multi-track input processing and X-pred mean flow matching to reach 325 ms latency in two function evaluations while retaining zero-shot voice cloning.

  14. X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

    cs.SD 2026-05 unverdicted novelty 5.0 of 10

    X-Voice achieves zero-shot cross-lingual voice cloning across 30 languages via IPA-based training on 420K hours of data and a two-stage paradigm that synthesizes its own audio prompts for text-masked fine-tuning.

  15. Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning

    eess.AS 2026-04 unverdicted novelty 5.0 of 10

    A cascaded audio-prompting and ICL-based online RL method improves naturalness and expressivity in conversational TTS with reduced data needs.

  16. Large Language Model Data Generation for Enhanced Intent Recognition in German Speech

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    LLM-generated German text data improves intent recognition for elderly German speakers, and the smaller German-focused LeoLM outperforms the much larger ChatGPT as a data generator.

  17. Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A four-modality contrastive model with token-level alignment improves text, video, and newly introduced audio-to-motion retrieval on HumanML3D and KIT-ML.

  18. Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

    cs.SD 2025-07 reject novelty 5.0 of 10

    QTTS models speech as sequences from a multi-codebook RVQ audio codec whose first codebook is trained with ASR supervision, aiming for higher-fidelity TTS than single-codebook systems.

  19. De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Existing voice-protection perturbations succeed only against naive attackers; a phoneme-guided purification-refinement pipeline restores cloneability of protected speech for most VC models.

  20. Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Rectified flows with a tunable sampling temperature offer the best naturalness-diversity trade-off among stochastic prosody predictors for text-to-speech.

  21. StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding

    cs.SD 2025-06 conditional novelty 5.0 of 10

    StreamFlow achieves streaming speech token decoding with 180 ms first-packet latency by using hierarchical block-wise attention masks in a DiT flow matching model.

  22. JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1

    cs.CV 2025-07 reject novelty 4.0 of 10

    A paper announcing a large-scale whole-body talking avatar benchmark and evaluation protocol, but with insufficient details to verify the dataset or the joint audio-video evaluation.

  23. AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

    cs.SD 2026-04 unverdicted novelty 3.0 of 10

    AT-ADD introduces standardized tracks and datasets for evaluating audio deepfake detectors on speech under real-world conditions and on diverse unknown audio types to promote generalization beyond speech-centric methods.

Pith tools