Pith. sign in

REVIEW 33 cited by

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09116 v3 pith:5LMWFQSH submitted 2023-04-18 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords speechnaturalspeechsingingzero-shotdiffusionlatentqualityvoice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually quantize speech into discrete tokens and use language models to generate these tokens one by one, which suffer from unstable prosody, word skipping/repeating issue, and poor voice quality. In this paper, we develop NaturalSpeech 2, a TTS system that leverages a neural audio codec with residual vector quantizers to get the quantized latent vectors and uses a diffusion model to generate these latent vectors conditioned on text input. To enhance the zero-shot capability that is important to achieve diverse speech synthesis, we design a speech prompting mechanism to facilitate in-context learning in the diffusion model and the duration/pitch predictor. We scale NaturalSpeech 2 to large-scale datasets with 44K hours of speech and singing data and evaluate its voice quality on unseen speakers. NaturalSpeech 2 outperforms previous TTS systems by a large margin in terms of prosody/timbre similarity, robustness, and voice quality in a zero-shot setting, and performs novel zero-shot singing synthesis with only a speech prompt. Audio samples are available at https://speechresearch.github.io/naturalspeech2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 33 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

    cs.SD 2026-06 unverdicted novelty 7.0 of 10

    FlexiSLM is the first spoken language model supporting dynamic and controllable frame rates on speech input and output, outperforming fixed-rate 7B models at high quality and enabling faster inference at lower rates l...

  2. MeshFlow: Mesh Generation with Equivariant Flow Matching

    cs.GR 2026-06 unverdicted novelty 7.0 of 10

    MeshFlow applies equivariant optimal-transport flow matching to generate triangle meshes as soups, matching autoregressive quality with an 18x inference speedup.

  3. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  4. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A mobile-oriented 83M-parameter masked transformer with sparse phone-anchored temporal embeddings achieves RTF 0.08 and lower WER than MaskGCT/F5-TTS on Seed-TTS test sets.

  5. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  6. EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    EmoInstruct-TTS uses Emotion2embed and an Instruction-Conditioned Emotion Flow Model (ICE-Flow) to generate acoustically grounded emotion representations from free-form instructions and integrate them into an LLM-base...

  7. SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    SemaVoice adds SFM-guided alignment to refine continuous speech representations in autoregressive TTS, reporting 1.71% English WER on Seed-TTS and competitiveness with open-source SOTA.

  8. Voice "Cloning" is Style Transfer

    cs.SD 2026-05 conditional novelty 6.0 of 10

    Voice cloning models perform style transfer rather than faithful replication, producing voices rated higher in authority, warmth, and trust while reducing variance in accent and speaking rate.

  9. Voice "Cloning" is Style Transfer

    cs.SD 2026-05 conditional novelty 6.0 of 10

    Voice cloning applies systematic style transfer rather than faithful replication, producing voices rated higher on authority and trust with reduced variance in accent and rate.

  10. A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    A framework detects speaker drift in TTS outputs by computing cosine similarities across speech segments and using LLMs for binary classification, supported by a human-validated synthetic benchmark.

  11. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  12. Qwen3-TTS Technical Report

    cs.SD 2026-01 unverdicted novelty 6.0 of 10

    Qwen3-TTS delivers state-of-the-art multilingual TTS performance with 3-second voice cloning, description control, and ultra-low-latency streaming via dual tokenizers and a dual-track LM architecture trained on over 5...

  13. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  14. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  15. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  16. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  17. JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching

    cs.CV 2025-06 unverdicted novelty 6.0 of 10

    JAM-Flow introduces a unified flow-matching model with a Multi-Modal Diffusion Transformer that jointly synthesizes facial motion and speech from text, audio, or motion inputs.

  18. Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

    eess.AS 2026-08 reject novelty 5.0 of 10

    An interactive genetic algorithm tunes arousal-valence coordinates per listener in emotional TTS, and personalized or culture-specific coordinates beat a generic U.S.-average baseline in small A/B tests.

  19. UniVoice: A Unified Model for Speech and Singing Voice Generation

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    UniVoice is a conditional flow matching model with a Diffusion Transformer backbone that unifies TTS and SVS via modality-specific encoders and a null melody token for speech, achieving 5.26% speech PER and 16.22% sin...

  20. Voice "Cloning" is Style Transfer

    cs.SD 2026-05 unverdicted novelty 5.0 of 10

    Voice cloning models perform style transfer rather than faithful cloning, producing voices rated as more authoritative and warm with reduced variance in accent and speaking rate.

  21. Scaling Properties of Continuous Diffusion Spoken Language Models

    cs.CL 2026-04 unverdicted novelty 5.0 of 10

    Continuous diffusion spoken language models follow scaling laws for loss and phoneme divergence and generate emotive multi-speaker speech at 16B scale, though long-form coherence stays difficult.

  22. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 accept novelty 5.0 of 10

    MLLM-enabled video translation is usefully framed as three roles—Semantic Reasoner, Expressive Performer, and Visual Synthesizer—rather than a cascade of ASR, MT, TTS, and lip-sync.

  23. TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for \"U-Tsang, Amdo and Kham Speech Dataset Generation

    cs.CL 2025-09 unverdicted novelty 5.0 of 10

    TMD-TTS is a unified multi-dialect TTS framework for Tibetan that uses explicit dialect labels, a fusion module, and DSDR-Net to synthesize expressive speech and enable speech-to-speech dialect conversion.

  24. SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A three-stage diffusion pipeline maps 3D Gaussian Splatting object representations to position-dependent impact sounds, trained first on text captions and then on real recordings.

  25. SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

    eess.AS 2025-07 conditional novelty 5.0 of 10

    SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.

  26. Movie Gen: A Cast of Media Foundation Models

    cs.CV 2024-10 unverdicted novelty 5.0 of 10

    A 30B-parameter transformer and related models generate high-quality videos and audio, claiming state-of-the-art results on text-to-video, video editing, personalization, and audio generation tasks.

  27. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

    eess.AS 2024-10 unverdicted novelty 5.0 of 10

    F5-TTS generates natural speech from text via flow matching on DiT with simple text padding, ConvNeXt refinement, and sway sampling, trained on 100K hours multilingual data.

  28. ZONOS2 Technical Report

    cs.SD 2026-06 unverdicted novelty 4.0 of 10

    ZONOS2 8B is a scaled MoE TTS model with 900M active parameters trained on 6M hours of data that reports competitive SOTA results on naturalness, speaker similarity, WER, and a new ZTTS1-Eval benchmark while releasing...

  29. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  30. FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks

    cs.CR 2025-08 conditional novelty 4.0 of 10

    FreeTalk adds masked, smoothed frequency-domain noise, optimized against a speaker-embedding model, to keep voice-cloning models from reproducing a victim's voice, while preserving speech-to-text accuracy.

  31. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

  32. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

  33. ZONOS2 Technical Report

    cs.SD 2026-06 unverdicted novelty 3.0 of 10

    ZONOS2 8B scales a prior TTS system to 8B parameters with MoE architecture and 6M hours of data, reporting competitive benchmark performance on naturalness and speaker similarity while releasing weights.

Pith tools