Pith. sign in

REVIEW 15 cited by

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09116 v3 pith:5LMWFQSH submitted 2023-04-18 eess.AS cs.AIcs.CLcs.LGcs.SD

classification eess.AScs.AIcs.CLcs.LGcs.SD
keywords speechnaturalspeechsingingzero-shotdiffusionlatentqualityvoice
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large TTS systems usually quantize speech into discrete tokens and use language models to generate these tokens one by one, which suffer from unstable prosody, word skipping/repeating issue, and poor voice quality. In this paper, we develop NaturalSpeech 2, a TTS system that leverages a neural audio codec with residual vector quantizers to get the quantized latent vectors and uses a diffusion model to generate these latent vectors conditioned on text input. To enhance the zero-shot capability that is important to achieve diverse speech synthesis, we design a speech prompting mechanism to facilitate in-context learning in the diffusion model and the duration/pitch predictor. We scale NaturalSpeech 2 to large-scale datasets with 44K hours of speech and singing data and evaluate its voice quality on unseen speakers. NaturalSpeech 2 outperforms previous TTS systems by a large margin in terms of prosody/timbre similarity, robustness, and voice quality in a zero-shot setting, and performs novel zero-shot singing synthesis with only a speech prompt. Audio samples are available at https://speechresearch.github.io/naturalspeech2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  2. StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A mobile-oriented 83M-parameter masked transformer with sparse phone-anchored temporal embeddings achieves RTF 0.08 and lower WER than MaskGCT/F5-TTS on Seed-TTS test sets.

  3. SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision

    cs.SD 2026-07 conditional novelty 6.0 of 10

    A joint text-code trajectory supervision recipe lets a two-stream speech LM achieve competitive long-form streaming S2ST with ~2k hours of paired speech.

  4. OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model

    cs.SD 2026-02 conditional novelty 6.0 of 10

    A zero-shot model that generates a video of a reference face speaking user-chosen text with a reference voice timbre.

  5. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  6. Audio-Guided Visual Editing with Complex Multi-Modal Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A training-free framework that maps audio embeddings into Stable Diffusion's text space and fuses multiple audio/text prompts via per-patch residual noise selection, outperforming text-only editors on new audio-visual...

  7. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  8. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  9. Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

    eess.AS 2026-08 reject novelty 5.0 of 10

    An interactive genetic algorithm tunes arousal-valence coordinates per listener in emotional TTS, and personalized or culture-specific coordinates beat a generic U.S.-average baseline in small A/B tests.

  10. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  11. SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A three-stage diffusion pipeline maps 3D Gaussian Splatting object representations to position-dependent impact sounds, trained first on text captions and then on real recordings.

  12. SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

    eess.AS 2025-07 conditional novelty 5.0 of 10

    SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.

  13. FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks

    cs.CR 2025-08 conditional novelty 4.0 of 10

    FreeTalk adds masked, smoothed frequency-domain noise, optimized against a speaker-embedding model, to keep voice-cloning models from reproducing a victim's voice, while preserving speech-to-text accuracy.

  14. Inference-time Scaling for Diffusion-based Audio Super-resolution

    cs.SD 2025-08 conditional novelty 4.0 of 10

    Generating 120 candidate super-resolved audios and choosing the best by task-specific verifiers improves speech, music, and sound effects over single-sample diffusion output, at 120x compute.

  15. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

Pith tools