Pith. sign in

REVIEW 27 cited by

VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05370 v2 pith:KYN77HD5 submitted 2024-06-08 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords vall-espeechcodechumanparitydecodingfirstintroduces
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity for the first time. Based on its predecessor, VALL-E, the new iteration introduces two significant enhancements: Repetition Aware Sampling refines the original nucleus sampling process by accounting for token repetition in the decoding history. It not only stabilizes the decoding but also circumvents the infinite loop issue. Grouped Code Modeling organizes codec codes into groups to effectively shorten the sequence length, which not only boosts inference speed but also addresses the challenges of long sequence modeling. Our experiments on the LibriSpeech and VCTK datasets show that VALL-E 2 surpasses previous systems in speech robustness, naturalness, and speaker similarity. It is the first of its kind to reach human parity on these benchmarks. Moreover, VALL-E 2 consistently synthesizes high-quality speech, even for sentences that are traditionally challenging due to their complexity or repetitive phrases. The advantages of this work could contribute to valuable endeavors, such as generating speech for individuals with aphasia or people with amyotrophic lateral sclerosis. See https://aka.ms/valle2 for demos of VALL-E 2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 27 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AST: Adaptive, Seamless, and Training-Free Precise Speech Editing

    cs.SD 2026-04 unverdicted novelty 7.0 of 10

    AST performs training-free text-based speech editing by stitching inverted source latents with synthesized targets and adaptively guiding the flow-matching decoder, achieving state-of-the-art temporal fidelity and spe...

  2. Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

    eess.AS 2025-05 conditional novelty 7.0 of 10

    A neural speech codec that dynamically varies frame rate per segment using waveform entropy achieves competitive or better reconstruction quality at lower average frame rates than constant frame rate baselines.

  3. Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...

  4. Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

    cs.HC 2026-07 conditional novelty 6.0 of 10

    AU-supervised single-token face encoding plus dual visual–speech DPO on a large real-conversation dataset improves empathetic conversational TTS over text/speech-only and prior visual CSS systems.

  5. SSTMark: Robust Training-Free Semantic-Level Speech Watermarking

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SSTMark embeds a watermark in AI speech by rewriting its transcript and resynthesizing it, achieving strong average robustness to audio distortions but at the cost of altering the spoken content.

  6. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  7. DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis

    cs.CL 2026-06 conditional novelty 6.0 of 10

    Block discrete diffusion over X-Codec2 tokens yields competitive zero-shot TTS with 0.6B parameters, 20K training hours, and a 0.15 real-time factor.

  8. Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.

  9. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models

    eess.AS 2025-10 conditional novelty 6.0 of 10

    A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.

  10. DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A compact zero-shot TTS that applies discrete flow matching with separate prediction heads for prosody and acoustic tokens, reporting near-best quality, best prosody/energy metrics, and up to 25.8x faster inference.

  11. DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DiTReducio is a training-free, pattern-guided layer and branch skipping method that accelerates DiT-based TTS, reporting significant FLOP and RTF reductions with modest quality loss at tuned thresholds.

  12. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  13. DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis

    eess.AS 2025-07 conditional novelty 6.0 of 10

    Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.

  14. Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

    cs.SD 2025-06 conditional novelty 6.0 of 10

    DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.

  15. A Variational Framework for Improving Naturalness in Generative Spoken Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Integrating a VAE with an autoregressive prior into token-based spoken language modeling learns continuous variational features that improve naturalness without hand-engineered pitch, with human raters preferring the ...

  16. Dataset of News Articles with Provenance Metadata for Media Relevance Assessment

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark dataset and two tasks let researchers test whether AI systems can judge if a news image's recorded location and date match the article, with current chatbots scoring 64-81% on location but 42-58% on date.

  17. NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

    cs.CL 2025-06 conditional novelty 6.0 of 10

    NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.

  18. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  19. Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

    cs.LG 2025-05 conditional novelty 6.0 of 10

    SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.

  20. Position: Towards Responsible Evaluation for Text-to-Speech

    eess.AS 2025-10 conditional novelty 5.0 of 10

    A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.

  21. DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A dual-branch pooling method, DRASP, combines global statistics with segment-level attention and improves MOS prediction correlation with human ratings.

  22. CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.

  23. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...

  24. DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A latent diffusion model with SSL-based content restoration and in-context speaker prompts improves dysarthric speech intelligibility and speaker similarity on UASpeech.

  25. Speaking images. A novel framework for the automated self-description of artworks

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.

  26. Voice Adaptation for Swiss German

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...

  27. Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English

    eess.AS 2026-07 conditional novelty 4.0 of 10

    Fine-tuning Chatterbox and CosyVoice 3 on 50 Singlish speakers measurably raises accent similarity, and the gain persists on held-out speakers.

Pith tools