REVIEW 27 cited by
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces VALL-E 2, the latest advancement in neural codec language models that marks a milestone in zero-shot text-to-speech synthesis (TTS), achieving human parity for the first time. Based on its predecessor, VALL-E, the new iteration introduces two significant enhancements: Repetition Aware Sampling refines the original nucleus sampling process by accounting for token repetition in the decoding history. It not only stabilizes the decoding but also circumvents the infinite loop issue. Grouped Code Modeling organizes codec codes into groups to effectively shorten the sequence length, which not only boosts inference speed but also addresses the challenges of long sequence modeling. Our experiments on the LibriSpeech and VCTK datasets show that VALL-E 2 surpasses previous systems in speech robustness, naturalness, and speaker similarity. It is the first of its kind to reach human parity on these benchmarks. Moreover, VALL-E 2 consistently synthesizes high-quality speech, even for sentences that are traditionally challenging due to their complexity or repetitive phrases. The advantages of this work could contribute to valuable endeavors, such as generating speech for individuals with aphasia or people with amyotrophic lateral sclerosis. See https://aka.ms/valle2 for demos of VALL-E 2.
Forward citations
Cited by 27 Pith papers
-
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
AST performs training-free text-based speech editing by stitching inverted source latents with synthesized targets and adaptively guiding the flow-matching decoder, achieving state-of-the-art temporal fidelity and spe...
-
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
A neural speech codec that dynamically varies frame rate per segment using waveform entropy achieves competitive or better reconstruction quality at lower average frame rates than constant frame rate baselines.
-
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Autoregressive TTS from 8-Hz, 768-dimensional continuous tokens works when the tokenizer shapes its latent space with a low-dimensional core and an energy hierarchy, and the generator separates guidance into local, se...
-
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
AU-supervised single-token face encoding plus dual visual–speech DPO on a large real-conversation dataset improves empathetic conversational TTS over text/speech-only and prior visual CSS systems.
-
SSTMark: Robust Training-Free Semantic-Level Speech Watermarking
SSTMark embeds a watermark in AI speech by rewriting its transcript and resynthesizing it, achieving strong average robustness to audio distortions but at the cost of altering the spoken content.
-
Unified Audio Intelligence Without Regressing on Text Intelligence
A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.
-
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Block discrete diffusion over X-Codec2 tokens yields competitive zero-shot TTS with 0.6B parameters, 20K training hours, and a 0.15 real-time factor.
-
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
A new metric, MCLP, uses a frozen audio-LLM's continuation likelihood to grade speaking-style consistency and doubles as a reward that improves role-play TTS on a new 1,435-hour drama dataset.
-
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
A single LLM can do ASR and zero-shot TTS on continuous speech features by switching between causal and bidirectional attention, reaching competitive but not state-of-the-art results.
-
DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
A compact zero-shot TTS that applies discrete flow matching with separate prediction heads for prosody and acoustic tokens, reporting near-best quality, best prosody/energy metrics, and up to 25.8x faster inference.
-
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
DiTReducio is a training-free, pattern-guided layer and branch skipping method that accelerates DiT-based TTS, reporting significant FLOP and RTF reductions with modest quality loss at tuned thresholds.
-
Next Tokens Denoising for Speech Synthesis
Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.
-
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
Reinforcement learning on duration prediction improves intelligibility and speaker similarity in a 4-step distilled text-to-speech model, and teacher-guided sampling recovers prosodic diversity.
-
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.
-
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
Integrating a VAE with an autoregressive prior into token-based spoken language modeling learns continuous variational features that improve naturalness without hand-engineered pitch, with human raters preferring the ...
-
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment
A new benchmark dataset and two tasks let researchers test whether AI systems can judge if a news image's recorded location and date match the article, with current chatbots scoring 64-81% on location but 42-58% on date.
-
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
NTPP models dual-channel spoken dialogue by predicting both speakers' next speech tokens as a pair, achieving speaker-independent full-duplex generation in a decoder-only transformer.
-
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.
-
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
SMLLE generates speech frame-by-frame using a Transducer for streaming semantic tokens plus a fully autoregressive mel-spectrogram model, reaching quality close to sentence-level zero-shot TTS.
-
Position: Towards Responsible Evaluation for Text-to-Speech
A call to reform text-to-speech evaluation around a three-level framework covering metric fidelity, comparability, and ethical oversight.
-
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
A dual-branch pooling method, DRASP, combines global statistics with segment-level attention and improves MOS prediction correlation with human ratings.
-
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.
-
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...
-
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
A latent diffusion model with SSL-based content restoration and in-context speaker prompts improves dysarthric speech intelligibility and speaker similarity on UASpeech.
-
Speaking images. A novel framework for the automated self-description of artworks
A four-stage open-source AI pipeline turns a digitized artwork into a short video where a depicted person animates and narrates the scene.
-
Voice Adaptation for Swiss German
Fine-tuning XTTS-v2 on about 5,000 hours of weakly labeled Swiss podcast audio produces a voice adaptation model that renders Standard German text in seven Swiss German dialect regions with near-reference quality in h...
-
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
Fine-tuning Chatterbox and CosyVoice 3 on 50 Singlish speakers measurably raises accent similarity, and the gain persists on held-out speakers.
Discussion (0). Sign in to comment.