Pith. sign in

REVIEW 2 cited by

WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09660 v2 pith:IE5FR6SY submitted 2021-06-17 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords wavegradmodelrefinementaudioiterativesequencestepsphoneme
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces WaveGrad 2, a non-autoregressive generative model for text-to-speech synthesis. WaveGrad 2 is trained to estimate the gradient of the log conditional density of the waveform given a phoneme sequence. The model takes an input phoneme sequence, and through an iterative refinement process, generates an audio waveform. This contrasts to the original WaveGrad vocoder which conditions on mel-spectrogram features, generated by a separate model. The iterative refinement process starts from Gaussian noise, and through a series of refinement steps (e.g., 50 steps), progressively recovers the audio sequence. WaveGrad 2 offers a natural way to trade-off between inference speed and sample quality, through adjusting the number of refinement steps. Experiments show that the model can generate high fidelity audio, approaching the performance of a state-of-the-art neural TTS system. We also report various ablation studies over different model configurations. Audio samples are available at https://wavegrad.github.io/v2.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow Matching Policy Gradients

    cs.LG 2025-07 conditional novelty 7.0 of 10

    FPO trains flow-based policies with PPO by replacing the likelihood ratio with an exponentiated flow matching loss difference.

  2. SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture

    cs.SD 2025-06 conditional novelty 5.0 of 10

    A reference-guided diffusion model with a low-frequency upsampling module achieves marginal quality improvements over prior SVS baselines on Opencpop, with significant caveats about statistical significance and reprod...

Pith tools