Pith. sign in

REVIEW 6 cited by

WaveGrad: Estimating Gradients for Waveform Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.00713 v2 pith:5UQU6FN6 submitted 2020-09-02 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords wavegradaudioautoregressivefidelitygenerategenerationgradientshigh
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density. The model is built on prior work on score matching and diffusion probabilistic models. It starts from a Gaussian white noise signal and iteratively refines the signal via a gradient-based sampler conditioned on the mel-spectrogram. WaveGrad offers a natural way to trade inference speed for sample quality by adjusting the number of refinement steps, and bridges the gap between non-autoregressive and autoregressive models in terms of audio quality. We find that it can generate high fidelity audio samples using as few as six iterations. Experiments reveal WaveGrad to generate high fidelity audio, outperforming adversarial non-autoregressive baselines and matching a strong likelihood-based autoregressive baseline using fewer sequential operations. Audio samples are available at https://wavegrad.github.io/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Strong Gravitational Lensing Posterior Sampling in Pixel-Space Using Diffusion Models and Recurrent Inference Machines

    astro-ph.IM 2026-07 conditional novelty 7.0 of 10

    DiRIM uses a diffusion model with recurrent score refinement to sample pixel-space joint posteriors of the lensed source and foreground mass map, reproducing mock strong-lens observations to the noise level.

  2. O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion

    cs.SD 2025-10 conditional novelty 6.0 of 10

    Training a FreeVC-style voice converter on TTS-generated same-text/different-speaker pairs, then fine-tuning on real speech, yields improved zero-shot voice conversion on LibriSpeech.

  3. RealDeal: Enhancing Realism and Details in Brain Image Generation via Image-to-Image Diffusion Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    RealDeal refines smoothed latent-diffusion brain MRIs with patch-based image-to-image diffusion, cutting FID from 46.6 to 17.3 and improving LPIPS, noise, and sharpness metrics.

  4. MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection

    cs.SD 2025-09 conditional novelty 5.0 of 10

    MoLEx combines LoRA adapters with a top-K expert router inside a frozen WavLM model, achieving 5.56% EER on ASVSpoof 5 without augmentation.

  5. ArtifactGen: Benchmarking WGAN-GP vs Diffusion for Label-Aware EEG Artifact Synthesis

    cs.LG 2025-09 reject novelty 4.0 of 10

    A WGAN-GP achieves closer spectral alignment and lower MMD than a diffusion model for EEG artifact synthesis, but class-conditional recovery is weak for both.

  6. DualFast: Dual-Speedup Framework for Fast Sampling of Diffusion Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A training-free correction that blends each step's noise estimate with the initial noise estimate improves few-step diffusion sampling across DDIM, DPM-Solver, and DPM-Solver++.

Pith tools