Pith. sign in

REVIEW 7 cited by

A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13336 v2 pith:OXWRS5FD submitted 2023-03-23 cs.SD cs.AIcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.LGcs.MMeess.AS
keywords modelspeechdiffusionaudioenhancementgenerativesurveysynthesis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active tasks: text to speech and speech enhancement. This work conducts a survey on audio diffusion model, which is complementary to existing surveys that either lack the recent progress of diffusion-based speech synthesis or highlight an overall picture of applying diffusion model in multiple fields. Specifically, this work first briefly introduces the background of audio and diffusion model. As for the text-to-speech task, we divide the methods into three categories based on the stage where diffusion model is adopted: acoustic model, vocoder and end-to-end framework. Moreover, we categorize various speech enhancement tasks by either certain signals are removed or added into the input speech. Comparisons of experimental results and discussions are also covered in this survey.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Permutation-Invariant Spectral Learning via Dyson Diffusion

    stat.ML 2025-10 conditional novelty 7.0 of 10

    Dyson Diffusion Model learns graph spectra by scoring Dyson Brownian Motion and matches or beats GNN/transformer diffusion baselines without ad-hoc spectral features.

  2. The Effect of Stochasticity in Score-Based Diffusion Sampling: a KL Divergence Analysis

    cs.LG 2025-06 conditional novelty 7.0 of 10

    KL divergence bounds show stochasticity in diffusion sampling contracts error with exact scores, but for learned scores it can help or hurt depending on the time profile of the score error.

  3. MPQ-DMv2: Flexible Residual Mixed Precision Quantization for Low-Bit Diffusion Models with Temporal Distillation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MPQ-DMv2 adds binary residual quantization, temporal relation distillation, and SVD-initialized LoRA to mixed-precision quantization, improving low-bit diffusion model generation quality.

  4. UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

    cs.SD 2025-06 conditional novelty 5.0 of 10

    UmbraTTS jointly synthesizes speech and environmental audio via conditional flow matching, conditioned on text and acoustic context, with controllable background volume.

  5. Marco-Voice Technical Report

    cs.CL 2025-08 reject novelty 4.0 of 10

    Marco-Voice is a TTS system combining voice cloning and emotional speech generation via speaker-emotion disentanglement, contrastive learning, and a new Mandarin emotional dataset, with claimed quality gains over Cosy...

  6. Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.

  7. SyntheticPop: Attacking Speaker Verification Systems With Synthetic VoicePops

    cs.CR 2025-02 conditional novelty 4.0 of 10

    SyntheticPop adds a low-frequency sine tone to spoofed training audio and drops a VoicePop-based voice authentication system from 69% to 14% accuracy.

Pith tools