Pith. sign in

REVIEW

DiffVoice: Text-to-Speech with Latent Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.11750 v1 pith:R2B5X4D3 submitted 2023-04-23 eess.AS cs.AIcs.HCcs.LGcs.SD

classification eess.AScs.AIcs.HCcs.LGcs.SD
keywords diffusionlatentdiffvoicemodelrepresentationspeechtext-to-speechachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we present DiffVoice, a novel text-to-speech model based on latent diffusion. We propose to first encode speech signals into a phoneme-rate latent representation with a variational autoencoder enhanced by adversarial training, and then jointly model the duration and the latent representation with a diffusion model. Subjective evaluations on LJSpeech and LibriTTS datasets demonstrate that our method beats the best publicly available systems in naturalness. By adopting recent generative inverse problem solving algorithms for diffusion models, DiffVoice achieves the state-of-the-art performance in text-based speech editing, and zero-shot adaptation.

Discussion (0). Sign in to comment.

Pith tools