Pith. sign in

REVIEW

Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21951 v2 pith:GCDR5EIX submitted 2024-10-29 eess.AS cs.AIcs.SD

classification eess.AScs.AIcs.SD
keywords speechauto-regressiveinferencedecodingspeculativetokensvadusaaccelerate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The auto-regressive architecture, like GPTs, is widely used in modern Text-to-Speech (TTS) systems. However, it incurs substantial inference time, particularly due to the challenges in the next-token prediction posed by lengthy sequences of speech tokens. In this work, we introduce VADUSA, one of the first approaches to accelerate auto-regressive TTS through speculative decoding. Our results show that VADUSA not only significantly improves inference speed but also enhances performance by incorporating draft heads to predict future speech content auto-regressively. Furthermore, the inclusion of a tolerance mechanism during sampling accelerates inference without compromising quality. Our approach demonstrates strong generalization across large datasets and various types of speech tokens.

Discussion (0). Continue with ORCID to comment.

Pith tools