Pith. sign in

REVIEW

VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05027 v3 pith:PIK2KXEA submitted 2023-09-10 eess.AS cs.AIcs.HCcs.SD

classification eess.AScs.AIcs.HCcs.SD
keywords voiceflowflowrectifieddiffusionsamplingsynthesisefficientmatching
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an acoustic model that utilizes a rectified flow matching algorithm to achieve high synthesis quality with a limited number of sampling steps. VoiceFlow formulates the process of generating mel-spectrograms into an ordinary differential equation conditional on text inputs, whose vector field is then estimated. The rectified flow technique then effectively straightens its sampling trajectory for efficient synthesis. Subjective and objective evaluations on both single and multi-speaker corpora showed the superior synthesis quality of VoiceFlow compared to the diffusion counterpart. Ablation studies further verified the validity of the rectified flow technique in VoiceFlow.

Discussion (0). Sign in to comment.

Pith tools