Learning latent representations for style control and transfer in end-to-end speech synthesis

Lei He; Shifeng Pan; Ya-Jie Zhang; Zhen-Hua Ling

arxiv: 1812.04342 · v2 · pith:B3KVZEIQnew · submitted 2018-12-11 · 💻 cs.CL · cs.SD· eess.AS

Learning latent representations for style control and transfer in end-to-end speech synthesis

Ya-Jie Zhang , Shifeng Pan , Lei He , Zhen-Hua Ling This is my paper

classification 💻 cs.CL cs.SDeess.AS

keywords stylecontrolmodelrepresentationspeechtransferend-to-endgood

0 comments

read the original abstract

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good properties such as disentangling, scaling, and combination, which makes it easy for style control. Style transfer can be achieved in this framework by first inferring style representation through the recognition network of VAE, then feeding it into TTS network to guide the style in synthesizing speech. To avoid Kullback-Leibler (KL) divergence collapse in training, several techniques are adopted. Finally, the proposed model shows good performance of style control and outperforms Global Style Token (GST) model in ABX preference tests on style transfer.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech
eess.AS 2019-07 unverdicted novelty 6.0

Decouples prosody alignment via pre-computed phoneme timestamps and adds VAE to achieve robust fine-grained prosody transfer in single-speaker neural TTS from unseen speakers.