EmoReg controls emotional intensity in diffusion-based voice conversion by scaling a PCA-projected direction vector in a fine-tuned self-supervised emotion embedding space.
CycleTransGAN-EVC: A CycleGAN-based Emotional Voice Conversion Model with Transformer
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this study, we explore the transformer's ability to capture intra-relations among frames by augmenting the receptive field of models. Concretely, we propose a CycleGAN-based model with the transformer and investigate its ability in the emotional voice conversion task. In the training procedure, we adopt curriculum learning to gradually increase the frame length so that the model can see from the short segment till the entire speech. The proposed method was evaluated on the Japanese emotional speech dataset and compared to several baselines (ACVAE, CycleGAN) with objective and subjective evaluations. The results show that our proposed model is able to convert emotion with higher strength and quality.
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
EmoReg controls emotional intensity in diffusion-based voice conversion by scaling a PCA-projected direction vector in a fine-tuned self-supervised emotion embedding space.