CycleGAN-synthesized emotional speech, added to training data, improves speaker verification on emotional utterances by up to 3.64% relative EER on one internal dataset.
Improving speaker verification robustness with synthetic emotional utterances
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
A speaker verification (SV) system offers an authentication service designed to confirm whether a given speech sample originates from a specific speaker. This technology has paved the way for various personalized applications that cater to individual preferences. A noteworthy challenge faced by SV systems is their ability to perform consistently across a range of emotional spectra. Most existing models exhibit high error rates when dealing with emotional utterances compared to neutral ones. Consequently, this phenomenon often leads to missing out on speech of interest. This issue primarily stems from the limited availability of labeled emotional speech data, impeding the development of robust speaker representations that encompass diverse emotional states. To address this concern, we propose a novel approach employing the CycleGAN framework to serve as a data augmentation method. This technique synthesizes emotional speech segments for each specific speaker while preserving the unique vocal identity. Our experimental findings underscore the effectiveness of incorporating synthetic emotional data into the training process. The models trained using this augmented dataset consistently outperform the baseline models on the task of verifying speakers in emotional speech scenarios, reducing equal error rate by as much as 3.64% relative.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Improving speaker verification robustness with synthetic emotional utterances
CycleGAN-synthesized emotional speech, added to training data, improves speaker verification on emotional utterances by up to 3.64% relative EER on one internal dataset.