How to Improve Your Speaker Embeddings Extractor in Generic Toolkits

Hossein Zeinali; Jan Cernocky; Johan Rohdin; Lukas Burget; Themos Stafylakis

arxiv: 1811.02066 · v1 · pith:2D5GRZ7Wnew · submitted 2018-11-05 · 💻 cs.SD · cs.CL· eess.AS

How to Improve Your Speaker Embeddings Extractor in Generic Toolkits

Hossein Zeinali , Lukas Burget , Johan Rohdin , Themos Stafylakis , Jan Cernocky This is my paper

classification 💻 cs.SD cs.CLeess.AS

keywords speakerembeddingsgenericimplementationmethodadditionalternativeanticipate

0 comments

read the original abstract

Recently, speaker embeddings extracted with deep neural networks became the state-of-the-art method for speaker verification. In this paper we aim to facilitate its implementation on a more generic toolkit than Kaldi, which we anticipate to enable further improvements on the method. We examine several tricks in training, such as the effects of normalizing input features and pooled statistics, different methods for preventing overfitting as well as alternative non-linearities that can be used instead of Rectifier Linear Units. In addition, we investigate the difference in performance between TDNN and CNN, and between two types of attention mechanism. Experimental results on Speaker in the Wild, SRE 2016 and SRE 2018 datasets demonstrate the effectiveness of the proposed implementation.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Self Multi-Head Attention for Speaker Recognition
cs.SD 2019-06 unverdicted novelty 6.0

Self multi-head attention applied after CNN encoding of spectrograms outperforms temporal and statistical pooling for speaker verification on VoxCeleb1 with 18% relative EER reduction.