Pith. sign in

REVIEW 1 cited by

Self-attention encoding and pooling for speaker recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.01077 v1 pith:NBKR2C5Z submitted 2020-08-03 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords speakerself-attentionarchitectureapplicationsapproachefficientencodingfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The computing power of mobile devices limits the end-user applications in terms of storage size, processing, memory and energy consumption. These limitations motivate researchers for the design of more efficient deep models. On the other hand, self-attention networks based on Transformer architecture have attracted remarkable interests due to their high parallelization capabilities and strong performance on a variety of Natural Language Processing (NLP) applications. Inspired by the Transformer, we propose a tandem Self-Attention Encoding and Pooling (SAEP) mechanism to obtain a discriminative speaker embedding given non-fixed length speech utterances. SAEP is a stack of identical blocks solely relied on self-attention and position-wise feed-forward networks to create vector representation of speakers. This approach encodes short-term speaker spectral features into speaker embeddings to be used in text-independent speaker verification. We have evaluated this approach on both VoxCeleb1 & 2 datasets. The proposed architecture is able to outperform the baseline x-vector, and shows competitive performance to some other benchmarks based on convolutions, with a significant reduction in model size. It employs 94%, 95%, and 73% less parameters compared to ResNet-34, ResNet-50, and x-vector, respectively. This indicates that the proposed fully attention based architecture is more efficient in extracting time-invariant features from speaker utterances.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Cross-Corpus Speech Emotion Recognition Method Based on Supervised Contrastive Learning

    cs.SD 2024-11 conditional novelty 4.0 of 10

    A supervised contrastive learning fine-tuning stage on two speech emotion datasets improves cross-corpus emotion recognition accuracy over direct fine-tuning.

Pith tools