Pith. sign in

REVIEW 1 cited by

Transforming the Embeddings: A Lightweight Technique for Speech Emotion Recognition Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.18640 v1 pith:XH43SOJM submitted 2023-05-29 eess.AS

classification eess.AS
keywords embeddingsrecognitionspeakerspeechapproachemotionfeaturesinput
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech emotion recognition (SER) is a field that has drawn a lot of attention due to its applications in diverse fields. A current trend in methods used for SER is to leverage embeddings from pre-trained models (PTMs) as input features to downstream models. However, the use of embeddings from speaker recognition PTMs hasn't garnered much focus in comparison to other PTM embeddings. To fill this gap and in order to understand the efficacy of speaker recognition PTM embeddings, we perform a comparative analysis of five PTM embeddings. Among all, x-vector embeddings performed the best possibly due to its training for speaker recognition leading to capturing various components of speech such as tone, pitch, etc. Our modeling approach which utilizes x-vector embeddings and mel-frequency cepstral coefficients (MFCC) as input features is the most lightweight approach while achieving comparable accuracy to previous state-of-the-art (SOTA) methods in the CREMA-D benchmark.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

    eess.AS 2025-06 conditional novelty 4.0 of 10

    Audio-Mamba models outperform attention-based models such as WavLM and HuBERT on non-verbal emotion recognition, and a Renyi-divergence fusion approach called RENO further improves accuracy.

Pith tools