Adding speaker-contrastive or BYOL self-supervised pretraining to a Whisper-based model improves low-resource speech emotion recognition on Urdu, German, and Bangla.
Speech Emotion Recognition using Support Vector Machine
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
In this project, we aim to classify the speech taken as one of the four emotions namely, sadness, anger, fear and happiness. The samples that have been taken to complete this project are taken from Linguistic Data Consortium (LDC) and UGA database. The important characteristics determined from the samples are energy, pitch, MFCC coefficients, LPCC coefficients and speaker rate. The classifier used to classify these emotional states is Support Vector Machine (SVM) and this is done using two classification strategies: One against All (OAA) and Gender Dependent Classification. Furthermore, a comparative analysis has been conducted between the two and LPCC and MFCC algorithms as well.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Learning More with Less: Self-Supervised Approaches for Low-Resource Speech Emotion Recognition
Adding speaker-contrastive or BYOL self-supervised pretraining to a Whisper-based model improves low-resource speech emotion recognition on Urdu, German, and Bangla.