Pith. sign in

REVIEW 1 cited by

Emotion Recognition from Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.10458 v1 pith:2J62NL5V submitted 2019-12-22 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords emotionfeaturesaudiorecognitionspeechclassificationlog-melnetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we conduct an extensive comparison of various approaches to speech based emotion recognition systems. The analyses were carried out on audio recordings from Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). After pre-processing the raw audio files, features such as Log-Mel Spectrogram, Mel-Frequency Cepstral Coefficients (MFCCs), pitch and energy were considered. The significance of these features for emotion classification was compared by applying methods such as Long Short Term Memory (LSTM), Convolutional Neural Networks (CNNs), Hidden Markov Models (HMMs) and Deep Neural Networks (DNNs). On the 14-class (2 genders x 7 emotions) classification task, an accuracy of 68% was achieved with a 4-layer 2 dimensional CNN using the Log-Mel Spectrogram features. We also observe that, in emotion recognition, the choice of audio features impacts the results much more than the model complexity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating the Impact of Word Informativeness on Speech Emotion Recognition

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Using GPT-2 surprisal to select a few words per sentence for acoustic feature extraction gives a small accuracy improvement over whole-utterance features in RAVDESS speech emotion recognition, though the corpus's two ...

Pith tools