Pith. sign in

REVIEW 1 cited by

Speech Emotion Detection Based on MFCC and CNN-LSTM Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.10666 v1 pith:KQHP6KTN submitted 2025-01-18 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords emotionsfeaturesarchitectureaudiodetectionemotionaccuracyanger
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotion detection techniques have been applied to multiple cases mainly from facial image features and vocal audio features, of which the latter aspect is disputed yet not only due to the complexity of speech audio processing but also the difficulties of extracting appropriate features. Part of the SAVEE and RAVDESS datasets are selected and combined as the dataset, containing seven sorts of common emotions (i.e. happy, neutral, sad, anger, disgust, fear, and surprise) and thousands of samples. Based on the Librosa package, this paper processes the initial audio input into waveplot and spectrum for analysis and concentrates on multiple features including MFCC as targets for feature extraction. The hybrid CNN-LSTM architecture is adopted by virtue of its strong capability to deal with sequential data and time series, which mainly consists of four convolutional layers and three long short-term memory layers. As a result, the architecture achieved an accuracy of 61.07% comprehensively for the test set, among which the detection of anger and neutral reaches a performance of 75.31% and 71.70% respectively. It can also be concluded that the classification accuracy is dependent on the properties of emotion to some extent, with frequently-used and distinct-featured emotions having less probability to be misclassified into other categories. Emotions like surprise whose meaning depends on the specific context are more likely to confuse with positive or negative emotions, and negative emotions also have a possibility to get mixed with each other.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Explainable Lightweight Compact Deep Models for Speech Emotion Recognition

    cs.SD 2026-07 conditional novelty 3.0 of 10

    A 33k-parameter CNN with attentive statistics pooling and Grad-CAM reaches 96.9% accuracy on SAVEE speech emotion recognition, but the evaluation rests on one speaker-independent split.

Pith tools