Pith. sign in

REVIEW 2 cited by

CNN+LSTM Architecture for Speech Emotion Recognition with Data Augmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.05630 v2 pith:4VZ6ALLA submitted 2018-02-15 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords accuracyarchitectureaugmentationdataemotionslayersrecurrentspeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we design a neural network for recognizing emotions in speech, using the IEMOCAP dataset. Following the latest advances in audio analysis, we use an architecture involving both convolutional layers, for extracting high-level features from raw spectrograms, and recurrent ones for aggregating long-term dependencies. We examine the techniques of data augmentation with vocal track length perturbation, layer-wise optimizer adjustment, batch normalization of recurrent layers and obtain highly competitive results of 64.5% for weighted accuracy and 61.7% for unweighted accuracy on four emotions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. THAI Speech Emotion Recognition (THAI-SER) corpus

    cs.SD 2025-07 conditional novelty 7.0 of 10

    THAI-SER is the first sizeable Thai speech emotion recognition corpus, with 41.6 hours of acted and elicited speech, 27,854 utterances, and crowdsourced labels for five emotions.

  2. Explainable Lightweight Compact Deep Models for Speech Emotion Recognition

    cs.SD 2026-07 conditional novelty 3.0 of 10

    A 33k-parameter CNN with attentive statistics pooling and Grad-CAM reaches 96.9% accuracy on SAVEE speech emotion recognition, but the evaluation rests on one speaker-independent split.

Pith tools