Pith. sign in

REVIEW 1 cited by

Leveraging Content and Acoustic Representations for Speech Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05566 v3 pith:N2WGP6TL submitted 2024-09-09 eess.AS

classification eess.AS
keywords speechacousticemotionmodelrepresentationstrainedcarecontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of labeled datasets further complicates the challenge where large models are prone to over-fitting. In this paper, we propose CARE (Content and Acoustic Representations of Emotions), where we design a dual encoding scheme which emphasizes semantic and acoustic factors of speech. While the semantic encoder is trained using distillation from utterance-level text representations, the acoustic encoder is trained to predict low-level frame-wise features of the speech signal. The proposed dual encoding scheme is a base-sized model trained only on unsupervised raw speech. With a simple light-weight classification model trained on the downstream task, we show that the CARE embeddings provide effective emotion recognition on a variety of datasets. We compare the proposal with several other self-supervised models as well as recent large-language model based approaches. In these evaluations, the proposed CARE is shown to be the best performing model based on average performance across 8 diverse datasets. We also conduct several ablation studies to analyze the importance of various design choices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Abhinaya, an ensemble of fine-tuned SSL, SLLM, and LLM models with majority voting, achieved state-of-the-art macro-F1 (44.02%) on the Interspeech 2025 naturalistic speech emotion recognition test set.

Pith tools