Pith. sign in

REVIEW 2 cited by

MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.05746 v1 pith:FMGOJNA7 submitted 2024-07-08 cs.AI cs.SDeess.AS

classification cs.AIcs.SDeess.AS
keywords emotionalspeechchallengeemotionmsp-podcastrecognitionstatessystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we detail our submission to the 2024 edition of the MSP-Podcast Speech Emotion Recognition (SER) Challenge. This challenge is divided into two distinct tasks: Categorical Emotion Recognition and Emotional Attribute Prediction. We concentrated our efforts on Task 1, which involves the categorical classification of eight emotional states using data from the MSP-Podcast dataset. Our approach employs an ensemble of models, each trained independently and then fused at the score level using a Support Vector Machine (SVM) classifier. The models were trained using various strategies, including Self-Supervised Learning (SSL) fine-tuning across different modalities: speech alone, text alone, and a combined speech and text approach. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. Thus, the system obtained F1-macro of 0.35\% on development set.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset

    eess.AS 2025-05 conditional novelty 6.0 of 10

    EmoCorrector retrieves emotional speech samples matching the edited text and uses them to post-correct the emotion of TSE output, supported by the new synthetic ECD-TSE dataset.

  2. "How to Explore Biases in Speech Emotion AI with Users?" A Speech-Emotion-Acting Study Exploring Age and Language Biases

    cs.HC 2025-07 conditional novelty 5.0 of 10

    In a 24-person Danish study, a speech emotion recognition model showed no significant age or language differences in recognizing deliberately acted happy, sad, angry, and calm speech, though high-arousal emotions were...

Pith tools