Pith. sign in

REVIEW 3 cited by

LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11468 v1 pith:KYM24DLW submitted 2025-01-20 eess.AS cs.SD

classification eess.AScs.SD
keywords modelemotionrecognitionconversationsproposedspeechtranscriptsdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotion recognition in conversations (ERC) is challenging due to the multimodal nature of the emotion expression. In this paper, we propose to pretrain a text-based recognition model from unsupervised speech transcripts with LLM guidance. These transcriptions are obtained from a raw speech dataset with a pre-trained ASR system. A text LLM model is queried to provide pseudo-labels for these transcripts, and these pseudo-labeled transcripts are subsequently used for learning an utterance level text-based emotion recognition model. We use the utterance level text embeddings for emotion recognition in conversations along with speech embeddings obtained from a recently proposed pre-trained model. A hierarchical way of training the speech-text model is proposed, keeping in mind the conversational nature of the dataset. We perform experiments on three established datasets, namely, IEMOCAP, MELD, and CMU- MOSI, where we illustrate that the proposed model improves over other benchmarks and achieves state-of-the-art results on two out of these three datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Using Rank-N-Contrast loss on valence-arousal rankings instead of symmetric cross-entropy improves ordinal consistency and cross-modal alignment for emotional speech and text.

  2. Multimodal Emotion Recognition in Conversations: A Survey of Methods, Trends, Challenges and Prospects

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A structured review of multimodal emotion recognition in conversations, covering datasets, feature processing, methods, and open challenges, with emphasis on recent LLM-based approaches.

  3. ABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge

    cs.SD 2025-05 conditional novelty 4.0 of 10

    Abhinaya, an ensemble of fine-tuned SSL, SLLM, and LLM models with majority voting, achieved state-of-the-art macro-F1 (44.02%) on the Interspeech 2025 naturalistic speech emotion recognition test set.

Pith tools