Pith. sign in

REVIEW 2 cited by

Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07078 v1 pith:XOT77KCG submitted 2024-09-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords emotionrecognitionmultimodalclipmodalitymodelpromptvideos
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present our solution for the Second Multimodal Emotion Recognition Challenge Track 1(MER2024-SEMI). To enhance the accuracy and generalization performance of emotion recognition, we propose several methods for Multimodal Emotion Recognition. Firstly, we introduce EmoVCLIP, a model fine-tuned based on CLIP using vision-language prompt learning, designed for video-based emotion recognition tasks. By leveraging prompt learning on CLIP, EmoVCLIP improves the performance of pre-trained CLIP on emotional videos. Additionally, to address the issue of modality dependence in multimodal fusion, we employ modality dropout for robust information fusion. Furthermore, to aid Baichuan in better extracting emotional information, we suggest using GPT-4 as the prompt for Baichuan. Lastly, we utilize a self-training strategy to leverage unlabeled videos. In this process, we use unlabeled videos with high-confidence pseudo-labels generated by our model and incorporate them into the training set. Experimental results demonstrate that our model ranks 1st in the MER2024-SEMI track, achieving an accuracy of 90.15% on the test set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VEGA aligns multimodal emotion features with CLIP-derived visual emotion prototypes and reports SOTA on IEMOCAP and MELD.

  2. ECMF: Enhanced Cross-Modal Fusion for Multimodal Emotion Recognition in MER-SEMI Challenge

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A multimodal fusion pipeline with pretrained encoders, LLM-enriched text, attention fusion, label refinement, and ensembling reaches 87.49% weighted F-score on MER2025-SEMI.

Pith tools