Pith. sign in

Multimodal lan- guage analysis in the wild: CMU-MOSEI dataset and inter- pretable dynamic fusion graph

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

dataset 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

dataset 1

polarities

use dataset 1

representative citing papers

Leveraging CLIP Encoder for Multimodal Emotion Recognition

cs.CV · 2025-06-01 · conditional · novelty 6.0

MER-CLIP uses a frozen CLIP text encoder as a label encoder and a cross-modal decoder to align video, audio, and text features, setting new state-of-the-art results on CMU-MOSI and CMU-MOSEI.

citing papers explorer

Showing 1 of 1 citing paper.

  • Leveraging CLIP Encoder for Multimodal Emotion Recognition cs.CV · 2025-06-01 · conditional · none · ref 2

    MER-CLIP uses a frozen CLIP text encoder as a label encoder and a cross-modal decoder to align video, audio, and text features, setting new state-of-the-art results on CMU-MOSI and CMU-MOSEI.