Pith. sign in

REVIEW 2 cited by

Revisiting Multimodal Emotion Recognition in Conversation from the Perspective of Graph Spectrum

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.17862 v2 pith:CGFD2PGV submitted 2024-04-27 cs.CL

classification cs.CL
keywords informationmultimodalgraphemotiongs-mcclow-frequencyconversationrecognition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Efficiently capturing consistent and complementary semantic features in a multimodal conversation context is crucial for Multimodal Emotion Recognition in Conversation (MERC). Existing methods mainly use graph structures to model dialogue context semantic dependencies and employ Graph Neural Networks (GNN) to capture multimodal semantic features for emotion recognition. However, these methods are limited by some inherent characteristics of GNN, such as over-smoothing and low-pass filtering, resulting in the inability to learn long-distance consistency information and complementary information efficiently. Since consistency and complementarity information correspond to low-frequency and high-frequency information, respectively, this paper revisits the problem of multimodal emotion recognition in conversation from the perspective of the graph spectrum. Specifically, we propose a Graph-Spectrum-based Multimodal Consistency and Complementary collaborative learning framework GS-MCC. First, GS-MCC uses a sliding window to construct a multimodal interaction graph to model conversational relationships and uses efficient Fourier graph operators to extract long-distance high-frequency and low-frequency information, respectively. Then, GS-MCC uses contrastive learning to construct self-supervised signals that reflect complementarity and consistent semantic collaboration with high and low-frequency signals, thereby improving the ability of high and low-frequency information to reflect real emotions. Finally, GS-MCC inputs the collaborative high and low-frequency information into the MLP network and softmax function for emotion prediction. Extensive experiments have proven the superiority of the GS-MCC architecture proposed in this paper on two benchmark data sets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generation or Judgement? A Paradigm Perspective on LLM-Based Emotion-Cause Pair Extraction in Conversation

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Pair-level judgement consistently outperforms dialogue-level generation for LLM-based ECPEC because models recognize pairs under explicit queries but fail at set-level discovery and shared-threshold decisions.

  2. A Novel Approach to for Multimodal Emotion Recognition : Multimodal semantic information fusion

    cs.CV 2025-02 reject novelty 4.0 of 10

    DeepMSI-MER combines contrastive learning with semantic-guided visual compression and reports improved emotion recognition accuracy on IEMOCAP and MELD.

Pith tools