Pith. sign in

REVIEW 2 cited by

COGMEN: COntextualized GNN based Multimodal Emotion recognitioN

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.02455 v1 pith:AB6X36YC submitted 2022-05-05 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords informationemotionsmodelcogmencontextualizedconversationemotionglobal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotions are an inherent part of human interactions, and consequently, it is imperative to develop AI systems that understand and recognize human emotions. During a conversation involving various people, a person's emotions are influenced by the other speaker's utterances and their own emotional state over the utterances. In this paper, we propose COntextualized Graph Neural Network based Multimodal Emotion recognitioN (COGMEN) system that leverages local information (i.e., inter/intra dependency between speakers) and global information (context). The proposed model uses Graph Neural Network (GNN) based architecture to model the complex dependencies (local and global information) in a conversation. Our model gives state-of-the-art (SOTA) results on IEMOCAP and MOSEI datasets, and detailed ablation experiments show the importance of modeling information at both levels.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VEGA aligns multimodal emotion features with CLIP-derived visual emotion prototypes and reports SOTA on IEMOCAP and MELD.

  2. Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion

    cs.MM 2025-07 conditional novelty 4.0 of 10

    Sync-TVA reports modest accuracy and weighted-F1 improvements over prior graph-based models on MELD and IEMOCAP, using modality-specific enhancement and cross-modal graph fusion.

Pith tools