Pith. sign in

REVIEW 2 cited by

CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.15432 v2 pith:BRPU57RE submitted 2023-07-28 cs.CL

classification cs.CL
keywords emotion-shiftinformationcfn-esaemotionmultimodalcross-modalemotionalmodalities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal emotion recognition in conversation (ERC) has garnered growing attention from research communities in various fields. In this paper, we propose a Cross-modal Fusion Network with Emotion-Shift Awareness (CFN-ESA) for ERC. Extant approaches employ each modality equally without distinguishing the amount of emotional information in these modalities, rendering it hard to adequately extract complementary information from multimodal data. To cope with this problem, in CFN-ESA, we treat textual modality as the primary source of emotional information, while visual and acoustic modalities are taken as the secondary sources. Besides, most multimodal ERC models ignore emotion-shift information and overfocus on contextual information, leading to the failure of emotion recognition under emotion-shift scenario. We elaborate an emotion-shift module to address this challenge. CFN-ESA mainly consists of unimodal encoder (RUME), cross-modal encoder (ACME), and emotion-shift module (LESM). RUME is applied to extract conversation-level contextual emotional cues while pulling together data distributions between modalities; ACME is utilized to perform multimodal interaction centered on textual modality; LESM is used to model emotion shift and capture emotion-shift information, thereby guiding the learning of the main task. Experimental results demonstrate that CFN-ESA can effectively promote performance for ERC and remarkably outperform state-of-the-art models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmoEUS: Uncertainty Supervision for Multimodal Emotion Recognition in Conversation

    cs.MM 2026-07 conditional novelty 6.0 of 10

    Modeling each modality as a Gaussian and supervising its variance with the 2-Wasserstein distance to emotion cluster centers improves IEMOCAP/MELD accuracy by about 0.5-0.8 points over listed baselines.

  2. EII-SCL: Harnessing Emotional Inertia for Multimodal Emotion Recognition in Conversation

    cs.MM 2026-07 conditional novelty 5.0 of 10

    A plug-in contrastive loss using speaker-local 'emotional inertia' hard negatives improves multimodal emotion-recognition accuracy and F1 on IEMOCAP and MELD by about 0.5–2 points.

Pith tools