Pith. sign in

REVIEW 5 cited by

AffectGPT: Dataset and Framework for Explainable Multimodal Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07653 v1 pith:CCNAEKEI submitted 2024-07-10 cs.HC

AffectGPT: Dataset and Framework for Explainable Multimodal Emotion Recognition

classification cs.HC
keywords datasetaffectgptannotationemotionmultimodalrecognitioncostemer
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Explainable Multimodal Emotion Recognition (EMER) is an emerging task that aims to achieve reliable and accurate emotion recognition. However, due to the high annotation cost, the existing dataset (denoted as EMER-Fine) is small, making it difficult to perform supervised training. To reduce the annotation cost and expand the dataset size, this paper reviews the previous dataset construction process. Then, we simplify the annotation pipeline, avoid manual checks, and replace the closed-source models with open-source models. Finally, we build \textbf{EMER-Coarse}, a coarsely-labeled dataset containing large-scale samples. Besides the dataset, we propose a two-stage training framework \textbf{AffectGPT}. The first stage exploits EMER-Coarse to learn a coarse mapping between multimodal inputs and emotion-related descriptions; the second stage uses EMER-Fine to better align with manually-checked results. Experimental results demonstrate the effectiveness of our proposed method on the challenging EMER task. To facilitate further research, we will make the code and dataset available at: https://github.com/zeroQiaoba/AffectGPT.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

    cs.AI 2026-07 conditional novelty 7.0

    A new in-cabin dataset pairs RGB/IR video, audio, and Chinese dialogue text with emotion, fatigue, and distraction labels, and baselines show fusion beats single modalities on the Chinese partition.

  2. MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models

    cs.CV 2026-05 conditional novelty 7.0

    MultiEmo-Bench supplies 10,344 images with aggregated multi-label emotion votes from 20 annotators each to evaluate MLLMs on dominant emotion and full distribution prediction.

  3. DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning

    cs.CV 2026-03 unverdicted novelty 7.0

    A new 1695-sample multicultural dataset plus two modules for stable multimodal fusion and modality consistency yield state-of-the-art deception detection with cross-cultural transfer.

  4. DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning

    cs.CV 2026-03 conditional novelty 6.0

    Schema-constrained MLLM reports plus SICS/DMC modules and the T4-Deception dataset yield SOTA in-domain, cross-domain, and cross-cultural multimodal deception detection.

  5. Multimodal Group Emotion Recognition In-the-Wild Towards a Privacy-Safe Non-Individual Approach

    cs.CV 2026-05 unverdicted novelty 4.0

    Proposes cross-attention audio-video fusion and VE-MD latent-space models for group emotion recognition that avoid individual cues and report competitive performance via ablation studies on synthetic and real data.