Pith. sign in

REVIEW 5 cited by

Explainable Multimodal Emotion Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.15401 v6 pith:V7WOGIYZ submitted 2023-06-27 cs.MM cs.HC

Explainable Multimodal Emotion Recognition

classification cs.MM cs.HC
keywords multimodalemotionlabelsrecognitiontaskemerlabelbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Multimodal emotion recognition is an important research topic in artificial intelligence, whose main goal is to integrate multimodal clues to identify human emotional states. Current works generally assume accurate labels for benchmark datasets and focus on developing more effective architectures. However, emotion annotation relies on subjective judgment. To obtain more reliable labels, existing datasets usually restrict the label space to some basic categories, then hire plenty of annotators and use majority voting to select the most likely label. However, this process may result in some correct but non-candidate or non-majority labels being ignored. To ensure reliability without ignoring subtle emotions, we propose a new task called ``Explainable Multimodal Emotion Recognition (EMER)''. Unlike traditional emotion recognition, EMER takes a step further by providing explanations for these predictions. Through this task, we can extract relatively reliable labels since each label has a certain basis. Meanwhile, we borrow large language models (LLMs) to disambiguate unimodal clues and generate more complete multimodal explanations. From them, we can extract richer emotions in an open-vocabulary manner. This paper presents our initial attempt at this task, including introducing a new dataset, establishing baselines, and defining evaluation metrics. In addition, EMER can serve as a benchmark task to evaluate the audio-video-text understanding performance of multimodal LLMs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AffectVerse: Emotional World Models for Multimodal Affective Computing

    cs.CV 2026-05 unverdicted novelty 7.0

    AffectVerse improves multimodal emotion recognition by at least 2.57% on nine benchmarks through an Emotion World Module that performs short-horizon latent affective prediction via cross-modal temporal imagination and...

  2. AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition

    cs.HC 2026-05 unverdicted novelty 7.0

    AffectGPT-RL applies reinforcement learning to optimize non-differentiable emotion wheel metrics in open-vocabulary multimodal emotion recognition, yielding performance gains and state-of-the-art results on basic emot...

  3. Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy

    cs.AI 2026-03 unverdicted novelty 6.0

    Nano-EmoX is a compact 2.2B multimodal model that unifies six core affective tasks across perception, understanding, and interaction levels via a curriculum framework, achieving competitive benchmark performance.

  4. OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

    cs.CV 2026-06 conditional novelty 5.0

    Rationale-privileged on-policy self-distillation reaches 84.19 mean on MER-UniBench by scoring student rollouts with a local teacher that alone sees frontier-generated multimodal evidence.

  5. MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding

    cs.HC 2026-04 unverdicted novelty 4.0

    MER2026 defines four tracks to advance generative emotion understanding from individual basic labels to dyadic, fine-grained, preference, and physiological scenarios.