REVIEW 7 cited by
OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-appraisal nature of human emotional experiences, as demonstrated by studies in psychology and cognitive science. To overcome this limitation, we advocate for introducing the concept of open vocabulary into MER. This paradigm shift aims to enable models to predict emotions beyond a fixed label space, accommodating a flexible set of categories to better reflect the nuanced spectrum of human emotions. To achieve this, we propose a novel paradigm: Open-Vocabulary MER (OV-MER), which enables emotion prediction without being confined to predefined spaces. However, constructing a dataset that encompasses the full range of emotions for OV-MER is practically infeasible; hence, we present a comprehensive solution including a newly curated database, novel evaluation metrics, and a preliminary benchmark. By advancing MER from basic emotions to more nuanced and diverse emotional states, we hope this work can inspire the next generation of MER, enhancing its generalizability and applicability in real-world scenarios. Code and dataset are available at: https://github.com/zeroQiaoba/AffectGPT.
Forward citations
Cited by 7 Pith papers
-
E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment
A new video benchmark jointly evaluates expressed and evoked emotion understanding in multimodal LLMs using perception, open-vocabulary recognition, and VAD rating tasks, with Bayesian pairwise alignment for scalable ...
-
Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
A new large-scale facial emotion caption dataset and a global-local contrastive training framework with positive mining improve zero-shot facial expression recognition.
-
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
Chain-Talker predicts empathetic captions from dialogue history, generates semantic speech codes, and renders expressive speech, outperforming prior conversational speech synthesis models in subjective and objective tests.
-
MER 2025: When Affective Computing Meets Large Language Models
A four-track multimodal emotion recognition benchmark for LLM-driven methods, with datasets, baselines, and evaluation protocols for categorical, fine-grained, descriptive, and personality tasks.
-
AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
AffectGPT introduces a 115K-sample descriptive emotion dataset, a pre-fusion multimodal LLM, and a unified benchmark, reporting large gains over existing MLLMs.
-
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
Multimodal LLMs frequently accept hallucinated emotion claims on a new adversarial benchmark, with the worst failures on image, audio, and video perception rather than on textbook emotion knowledge.
-
Affective-CARA: A Knowledge Graph Driven Framework for Culturally Adaptive Emotional Intelligence in HCI
Affective-CARA integrates a hyperbolic culture emotion graph, a PPO-style reward optimizer, and a response mediator for culturally adaptive chatbot replies, but its headline metrics do not measure the claimed system behavior.
Discussion (0). Continue with ORCID to comment.