AttuneBench introduces a multi-turn conversation benchmark using participant annotations to evaluate LLM emotional intelligence, finding that model performance on emotion recognition, behavior classification, preference prediction, and response quality are largely independent.
Liu, Jinfeng Zhou, Alvionna S
5 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
ESC uses emotional cues triggered by an external verifier to enable training-free self-correction in VLMs, improving reliability on safety, hallucination, and reasoning benchmarks.
Authors build an emotional intensity dataset and fine-tune generative LLMs to predict continuous 0-100 scores, claiming outperformance over classification baselines plus generalization to sentiment and arousal.
MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.
A holistic survey of affective computing for intelligent agents covering emotion understanding via multimodal data, affective cognition, emotional expression synthesis, key challenges, and future directions emphasizing generative technologies.
citing papers explorer
-
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
AttuneBench introduces a multi-turn conversation benchmark using participant annotations to evaluate LLM emotional intelligence, finding that model performance on emotion recognition, behavior classification, preference prediction, and response quality are largely independent.
-
ESC: Emotional Self-Correction for Reliable Vision-Language Models
ESC uses emotional cues triggered by an external verifier to enable training-free self-correction in VLMs, improving reliability on safety, hallucination, and reasoning benchmarks.
-
Beyond Sentiment Classification: A Generative Framework for Emotion Intensity Evaluation in Text
Authors build an emotional intensity dataset and fine-tune generative LLMs to predict continuous 0-100 scores, claiming outperformance over classification baselines plus generalization to sentiment and arousal.
-
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.
-
Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects
A holistic survey of affective computing for intelligent agents covering emotion understanding via multimodal data, affective cognition, emotional expression synthesis, key challenges, and future directions emphasizing generative technologies.