REVIEW 10 cited by
EmoBench: Evaluating the Emotional Intelligence of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion recognition, neglecting essential EI capabilities such as emotion regulation and thought facilitation through emotion understanding; second, they are primarily constructed from existing datasets, which include frequent patterns, explicit information, and annotation errors, leading to unreliable evaluation. We propose EmoBench, a benchmark that draws upon established psychological theories and proposes a comprehensive definition for machine EI, including Emotional Understanding and Emotional Application. EmoBench includes a set of 400 hand-crafted questions in English and Chinese, which are meticulously designed to require thorough reasoning and understanding. Our findings reveal a considerable gap between the EI of existing LLMs and the average human, highlighting a promising direction for future research. Our code and data are publicly available at https://github.com/Sahandfer/EmoBench.
Forward citations
Cited by 10 Pith papers
-
Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding
A new multi-label emotion benchmark for four Ethiopian languages shows that fine-tuned encoder-only models outperform zero-shot and few-shot large language models, with large gaps between resource-rich and resource-po...
-
CharacterBench: Benchmarking Character Customization of Large Language Models
CharacterBench provides a large Chinese-English benchmark and a fine-tuned judge model for measuring 11 dimensions of LLM character customization, with reported human-correlation improvements over GPT-4.
-
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.
-
An LLM's Apology: Outsourcing Awkwardness in the Age of AI
Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.
-
MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
MotiveBench shows that large language models still lag human consensus on motivational reasoning, with GPT-4o scoring 80.89 percent and chain-of-thought prompting usually reducing accuracy.
-
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.
-
EmoAssist: Emotional Assistant for Visual Impairment Community
A fine-tuned LLaVA model, trained with preference optimization on 800 emotional image-QA examples, beats GPT-4o on a new empathy-focused benchmark for visual impairment assistance.
-
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.
-
Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations
Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.
-
A Survey on Human-Centric LLMs
A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.
Discussion (0). Continue with ORCID to comment.