Pith. sign in

REVIEW 10 cited by

EmoBench: Evaluating the Emotional Intelligence of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12071 v3 pith:YUDNI7E7 submitted 2024-02-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords emobenchemotionalemotionexistingunderstandingbenchmarkscomprehensiveevaluating
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion recognition, neglecting essential EI capabilities such as emotion regulation and thought facilitation through emotion understanding; second, they are primarily constructed from existing datasets, which include frequent patterns, explicit information, and annotation errors, leading to unreliable evaluation. We propose EmoBench, a benchmark that draws upon established psychological theories and proposes a comprehensive definition for machine EI, including Emotional Understanding and Emotional Application. EmoBench includes a set of 400 hand-crafted questions in English and Chinese, which are meticulously designed to require thorough reasoning and understanding. Our findings reveal a considerable gap between the EI of existing LLMs and the average human, highlighting a promising direction for future research. Our code and data are publicly available at https://github.com/Sahandfer/EmoBench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding

    cs.CL 2024-12 conditional novelty 7.0 of 10

    A new multi-label emotion benchmark for four Ethiopian languages shows that fine-tuned encoder-only models outperform zero-shot and few-shot large language models, with large gaps between resource-rich and resource-po...

  2. CharacterBench: Benchmarking Character Customization of Large Language Models

    cs.CL 2024-12 conditional novelty 7.0 of 10

    CharacterBench provides a large Chinese-English benchmark and a fine-tuned judge model for measuring 11 dimensions of LLM character customization, with reported human-correlation improvements over GPT-4.

  3. MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.

  4. An LLM's Apology: Outsourcing Awkwardness in the Age of AI

    cs.CY 2025-06 reject novelty 6.0 of 10

    Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.

  5. MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MotiveBench shows that large language models still lag human consensus on motivational reasoning, with GPT-4o scoring 80.89 percent and chain-of-thought prompting usually reducing accuracy.

  6. LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.

  7. EmoAssist: Emotional Assistant for Visual Impairment Community

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A fine-tuned LLaVA model, trained with preference optimization on 800 emotional image-QA examples, beats GPT-4o on a new empathy-focused benchmark for visual impairment assistance.

  8. MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MEMO-Bench scores 12 text-to-image models and 16 multimodal LLMs on emotion generation and recognition, finding stronger performance on positive emotions and weak fine-grained intensity estimation.

  9. Boosting Self-Efficacy and Performance of Large Language Models via Verbal Efficacy Stimulations

    cs.CL 2025-02 conditional novelty 4.0 of 10

    Emotionally styled verbal prompts (encouraging, provocative, critical) modestly improve zero-shot LLM accuracy on many tasks, with the best style varying by model and task zone.

  10. A Survey on Human-Centric LLMs

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.

Pith tools