Pith. sign in

REVIEW 5 cited by

EmoBench: Evaluating the Emotional Intelligence of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.12071 v3 pith:YUDNI7E7 submitted 2024-02-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords emobenchemotionalemotionexistingunderstandingbenchmarkscomprehensiveevaluating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in Large Language Models (LLMs) have highlighted the need for robust, comprehensive, and challenging benchmarks. Yet, research on evaluating their Emotional Intelligence (EI) is considerably limited. Existing benchmarks have two major shortcomings: first, they mainly focus on emotion recognition, neglecting essential EI capabilities such as emotion regulation and thought facilitation through emotion understanding; second, they are primarily constructed from existing datasets, which include frequent patterns, explicit information, and annotation errors, leading to unreliable evaluation. We propose EmoBench, a benchmark that draws upon established psychological theories and proposes a comprehensive definition for machine EI, including Emotional Understanding and Emotional Application. EmoBench includes a set of 400 hand-crafted questions in English and Chinese, which are meticulously designed to require thorough reasoning and understanding. Our findings reveal a considerable gap between the EI of existing LLMs and the average human, highlighting a promising direction for future research. Our code and data are publicly available at https://github.com/Sahandfer/EmoBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

    cs.CL 2026-03 unverdicted novelty 6.0 of 10

    MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.

  2. An LLM's Apology: Outsourcing Awkwardness in the Age of AI

    cs.CY 2025-06 reject novelty 6.0 of 10

    Anthropic's Sonnet models score highest on FLAKE-Bench, a new benchmark for LLMs that write believable, kind, and human-sounding excuses for cancelling plans.

  3. MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MotiveBench shows that large language models still lag human consensus on motivational reasoning, with GPT-4o scoring 80.89 percent and chain-of-thought prompting usually reducing accuracy.

  4. LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Language agents' believability and goal achievement decline over multi-episode social interactions, and curated memory summaries only partially close the gap with humans.

  5. EmoAssist: Emotional Assistant for Visual Impairment Community

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A fine-tuned LLaVA model, trained with preference optimization on 800 emotional image-QA examples, beats GPT-4o on a new empathy-focused benchmark for visual impairment assistance.

Pith tools