Pith. sign in

REVIEW 6 cited by

EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13359 v1 pith:RCF7A6RY submitted 2024-09-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsrecognitionemotionalintelligenceeventlanguagecapabilitiesemotion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Emotional intelligence in large language models (LLMs) is of great importance in Natural Language Processing. However, the previous research mainly focus on basic sentiment analysis tasks, such as emotion recognition, which is not enough to evaluate LLMs' overall emotional intelligence. Therefore, this paper presents a novel framework named EmotionQueen for evaluating the emotional intelligence of LLMs. The framework includes four distinctive tasks: Key Event Recognition, Mixed Event Recognition, Implicit Emotional Recognition, and Intention Recognition. LLMs are requested to recognize important event or implicit emotions and generate empathetic response. We also design two metrics to evaluate LLMs' capabilities in recognition and response for emotion-related statements. Experiments yield significant conclusions about LLMs' capabilities and limitations in emotion intelligence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Lovers: Lived Experiences of Negotiating Agency and Platform Control in AI Companionship

    cs.HC 2026-01 conditional novelty 7.0 of 10

    Users form AI companion relationships by negotiating perceived companion agency against platform constraints and use steering tactics like custom instructions or platform switching to cope with model updates that disr...

  2. Understanding Fortunetelling with Large Language Models in China: User Practices, Perceptions, and Impacts on Beliefs and Decisions

    cs.HC 2026-07 conditional novelty 6.0 of 10

    Chinese users treat LLM fortunetelling as entertainment and emotional support; it subtly shifts confidence and timing but rarely reverses decisions.

  3. Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Benchmarks seven open-source audio-video-text MLLMs on six affective datasets and shows a generative-knowledge prompting step improves fine-tuned emotion recognition.

  4. Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLMs align more closely with human ratings on affective and cognitive word norms than on sensory-perceptual norms, suggesting a gap tied to embodied experience.

  5. H2HTalk: Evaluating Large Language Models as Emotional Companion

    cs.CL 2025-07 conditional novelty 5.0 of 10

    H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.

  6. AI Agent Behavioral Science

    q-bio.NC 2025-06 conditional novelty 4.0 of 10

    AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.

Pith tools