Pith. sign in

REVIEW 2 cited by

Are Large Language Models More Empathetic than Humans?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.05063 v1 pith:IFIFLNF3 submitted 2024-06-07 cs.CL

classification cs.CL
keywords empatheticllmshumansrespondingapproximatelyassessingcomparedemotions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the emergence of large language models (LLMs), investigating if they can surpass humans in areas such as emotion recognition and empathetic responding has become a focal point of research. This paper presents a comprehensive study exploring the empathetic responding capabilities of four state-of-the-art LLMs: GPT-4, LLaMA-2-70B-Chat, Gemini-1.0-Pro, and Mixtral-8x7B-Instruct in comparison to a human baseline. We engaged 1,000 participants in a between-subjects user study, assessing the empathetic quality of responses generated by humans and the four LLMs to 2,000 emotional dialogue prompts meticulously selected to cover a broad spectrum of 32 distinct positive and negative emotions. Our findings reveal a statistically significant superiority of the empathetic responding capability of LLMs over humans. GPT-4 emerged as the most empathetic, marking approximately 31% increase in responses rated as "Good" compared to the human benchmark. It was followed by LLaMA-2, Mixtral-8x7B, and Gemini-Pro, which showed increases of approximately 24%, 21%, and 10% in "Good" ratings, respectively. We further analyzed the response ratings at a finer granularity and discovered that some LLMs are significantly better at responding to specific emotions compared to others. The suggested evaluation framework offers a scalable and adaptable approach for assessing the empathy of new LLMs, avoiding the need to replicate this study's findings in future research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

    cs.HC 2025-09 reject novelty 6.0 of 10

    E-THER is a small annotated therapy-video dataset for verbal-visual incongruence, but the claimed empathy gains are supported mainly by author-built keyword metrics with statistical inconsistencies.

  2. Distilling Empathy from Large Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Fine-tuning small models with SFT plus DPO on empathy-improved responses from LLMs yields large judged win-rate gains over base small models.

Pith tools