REVIEW 22 cited by
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com
Forward citations
Cited by 22 Pith papers
-
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
AlpsBench supplies 2500 real-dialogue sequences with verified memories to benchmark LLM extraction, updating, retrieval, and utilization of personalized information.
-
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
AttuneBench introduces a multi-turn conversation benchmark with participant annotations showing that LLM emotional intelligence decomposes into independent capabilities with preference prediction being most discriminative.
-
StoryAlign: Evaluating and Training Reward Models for Story Generation
StoryReward, trained on a new 100k story preference dataset, sets state-of-the-art performance on the introduced StoryRMB benchmark for aligning LLM stories with human preferences.
-
Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
A forecasting-based metric called 100-Endings quantifies story tension via prediction mismatches, correctly ranks human stories above AI ones, and supports a structural pipeline for higher-tension LLM story generation.
-
StoryScope: Investigating idiosyncrasies in AI fiction
Discourse-level narrative features, extracted by LLMs across 10 dimensions, distinguish AI-generated from human-written stories at 93.2% macro-F1 and attribute authorship at 68.4% macro-F1.
-
StoryScope: Investigating idiosyncrasies in AI fiction
StoryScope extracts narrative features showing AI stories favor tidy plots and over-explain themes while human stories show more moral ambiguity and temporal complexity, enabling strong detection and attribution.
-
HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
HSS-Synth generates 230k instruction-tuning samples for 14 humanities/social-science fields and reports state-of-the-art fine-tuning results on 16 benchmarks.
-
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
BridgeAlign's rubric-guided 'bridge' degradation creates near-boundary hard preference pairs, letting Qwen3-8B beat 11 baselines on average across 17 human-preference and knowledge benchmarks.
-
PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
Personality-aware prompting plus contrastive retrieval of aligned scenarios measurably lifts LLM accuracy on EmoBench emotional-understanding tasks, especially for GPT models with CoT.
-
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
A new evaluation framework using MMD on Biber features shows LLMs deviate from human linguistic distributions across registers, with closest models varying by register rather than size.
-
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
The authors introduce a register-aware evaluation framework that compares LLM outputs to human reference corpora via Biber's lexico-grammatical features and MMD across five English registers.
-
AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
AttuneBench introduces a multi-turn conversation benchmark using participant annotations to evaluate LLM emotional intelligence, finding that model performance on emotion recognition, behavior classification, preferen...
-
Information-Theoretic Limits of Reliability and Scaling in Language Models
A theoretical framework derives a reliability ceiling and a max-form Chinchilla-type scaling law for LLMs from task entropy and dependency spectra.
-
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
MICA combines incremental per-turn distance rewards and Monte Carlo returns from a shared potential function over user support states to create a mixed advantage signal that enables stable multi-turn RL optimization f...
-
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation
A skeleton-first reasoning generation method reduces answer anchoring in reverse chain-of-thought traces, while semantic suppression increases latent anchoring.
-
Escaping the Verifier: Learning to Reason via Demonstrations
RARO trains reasoning LLMs from expert demonstrations alone using an adversarial game between a policy and a relativistic critic, outperforming verifier-free baselines and approaching verifier-based RL.
-
Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
Parallel inference rollouts aggregated into pseudo-references enable reference-free RL supervision that matches expert-annotated performance on health tasks while using 9x less test-time compute.
-
Emotional intelligence in large language models is fragmented across perception, cognition, and interaction
FACET evaluation of nine frontier LLMs shows emotional intelligence is fragmented into cognitive-dominant, interactive-dominant, and context-dependent profiles, with hidden emotion recognition as a universal bottleneck.
-
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
MICA mixes per-turn and whole-trajectory normalized reward signals to train emotional-support chatbots, outperforming GRPO and REINFORCE++ on EMPA, EQ-Bench, and EmoBench.
-
RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
RECAP is an inference-time framework using cognitive appraisal theory to enhance emotional alignment and transparency in medical dialogue systems across model scales.
-
PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
A post-training framework with persona-specific LoRA experts and a situation-aware router improves LLM emotional responses, but the evidence on preserving general ability is undercut by missing base-model comparisons.
-
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
Focal Reward balances rubric-based RL by saturation-aware reweighting derived from inverse reward projection, outperforming static aggregation on 18 model-benchmark pairs.
Discussion (0). Sign in to comment.