ActTraitBench is a human-grounded benchmark using psychometric-to-behavior mappings and quantile calibration that reveals pervasive knowledge-decision gaps in 14 LLMs, larger in capable models, with CoCA proposed as mitigation.
The personality illusion: Revealing dissociation between self-reports & behavior in llms.arXiv preprint arXiv:2509.03730
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5verdicts
UNVERDICTED 5representative citing papers
LLMs conditioned on actual psychometric profiles produce life stories from which independent LLMs recover personality scores at mean r=0.75, 85% of human reliability, with emotional patterns replicating in real human data.
LLMs correct only 34.8% of zero-shot annotation errors via prompting, and Definition-Specific Familiarity correlates positively with performance (partial r = +0.41) while memorization metrics do not.
Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.
Persona-E² is a human-annotated dataset linking MBTI and Big Five personality traits to reader emotional responses across text domains, showing that personality data helps LLMs avoid surface-level stereotypes in emotion prediction.
citing papers explorer
-
ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation
ActTraitBench is a human-grounded benchmark using psychometric-to-behavior mappings and quantile calibration that reveals pervasive knowledge-decision gaps in 14 LLMs, larger in capable models, with CoCA proposed as mitigation.
-
Stories of Your Life as Others: A Round-Trip Evaluation of LLM-Generated Life Stories Conditioned on Rich Psychometric Profiles
LLMs conditioned on actual psychometric profiles produce life stories from which independent LLMs recover personality scores at mean r=0.75, 85% of human reliability, with emotional patterns replicating in real human data.
-
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
LLMs correct only 34.8% of zero-shot annotation errors via prompting, and Definition-Specific Familiarity correlates positively with performance (partial r = +0.41) while memorization metrics do not.
-
Interactive Evaluation Requires a Design Science
Interactive evaluation of AI must be reframed as a distinct paradigm that maps interaction trajectories to judgments on process, recoverability, coordination, robustness, and system performance, supported by a two-axis taxonomy and design principles.
-
Persona-E$^2$: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events
Persona-E² is a human-annotated dataset linking MBTI and Big Five personality traits to reader emotional responses across text domains, showing that personality data helps LLMs avoid surface-level stereotypes in emotion prediction.