Agent-ValueBench is the first dedicated benchmark for agent values, showing they diverge from LLM values, form a homogeneous 'Value Tide' across models, and bend under harnesses and skill steering.
Largelanguagemodelpsychomet- rics: A systematic review of evaluation, validation, and enhancement.CoRR, abs/2505.08245
9 Pith papers cite this work, alongside 5 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
An LLM-native five-factor psychometric instrument shows self-reports fail to predict behavior even on constructs derived from LLM behavior, and LLM judges share a variance source humans do not.
Apparent psychological profiles of LLMs are largely measurement artifacts driven by directional response bias rather than actual traits.
The primary axis of psychometric variation among LLMs is the degree to which they represent themselves as loci of phenomenal experience rather than systems of behavioral responses.
Validity indices adapted from clinical assessment classify four frontier LLMs as construct-level invalid on metacognitive probes, with valid models showing positive item-sensitive confidence (r=.18) while invalid ones show the opposite (r=-.20).
LLM self-reports predict behavior selectively: TPB reaches human-level coherence within shared conversations but collapses across sessions for primed behaviors, unlike Big 5, with persona prompting stabilizing reports but not actions.
Fine-tuning LLMs on small pilot survey data balances structural, marginal, and individual fidelity better than prompting or rectification, but fidelity levels vary across subsamples in a COVID-19 misinformation case study.
A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.
Standard psychometric questionnaires like the Big Five and PVQ produce different and more consistent results than ecologically valid questions drawn from real user conversations, suggesting the former may mischaracterize LLM behavior.
citing papers explorer
-
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
Agent-ValueBench is the first dedicated benchmark for agent values, showing they diverge from LLM values, form a homogeneous 'Value Tide' across models, and bend under harnesses and skill steering.
-
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
An LLM-native five-factor psychometric instrument shows self-reports fail to predict behavior even on constructs derived from LLM behavior, and LLM judges share a variance source humans do not.
-
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
Apparent psychological profiles of LLMs are largely measurement artifacts driven by directional response bias rather than actual traits.
-
The Pinocchio Dimension: Phenomenality of Experience as the Primary Axis of LLM Psychometric Differences
The primary axis of psychometric variation among LLMs is the degree to which they represent themselves as loci of phenomenal experience rather than systems of behavioral responses.
-
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report
Validity indices adapted from clinical assessment classify four frontier LLMs as construct-level invalid on metacognitive probes, with valid models showing positive item-sensitive confidence (r=.18) while invalid ones show the opposite (r=-.20).
-
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
LLM self-reports predict behavior selectively: TPB reaches human-level coherence within shared conversations but collapses across sessions for primed behaviors, unlike Big 5, with persona prompting stabilizing reports but not actions.
-
Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data
Fine-tuning LLMs on small pilot survey data balances structural, marginal, and individual fidelity better than prompting or rectification, but fidelity levels vary across subsamples in a COVID-19 misinformation case study.
-
A Survey of Large Language Models for Perception and Measurement of Human Psychology
A survey proposing a three-pillar framework to evaluate LLMs as tools for measuring latent psychological constructs and reviewing applications in personality and mental health.
-
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Standard psychometric questionnaires like the Big Five and PVQ produce different and more consistent results than ecologically valid questions drawn from real user conversations, suggesting the former may mischaracterize LLM behavior.