REVIEW 12 cited by
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent advancements in Large Language Models (LLMs) have led to their adaptation in various domains as conversational agents. We wonder: can personality tests be applied to these agents to analyze their behavior, similar to humans? We introduce TRAIT, a new benchmark consisting of 8K multi-choice questions designed to assess the personality of LLMs. TRAIT is built on two psychometrically validated small human questionnaires, Big Five Inventory (BFI) and Short Dark Triad (SD-3), enhanced with the ATOMIC-10X knowledge graph to a variety of real-world scenarios. TRAIT also outperforms existing personality tests for LLMs in terms of reliability and validity, achieving the highest scores across four key metrics: Content Validity, Internal Validity, Refusal Rate, and Reliability. Using TRAIT, we reveal two notable insights into personalities of LLMs: 1) LLMs exhibit distinct and consistent personality, which is highly influenced by their training data (e.g., data used for alignment tuning), and 2) current prompting techniques have limited effectiveness in eliciting certain traits, such as high psychopathy or low conscientiousness, suggesting the need for further research in this direction.
Forward citations
Cited by 12 Pith papers
-
Persona Cartography: Charting Language Model Personality Traits in Weight Space
Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.
-
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Instruction-tuned LLMs report stable and steerable personality traits, but these traits poorly predict their behavior on risk, bias, honesty, and sycophancy tasks.
-
On the Adaptive Psychological Persuasion of Large Language Models
An adaptive preference-optimization method helps LLM persuaders choose among 11 psychological strategies, improving persuasion success on counterfactual facts while preserving general capability.
-
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
Value Portrait links LLM value scores to human PVQ-validated items from real conversations, finding that LLMs emphasize Benevolence, Security, and Self-Direction while downplaying Tradition, Power, and Achievement.
-
Examining Identity Drift in Conversations of LLM Agents
Larger LLMs show more identity drift during long conversations, and assigning a persona does not reliably maintain identity.
-
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
Chatbot self-report personality scores correlate only weakly with human-perceived personality and interaction quality across 500 GPT-4o chatbots, undermining the validity of self-report scales in this context.
-
Effects of Personality- and Opinion-Alignment in Human-AI Interaction
People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.
-
Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?
LLM prompting can steer both speech and nonverbal cues of virtual agents toward intended extraversion levels, with human observers detecting the difference.
-
Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)
A proposed model says a generative social agent's self-, user-, and context-knowledge drives its persuasive behavior, which in turn shapes users' attitudes and behavior.
-
Scaling Personality Control in LLMs with Big Five Scaler Prompts
Numeric Big Five trait values placed in prompts shift LLMs' self-reported and dialogue-expressed personality, with simple prompts and low intensity scales working best.
-
A validity-guided workflow for robust large language model research in psychology
A six-stage workflow scales validity requirements to research ambition so that LLM-based psychological claims rest on demonstrated measurement quality rather than prompt artifacts.
-
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
A survey of six dimensions of LLM psychological assessment concludes that results are strongly affected by test design and remain inconsistent across models and settings.
Discussion (0). Continue with ORCID to comment.