Pith. sign in

REVIEW 12 cited by

Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14703 v3 pith:M4HKYC54 submitted 2024-06-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmspersonalitytraitvalidityagentsconsistentdatadesigned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recent advancements in Large Language Models (LLMs) have led to their adaptation in various domains as conversational agents. We wonder: can personality tests be applied to these agents to analyze their behavior, similar to humans? We introduce TRAIT, a new benchmark consisting of 8K multi-choice questions designed to assess the personality of LLMs. TRAIT is built on two psychometrically validated small human questionnaires, Big Five Inventory (BFI) and Short Dark Triad (SD-3), enhanced with the ATOMIC-10X knowledge graph to a variety of real-world scenarios. TRAIT also outperforms existing personality tests for LLMs in terms of reliability and validity, achieving the highest scores across four key metrics: Content Validity, Internal Validity, Refusal Rate, and Reliability. Using TRAIT, we reveal two notable insights into personalities of LLMs: 1) LLMs exhibit distinct and consistent personality, which is highly influenced by their training data (e.g., data used for alignment tuning), and 2) current prompting techniques have limited effectiveness in eliciting certain traits, such as high psychopathy or low conscientiousness, suggesting the need for further research in this direction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Persona Cartography: Charting Language Model Personality Traits in Weight Space

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.

  2. The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Instruction-tuned LLMs report stable and steerable personality traits, but these traits poorly predict their behavior on risk, bias, honesty, and sycophancy tasks.

  3. On the Adaptive Psychological Persuasion of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    An adaptive preference-optimization method helps LLM persuaders choose among 11 psychological strategies, improving persuasion success on counterfactual facts while preserving general capability.

  4. Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Value Portrait links LLM value scores to human PVQ-validated items from real conversations, finding that LLMs emphasize Benevolence, Security, and Self-Direction while downplaying Tradition, Power, and Achievement.

  5. Examining Identity Drift in Conversations of LLM Agents

    cs.CY 2024-12 conditional novelty 6.0 of 10

    Larger LLMs show more identity drift during long conversations, and assigning a persona does not reliably maintain identity.

  6. Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

    cs.HC 2024-11 conditional novelty 6.0 of 10

    Chatbot self-report personality scores correlate only weakly with human-perceived personality and interaction quality across 500 GPT-4o chatbots, undermining the validity of self-report scales in this context.

  7. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  8. Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?

    cs.HC 2025-08 conditional novelty 5.0 of 10

    LLM prompting can steer both speech and nonverbal cues of virtual agents toward intended extraversion levels, with human observers detecting the difference.

  9. Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)

    cs.HC 2026-02 conditional novelty 4.0 of 10

    A proposed model says a generative social agent's self-, user-, and context-knowledge drives its persuasive behavior, which in turn shapes users' attitudes and behavior.

  10. Scaling Personality Control in LLMs with Big Five Scaler Prompts

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Numeric Big Five trait values placed in prompts shift LLMs' self-reported and dialogue-expressed personality, with simple prompts and low intensity scales working best.

  11. A validity-guided workflow for robust large language model research in psychology

    cs.HC 2025-07 conditional novelty 4.0 of 10

    A six-stage workflow scales validity requirements to research ambition so that LLM-based psychological claims rest on demonstrated measurement quality rather than prompt artifacts.

  12. Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications

    cs.CY 2025-04 conditional novelty 4.0 of 10

    A survey of six dimensions of LLM psychological assessment concludes that results are strongly affected by test design and remain inconsistent across models and settings.

Pith tools