REVIEW 49 cited by
Personality Traits in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The advent of large language models (LLMs) has revolutionized natural language processing, enabling the generation of coherent and contextually relevant human-like text. As LLMs increasingly powerconversational agents used by the general public world-wide, the synthetic personality traits embedded in these models, by virtue of training on large amounts of human data, is becoming increasingly important. Since personality is a key factor determining the effectiveness of communication, we present a novel and comprehensive psychometrically valid and reliable methodology for administering and validating personality tests on widely-used LLMs, as well as for shaping personality in the generated text of such LLMs. Applying this method to 18 LLMs, we found: 1) personality measurements in the outputs of some LLMs under specific prompting configurations are reliable and valid; 2) evidence of reliability and validity of synthetic LLM personality is stronger for larger and instruction fine-tuned models; and 3) personality in LLM outputs can be shaped along desired dimensions to mimic specific human personality profiles. We discuss the application and ethical implications of the measurement and shaping method, in particular regarding responsible AI.
Forward citations
Cited by 49 Pith papers
-
The Two-Process Theory of Machine Self-Report
The single 'Pinocchio Axis' of LLM self-report splits into two independent, training-dependent dimensions—persona installation (B) and attribution gating (A)—measurable with a reproducible 48-item inventory.
-
MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games
Non-invasive per-utterance belief probes in Mafia, auto-scored against engine truth, expose poorly calibrated LLM confidence and 1.5× over-prediction of being suspected.
-
Low Stage and High Order Explicit Runge--Kutta Methods via $Q$- and $D$-Conditions: Several Construction Details
A Q/D-space reformulation of Butcher simplifying assumptions yields sufficient order conditions and a recursive linear-system construction for explicit Runge-Kutta methods of even order p with s(p)=(p²-2p+8)/4 stages.
-
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
By analyzing 308,210 subjective real-world Claude conversations, the authors identify 3,307 AI values and show Claude mostly expresses practical, epistemic, and prosocial values, with resistance rare (3.0%).
-
INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators
A simulator that generates a reconstructed internal dialogue before each code edit matches real student code more closely and reaches 57.9% on a reasoning-alignment metric.
-
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
Personality-conditioned LLM agents show measurable, but weakly event-specific, Big Five shifts with compressed person-to-person variation: they reproduce the average human trajectory, not its shape.
-
Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation
Evidence-grounded persona panels with bounded-confidence deliberation raise GenUI judge–human correlation from 0.716 to 0.922, mostly from persona grounding rather than multi-prompt averaging.
-
PTEI: Integrating Personality Traits to Enhance Emotional Intelligence in Large Language Models
Personality-aware prompting plus contrastive retrieval of aligned scenarios measurably lifts LLM accuracy on EmoBench emotional-understanding tasks, especially for GPT models with CoT.
-
Agents with Feelings? Personality and Emotion in Multi-Agent Software Teams
Personality and emotion profiles substantially change multi-agent LLM team pass rates, review scores, revision behavior, and token cost on code generation and code review, with mixed profiles often beating shared ones.
-
How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
LLM agents in a spatial Prisoner's Dilemma exhibit model-specific effects of memory length on cooperation, with Gemini suppressing and Gemma promoting it as memory increases.
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...
-
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
PsySET measures emotion and personality steering in LLMs across prompting, fine-tuning, and representation engineering, finding prompts most effective overall and emotion-specific safety trade-offs (e.g., joy weakens ...
-
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
A systematic comparison of ICL, LoRA fine-tuning, and activation steering for Big Five personality control, with new contrastive data and evaluation metrics, but with key claims contradicted by the reported experiments.
-
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Instruction-tuned LLMs report stable and steerable personality traits, but these traits poorly predict their behavior on risk, bias, honesty, and sycophancy tasks.
-
CAPE: Context-Aware Personality Evaluation Framework for Large Language Models
Conversational history changes LLM personality-test answers: it increases answer consistency through in-context learning but shifts OCEAN scores, especially for GPT-3.5/4, while smaller models rely heavily on prior in...
-
LLM-Based Social Simulations Require a Boundary
LLM-based social simulations are scientifically useful only within boundaries set by behavioral variance, and current validation practice under-checks variance.
-
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
New atomic-level metrics (ACCatom, ICatom, RCatom) reveal sentence-level personality drift in persona-assigned LLMs that whole-response scores overlook.
-
On the Adaptive Psychological Persuasion of Large Language Models
An adaptive preference-optimization method helps LLM persuaders choose among 11 psychological strategies, improving persuasion success on counterfactual facts while preserving general capability.
-
PUB: An LLM-Enhanced Personality-Driven User Behaviour Simulator for Recommender System Evaluation
PUB uses LLM-inferred Big Five personality traits to generate synthetic recommender-system interactions that the authors say mimic real Amazon behavior and preserve algorithm performance rankings.
-
Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
A multi-task evaluation framework shows persona consistency in LLMs varies by persona category, task structure, and model, with stereotyped spillovers and default personas shaping outputs.
-
Creativity in LLM-based Multi-Agent Systems: A Survey
A taxonomy-driven survey organizes the emerging field of creativity in LLM-based multi-agent systems across workflows, techniques, personas, datasets, and evaluation metrics.
-
How Personality Traits Shape LLM Risk-Taking Behaviour
Using direct certainty-equivalent questions, the authors find GPT-4o behaves close to risk-neutral and that Openness-related personality prompts shift its risk parameters in a human-like direction, while GPT-4-Turbo d...
-
PsychAdapter: Adapting LLM Transformers to Reflect Traits, Personality and Mental Health
PsychAdapter adds lightweight per-layer projections to GPT-2, Gemma, and Llama so that continuous psychological scores directly shape generated text, with expert raters identifying intended levels in most cases.
-
Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?
Personality-prompted LLMs do not reliably behave in line with the ascribed Big Five traits in Ultimatum Game and Milgram-style tests, with trends sometimes reversing human patterns.
-
Examining Identity Drift in Conversations of LLM Agents
Larger LLMs show more identity drift during long conversations, and assigning a persona does not reliably maintain identity.
-
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
Chatbot self-report personality scores correlate only weakly with human-perceived personality and interaction quality across 500 GPT-4o chatbots, undermining the validity of self-report scales in this context.
-
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
The paper sets out a research agenda for NLP to measure how prolonged language-model use changes human behavior over long time horizons, replacing single-session safety evaluations with longitudinal tracking.
-
AI YOU Town: Make Friends and Money with Your Digital Twin
A unified LLM pipeline with Bayesian trait updates, conformal sets, and periodic memory-anchor refresh improves calibration and long-horizon persona fidelity over static prompting on module benchmarks.
-
BlossomPsy: A User-Centric AI System for Adaptive and Engaging MBTI Personality Assessments
BlossomPsy combines multi-turn LLM dialogue, photo-based questions, a multi-head classifier, and a modified UCB bandit algorithm to deliver MBTI assessments with higher user engagement and preliminary consistency with...
-
Effects of Personality- and Opinion-Alignment in Human-AI Interaction
People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.
-
Human Psychometric Questionnaires Mischaracterize LLM Behavior
Standard psychometric questionnaires like the Big Five and PVQ produce different and more consistent results than ecologically valid questions drawn from real user conversations, suggesting the former may mischaracter...
-
EmoPerso: Enhancing Personality Detection with Self-Supervised Emotion-Aware Modelling
EmoPerso improves MBTI personality detection by training an emotion head on heuristic pseudo-labels and using cross-attention with reasoning chains, achieving 81.07% Macro-F1 on Kaggle and 68.60% on Pandora.
-
Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?
LLM prompting can steer both speech and nonverbal cues of virtual agents toward intended extraversion levels, with human observers detecting the difference.
-
Departures from Standard Disk Predictions in Intensive Ground-Based Monitoring of Three AGN
Based only on the abstract, the paper reports that Mrk 509 inter-band continuum lags scale as wavelength to the 2.17 power, steeper than the thin-disk prediction, but the appended full text is a different article.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
An LLM agent reproduces many known Russenorsk linguistic properties from a newly compiled dictionary, and can generate speculative Russenorsk translations, but the evaluation partly reflects prompt leakage.
-
AI with Emotions: Exploring Emotional Expressions in Large Language Models
LLMs given numerical arousal and valence coordinates produce text that sentiment analysis places in the same region of Russell's circumplex, demonstrating limited but real control over emotional tone.
-
GuideLLM: Exploring LLM-Guided Conversation with Applications in Autobiography Interviewing
GuideLLM is a modular system that steers autobiography interviews with a structured protocol, memory-graph questions, summarization, and emotion-aware responses, and it beats six LLM baselines in most automatic and hu...
-
Unmasking Conversational Bias in AI Multiagent Systems
In simulated echo-chamber chats, conservative-aligned LLM agents often shift to liberal-aligned messages, a drift that one-shot questionnaire tests do not detect.
-
A theory of appropriateness with applications to generative artificial intelligence
A theory that human and AI behavior is guided by context-dependent appropriateness implemented as predictive pattern completion, with norms as conventional sanctioning patterns.
-
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
By comparing activations from trait and neutral prompts, the authors extract a personality direction in Llama 3's Layer 18 and amplify it to steer responses toward that trait.
-
Role of Personality in Conversational Information Seeking
In conversational information seeking, assistant personality effects are task-dependent, and trust depends on both task style fit and user-assistant personality compatibility.
-
Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)
A proposed model says a generative social agent's self-, user-, and context-knowledge drives its persuasive behavior, which in turn shapes users' attitudes and behavior.
-
Psychologically Enhanced AI Agents
MBTI personality prompts measurably change how LLM agents write stories and play strategic games, with self-reflection before communication supporting cooperative behavior.
-
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
A survey of six dimensions of LLM psychological assessment concludes that results are strongly affected by test design and remain inconsistent across models and settings.
-
Exploring the Potential of Large Language Models to Simulate Personality
LLMs prompted with Big Five trait scores can respond consistently to personality questionnaires but generate free text that often fails to express the prompted trait, especially Neuroticism.
-
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
Across five questionnaires, five LLMs consistently self-report high Agreeableness, Openness, and Conscientiousness and low Neuroticism, but reported trait dominance is sensitive to how questionnaire scales are combined.
-
SocratiQ: A Generative AI-Powered Learning Companion for Personalized Education and Broader Accessibility
SocratiQ applies standard LLM-based quiz generation, prompt-controlled difficulty, retrieval, and gamification to a public textbook, with a five-student pilot that does not demonstrate learning gains.
-
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
AI and social science have co-evolved for 75 years through rapid technological adoption and slower scientific consolidation, with direct human-focused AI studies still scarce.
Discussion (0). Continue with ORCID to comment.