Pith. sign in

REVIEW 3 cited by

Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14693 v1 pith:IKI3AK5Q submitted 2023-05-24 cs.CL cs.LG

classification cs.CLcs.LG
keywords personalityllmsself-assessmenttestsmodelsanswerdevelopedlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Have Large Language Models (LLMs) developed a personality? The short answer is a resounding "We Don't Know!". In this paper, we show that we do not yet have the right tools to measure personality in language models. Personality is an important characteristic that influences behavior. As LLMs emulate human-like intelligence and performance in various tasks, a natural question to ask is whether these models have developed a personality. Previous works have evaluated machine personality through self-assessment personality tests, which are a set of multiple-choice questions created to evaluate personality in humans. A fundamental assumption here is that human personality tests can accurately measure personality in machines. In this paper, we investigate the emergence of personality in five LLMs of different sizes ranging from 1.5B to 30B. We propose the Option-Order Symmetry property as a necessary condition for the reliability of these self-assessment tests. Under this condition, the answer to self-assessment questions is invariant to the order in which the options are presented. We find that many LLMs personality test responses do not preserve option-order symmetry. We take a deeper look at LLMs test responses where option-order symmetry is preserved to find that in these cases, LLMs do not take into account the situational statement being tested and produce the exact same answer irrespective of the situation being tested. We also identify the existence of inherent biases in these LLMs which is the root cause of the aforementioned phenomenon and makes self-assessment tests unreliable. These observations indicate that self-assessment tests are not the correct tools to measure personality in LLMs. Through this paper, we hope to draw attention to the shortcomings of current literature in measuring personality in LLMs and call for developing tools for machine personality measurement.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    SAE features recovered from matched high–low trait behaviors can be steered to bidirectionally shift situational personality expression and produce human-like social benefit–cost patterns in an 8B LLM.

  2. CAPE: Context-Aware Personality Evaluation Framework for Large Language Models

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Conversational history changes LLM personality-test answers: it increases answer consistency through in-context learning but shifts OCEAN scores, especially for GPT-3.5/4, while smaller models rely heavily on prior in...

  3. AI YOU Town: Make Friends and Money with Your Digital Twin

    cs.AI 2026-07 conditional novelty 5.0 of 10

    A unified LLM pipeline with Bayesian trait updates, conformal sets, and periodic memory-anchor refresh improves calibration and long-horizon persona fidelity over static prompting on module benchmarks.

Pith tools