Pith. sign in

REVIEW 12 cited by

PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.02547 v5 pith:VQBHAI7L submitted 2023-05-04 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords personalitytraitspersonasfivehumanlargellmsaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the many use cases for large language models (LLMs) in creating personalized chatbots, there has been limited research on evaluating the extent to which the behaviors of personalized LLMs accurately and consistently reflect specific personality traits. We consider studying the behavior of LLM-based agents which we refer to as LLM personas and present a case study with GPT-3.5 and GPT-4 to investigate whether LLMs can generate content that aligns with their assigned personality profiles. To this end, we simulate distinct LLM personas based on the Big Five personality model, have them complete the 44-item Big Five Inventory (BFI) personality test and a story writing task, and then assess their essays with automatic and human evaluations. Results show that LLM personas' self-reported BFI scores are consistent with their designated personality types, with large effect sizes observed across five traits. Additionally, LLM personas' writings have emerging representative linguistic patterns for personality traits when compared with a human writing corpus. Furthermore, human evaluation shows that humans can perceive some personality traits with an accuracy of up to 80%. Interestingly, the accuracy drops significantly when the annotators were informed of AI authorship.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 27 citations worldwide. Full citation record

  1. Low Stage and High Order Explicit Runge--Kutta Methods via $Q$- and $D$-Conditions: Several Construction Details

    math.NA 2026-05 unverdicted novelty 7.0 of 10

    A Q/D-space reformulation of Butcher simplifying assumptions yields sufficient order conditions and a recursive linear-system construction for explicit Runge-Kutta methods of even order p with s(p)=(p²-2p+8)/4 stages.

  2. More Is Not More: What Matters for Diversity in LLM Opinions?

    cs.CL 2026-05 conditional novelty 7.0 of 10

    Diversity in LLM opinions comes mostly from the first persona sentence and from combining different interaction architectures, not from richer personas, temperature, or diversity instructions.

  3. Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Small hyperbolic models (146M–3B) report 100% creative-seed preference, 90.7% compliance-gap detection, and a selective-gating skeleton–wallpaper memory pilot as a companion-AI stack.

  4. How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    LLM agents in a spatial Prisoner's Dilemma exhibit model-specific effects of memory length on cooperation, with Gemini suppressing and Gemma promoting it as memory increases.

  5. McKean-Vlasov SPDEs driven by Poisson random measure: Well-posedness and large deviation principle

    math.PR 2025-08 unverdicted novelty 6.0 of 10

    The abstract announces well-posedness and a large deviation principle for McKean-Vlasov SPDEs with Poisson jumps, but the supplied full text is a different paper.

  6. Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs

    cs.AI 2025-05 conditional novelty 6.0 of 10

    LLM confidence is less sensitive to task difficulty than human confidence and bends to persona stereotypes, and separating confidence prompts from answer prompts (AFCE) improves calibration on hard tasks.

  7. Localizing Persona Representations in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Persona information is most separable in the final third of LLM layers, and in Llama3's last layer ethical personas share 17.6% of salient activations while political personas have 2.1% to 5.5% unique activations.

  8. RESBev: Making BEV Perception More Robust

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    A latent world model predicts clean BEV semantic features from sequential observations to recover existing Lift-Splat-Shoot pipelines under natural and adversarial corruption with few-shot fine-tuning.

  9. Effects of Personality- and Opinion-Alignment in Human-AI Interaction

    cs.HC 2025-11 conditional novelty 5.0 of 10

    People rate AI chatbots as more trustworthy, competent, warm, and persuasive when the chatbots share their opinion, whereas matching the chatbot's personality to the user's has little or no effect.

  10. Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors

    cs.IR 2025-08 conditional novelty 5.0 of 10

    Persona-level LoRA fine-tuning lets a 3.8B small language model simulate MovieLens users about as accurately as a much larger frozen LLM, at lower cost.

  11. Exploring the Impact of Occupational Personas on Domain-Specific QA

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Profession-based personas slightly improve LLM accuracy on science QA, while occupational personality personas often reduce it, even when semantically related.

  12. Promoting Online Safety by Simulating Unsafe Conversations with LLMs

    cs.HC 2025-07 conditional novelty 4.0 of 10

    A pair of language models, one acting as scammer and one as target, can simulate realistic scam conversations, but the paper presents no user evaluation of whether this improves scam resilience.

Pith tools