Pith. sign in

REVIEW 3 cited by

Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.07248 v1 pith:7PD67JVO submitted 2024-05-12 cs.CL cs.AIcs.CYcs.HC

classification cs.CLcs.AIcs.CYcs.HC
keywords llmshumanresponsessimulatedescriptionspsychologicalabilitygeneric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The humanlike responses of large language models (LLMs) have prompted social scientists to investigate whether LLMs can be used to simulate human participants in experiments, opinion polls and surveys. Of central interest in this line of research has been mapping out the psychological profiles of LLMs by prompting them to respond to standardized questionnaires. The conflicting findings of this research are unsurprising given that mapping out underlying, or latent, traits from LLMs' text responses to questionnaires is no easy task. To address this, we use psychometrics, the science of psychological measurement. In this study, we prompt OpenAI's flagship models, GPT-3.5 and GPT-4, to assume different personas and respond to a range of standardized measures of personality constructs. We used two kinds of persona descriptions: either generic (four or five random person descriptions) or specific (mostly demographics of actual humans from a large-scale human dataset). We found that the responses from GPT-4, but not GPT-3.5, using generic persona descriptions show promising, albeit not perfect, psychometric properties, similar to human norms, but the data from both LLMs when using specific demographic profiles, show poor psychometrics properties. We conclude that, currently, when LLMs are asked to simulate silicon personas, their responses are poor signals of potentially underlying latent traits. Thus, our work casts doubt on LLMs' ability to simulate individual-level human behaviour across multiple-choice question answering tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 6 citations worldwide. Full citation record

  1. Will Scaling Improve Social Simulation with LLMs?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.

  2. Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A mediator-guided LLM simulation selects survey items whose virtual convergent validity matches human ground truth better than random or LLM-judge baselines.

  3. Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    SILIC uses LLM-guided inverse reinforcement learning and Theory of Planned Behavior chain reasoning to infer age, gender, income, and employment from travel trajectories, reportedly beating SVM, XGBoost, CatBoost, and...

Pith tools