Pith. sign in

REVIEW 5 cited by

Large Language Models Do Not Simulate Human Psychology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.06950 v3 pith:FAFXNX73 submitted 2025-08-09 cs.AI

Large Language Models Do Not Simulate Human Psychology

classification cs.AI
keywords humanllmspsychologyresponsessimulatelargepsychologicalarguments
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large Language Models (LLMs),such as ChatGPT, are increasingly used in research, ranging from simple writing assistance to complex data annotation tasks. Recently, some research has suggested that LLMs may even be able to simulate human psychology and can, hence, replace human participants in psychological studies. We caution against this approach. We provide conceptual arguments against the hypothesis that LLMs simulate human psychology. We then present empiric evidence illustrating our arguments by demonstrating that slight changes to wording that correspond to large changes in meaning lead to notable discrepancies between LLMs' and human responses, even for the recent CENTAUR model that was specifically fine-tuned on psychological responses. Additionally, different LLMs show very different responses to novel items, further illustrating their lack of reliability. We conclude that LLMs do not simulate human psychology and recommend that psychological researchers should treat LLMs as useful but fundamentally unreliable tools that need to be validated against human responses for every new application.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Building an Atlas of Social Experiments to Link Studies, Reconcile Conflicts, and Bridge Gaps

    cs.CY 2026-05 unverdicted novelty 7.0

    ExAtlas composes effects from locally close prior studies in treatment-outcome space to link, reconcile conflicts, or propose bridge experiments, recovering effect direction in 98.6% of held-out locally supported targ...

  2. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 7.0

    The paper delivers a taxonomy of seven LLM study types in software engineering along with eight guidelines that separate mandatory requirements from recommended practices to address reproducibility challenges.

  3. Guidelines for Empirical Studies in Software Engineering involving Large Language Models

    cs.SE 2025-08 accept novelty 6.0

    A group of 22 researchers proposes seven study types and eight guidelines for empirical software engineering studies involving LLMs to enhance reproducibility and replicability.

  4. Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating

    cs.CY 2026-07 conditional novelty 5.0

    Semantic Similarity Rating of LLM text responses substantially reduces mode collapse in silicon sampling of political thermometer scores versus direct numeric prompting, with one global temperature that generalizes fr...

  5. Addressing Longstanding Challenges in Cognitive Science with Language Models

    cs.AI 2025-10 conditional novelty 4.0

    A review proposes that LLMs can serve as tools for a more integrative and cumulative cognitive science when used under human oversight.