Pith. sign in

REVIEW 4 cited by

Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19599 v3 pith:5LYEDL4F submitted 2024-10-25 econ.GN cs.AIcs.CYcs.HCq-fin.EC

classification econ.GNcs.AIcs.CYcs.HCq-fin.EC
keywords llmshumanbehaviorsurrogatescautionhumanslanguagemany
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Will Scaling Improve Social Simulation with LLMs?

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.

  2. LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    LISTEN uses LLMs as zero-shot preference oracles, via iterative utility refinement (LISTEN-U) or tournament comparisons (LISTEN-T), to select preferred items from large multi-objective candidate sets.

  3. Static network structure cannot stabilize cooperation among Large Language Model agents

    cs.SI 2024-11 conditional novelty 6.0 of 10

    LLM agents playing repeated prisoner's dilemma did not show the network-stabilized cooperation seen in humans, and GPT-3.5 barely responded to network structure.

  4. LLM-Mirror: A Generated-Persona Approach for Survey Pre-Testing

    cs.CY 2024-12 conditional novelty 5.0 of 10

    LLM-generated survey responses, conditioned on a respondent's prior answers or a generated persona, align with human responses at the distributional level and, less reliably, at the individual level.

Pith tools