REVIEW 4 cited by
Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.
Forward citations
Cited by 4 Pith papers
-
Will Scaling Improve Social Simulation with LLMs?
Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.
-
LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
LISTEN uses LLMs as zero-shot preference oracles, via iterative utility refinement (LISTEN-U) or tournament comparisons (LISTEN-T), to select preferred items from large multi-objective candidate sets.
-
Static network structure cannot stabilize cooperation among Large Language Model agents
LLM agents playing repeated prisoner's dilemma did not show the network-stabilized cooperation seen in humans, and GPT-3.5 barely responded to network structure.
-
LLM-Mirror: A Generated-Persona Approach for Survey Pre-Testing
LLM-generated survey responses, conditioned on a respondent's prior answers or a generated persona, align with human responses at the distributional level and, less reliably, at the individual level.
Discussion (0). Continue with ORCID to comment.