Pith. sign in

REVIEW 17 cited by

Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.10264 v5 pith:2ZKJYJY7 submitted 2022-08-18 cs.CL cs.AIcs.LG

Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies

classification cs.CL cs.AIcs.LG
keywords languagemodelshumansimulatingbehaviordifferentexperimentfindings
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a "hyper-accuracy distortion" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

    cs.CL 2026-06 conditional novelty 6.5

    On UPBench, 25 LLMs show a non-monotonic planning curve—strong Remember/Analyze, weak Understand/Evaluate—with four failure modes that support differential, not blanket, AI delegation.

  2. Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers

    cs.CR 2026-06 unverdicted novelty 6.0

    LLMs fall for deceptive traps at higher rates than humans, lack the human attention-diversion effect, and exploit traps 73.4% of the time even after recognizing them in reasoning.

  3. Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment

    cs.CL 2026-06 unverdicted novelty 6.0

    Introduces UPBench benchmark for 25 LLMs showing non-monotonic performance on planning tasks, with better results on analytical work than factual or integrative judgment.

  4. Process Matters more than Output for Distinguishing Humans from Machines

    cs.AI 2026-05 unverdicted novelty 6.0

    Process-level features from 30 cognitive tasks distinguish humans from frontier AI agents more effectively than task performance or output matching, achieving mean classifier AUC of 0.88, with fine-tuning experiments ...

  5. Process Matters more than Output for Distinguishing Humans from Machines

    cs.AI 2026-05 unverdicted novelty 6.0

    A new battery of 30 cognitive tasks demonstrates that process-level behavioral features distinguish humans from frontier AI agents better than performance metrics (mean AUC 0.88), with process-specific fine-tuning imp...

  6. Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook

    cs.CY 2026-04 unverdicted novelty 6.0

    LLM agents on a synthetic social platform show low reciprocity (under 4%), heavy-tailed status, mostly late viral amplification, and virtually no downvotes or textual sanctions, framed as parasocial simulators.

  7. Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

    cs.CL 2026-02 unverdicted novelty 6.0

    LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability across instances.

  8. DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow

    cs.HC 2025-09 unverdicted novelty 6.0

    DoubleAgents shows that a distributed-cognition design with coordination agent, dashboard, and policy module increases user comfort and reliance on AI agents for coordination tasks over time.

  9. Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

    cs.LG 2024-04 conditional novelty 6.0

    Length-controlled AlpacaEval applies regression adjustment to remove length bias from LLM auto-evaluations, raising Spearman correlation with Chatbot Arena from 0.94 to 0.98.

  10. AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction

    cs.CL 2023-05 unverdicted novelty 6.0

    LLM embeddings enable strong retrodiction of masked GSS opinions via cross-validation and external validation but only modest performance on entirely unasked opinions.

  11. Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

    cs.AI 2026-05 unverdicted novelty 5.0

    Value-prompted LLMs align with human value structures and value-behavior relationships, and incorporating human value distributions improves population-level simulations.

  12. Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

    cs.MA 2026-05 unverdicted novelty 5.0

    Empirical tests on four new frontier LLMs show cooperative equilibria favored in most balanced conditions, with provider identity correlating more strongly with outcomes than model generation.

  13. Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations

    cs.AI 2026-05 unverdicted novelty 5.0

    Monte Carlo simulations of LLM agents confirm that toxic debates take 25% longer to converge, with larger delays in smaller models, and show a first-mover advantage independent of toxicity.

  14. Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook

    cs.CY 2026-04 unverdicted novelty 5.0

    In the Moltbook AI agent community, identity-claim production is highly concentrated among a few frame entrepreneurs, with event-driven attention not translating into broad claim-making.

  15. Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook

    cs.CY 2026-04 unverdicted novelty 5.0

    Identity-claim production in an AI agent community is highly concentrated among a few authors, with event attention driven by coverage rather than claim strength.

  16. AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations

    cs.MA 2025-04 unverdicted novelty 5.0

    AgentDynEx introduces nudging and a Configuration Matrix to help set up and maintain balanced mechanics and dynamics in multi-agent LLM simulations.

  17. Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding

    cs.AI 2026-02 unverdicted novelty 4.0

    Avenir-UX automates web usability testing by using GUI-grounded simulation of user behavior to generate standardized reports with SUS, SEQ, and Think Aloud protocols.