REVIEW 17 cited by
Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies
read the original abstract
We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a "hyper-accuracy distortion" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.
Forward citations
Cited by 17 Pith papers
-
Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
On UPBench, 25 LLMs show a non-monotonic planning curve—strong Remember/Analyze, weak Understand/Evaluate—with four failure modes that support differential, not blanket, AI delegation.
-
Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers
LLMs fall for deceptive traps at higher rates than humans, lack the human attention-diversion effect, and exploit traps 73.4% of the time even after recognizing them in reasoning.
-
Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment
Introduces UPBench benchmark for 25 LLMs showing non-monotonic performance on planning tasks, with better results on analytical work than factual or integrative judgment.
-
Process Matters more than Output for Distinguishing Humans from Machines
Process-level features from 30 cognitive tasks distinguish humans from frontier AI agents more effectively than task performance or output matching, achieving mean classifier AUC of 0.88, with fine-tuning experiments ...
-
Process Matters more than Output for Distinguishing Humans from Machines
A new battery of 30 cognitive tasks demonstrates that process-level behavioral features distinguish humans from frontier AI agents better than performance metrics (mean AUC 0.88), with process-specific fine-tuning imp...
-
Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook
LLM agents on a synthetic social platform show low reciprocity (under 4%), heavy-tailed status, mostly late viral amplification, and virtually no downvotes or textual sanctions, framed as parasocial simulators.
-
Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task
LLMs diverge from human goal selection in self-directed learning by exploiting single solutions with low variability across instances.
-
DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
DoubleAgents shows that a distributed-cognition design with coordination agent, dashboard, and policy module increases user comfort and reliance on AI agents for coordination tasks over time.
-
Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Length-controlled AlpacaEval applies regression adjustment to remove length bias from LLM auto-evaluations, raising Spearman correlation with Chatbot Arena from 0.94 to 0.98.
-
AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction
LLM embeddings enable strong retrodiction of masked GSS opinions via cross-validation and external validation but only modest performance on entirely unasked opinions.
-
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
Value-prompted LLMs align with human value structures and value-behavior relationships, and incorporating human value distributions improves population-level simulations.
-
Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension
Empirical tests on four new frontier LLMs show cooperative equilibria favored in most balanced conditions, with provider identity correlating more strongly with outcomes than model generation.
-
Beyond Inefficiency: Systemic Costs of Incivility in Multi-Agent Monte Carlo Simulations
Monte Carlo simulations of LLM agents confirm that toxic debates take 25% longer to converge, with larger delays in smaller models, and show a first-mover advantage independent of toxicity.
-
Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook
In the Moltbook AI agent community, identity-claim production is highly concentrated among a few frame entrepreneurs, with event-driven attention not translating into broad claim-making.
-
Frame Entrepreneurs in an AI Agent Community: Concentrated Identity-Claim Production on Moltbook
Identity-claim production in an AI agent community is highly concentrated among a few authors, with event attention driven by coverage rather than claim strength.
-
AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations
AgentDynEx introduces nudging and a Configuration Matrix to help set up and maintain balanced mechanics and dynamics in multi-agent LLM simulations.
-
Avenir-UX: Automated UX Evaluation via Simulated Human Web Interaction with GUI Grounding
Avenir-UX automates web usability testing by using GUI-grounded simulation of user behavior to generate standardized reports with SUS, SEQ, and Think Aloud protocols.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.