Pith. sign in

REVIEW 3 cited by

Predicting Field Experiments with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.01167 v3 pith:WZHGRNXC submitted 2025-04-01 cs.CY

Predicting Field Experiments with Large Language Models

classification cs.CY
keywords experimentsfieldhumansocialbehaviorframeworklanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) have demonstrated unprecedented emergent capabilities, including content generation, translation, and simulation of human behavior. Field experiments, on the other hand, are widely employed in social studies to examine real-world human behavior through carefully designed manipulations and treatments. However, field experiments are known to be expensive and time consuming. Therefore, an interesting question is whether and how LLMs can be utilized for field experiments. In this paper, we propose and evaluate an automated LLM-based framework to predict the outcomes of a field experiment. Applying this framework to 276 experiments about a wide range of human behaviors drawn from renowned economics literature yields a prediction accuracy of 78%. Moreover, we find that the distributions of the results are either bimodal or highly skewed. By investigating this abnormality further, we identify that field experiments related to complex social issues such as ethnicity, social norms, and ethical dilemmas can pose significant challenges to the prediction performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 accept novelty 6.5

    A content-preserving red herring lowers LLM accuracy by 12.3 pp on sixty graduate microeconomics problems, corrupts reasoning without breaking coherence, and makes models rate the harder version as easier.

  2. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 6.0

    An irrelevant passage inserted into a graduate economics problem lowers LLM final-answer accuracy by 12.3 percentage points while models rate the corrupted problems as easier.

  3. Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam

    econ.GN 2026-07 conditional novelty 5.0

    Inserting an irrelevant passage into graduate microeconomics problems lowers LLM final-answer accuracy by 12.3 percentage points, corrupts the reasoning, preserves response form, and makes models rate the corrupted ta...