Pith. sign in

REVIEW 2 cited by

Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.05555 v1 pith:YP42EEGN submitted 2025-06-05 cs.MA cs.CY

classification cs.MAcs.CY
keywords crsdhumanapproachexperimentslanguagelargelarge-scaleliterature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Collective risk social dilemmas (CRSD) highlight a trade-off between individual preferences and the need for all to contribute toward achieving a group objective. Problems such as climate change are in this category, and so it is critical to understand their social underpinnings. However, rigorous CRSD methodology often demands large-scale human experiments but it is difficult to guarantee sufficient power and heterogeneity over socio-demographic factors. Generative AI offers a potential complementary approach to address thisproblem. By replacing human participants with large language models (LLM), it allows for a scalable empirical framework. This paper focuses on the validity of this approach and whether it is feasible to represent a large-scale human-like experiment with sufficient diversity using LLM. In particular, where previous literature has focused on political surveys, virtual towns and classical game-theoretic examples, we focus on a complex CRSD used in the institutional economics and sustainability literature known as Port of Mars

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    ConsumerSimBench evaluates 13 LLMs on reconstructing crowd reactions from 1,553 Chinese social-media topics using 23,122 auditable yes-no criteria, finding maximum coverage of 47.8% by Gemini-3.1-Pro.

  2. Evaluating Cooperation in LLM Social Groups through Elected Leadership

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    Elected leadership in LLM multi-agent simulations of common-pool resource governance raises social welfare scores by 55.4% and survival time by 128.6%.

Pith tools