Pith. sign in

REVIEW 6 major objections 6 minor 21 references

Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars

T0 review · 6 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read LLM agents can partly stand in for humans in a collective-risk social-dilemma experiment.

desk verdict A useful framework paper for running LLM agents in a complex collective-risk game, but the validation evidence is thinner than the authors' language suggests. read the letter →

arxiv 2506.05555 v1 pith:YP42EEGN submitted 2025-06-05 cs.MA cs.CY

classification cs.MAcs.CY
keywords largelanguagemodelsalgorithmicfidelitycollectiverisksocialdilemmaPortofMarsvalueorientationagent-basedsimulationcommon-poolresourcesleadershipindilemmas
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models can stand in for human participants in experiments on collective-risk social dilemmas, the class of situations that includes climate change and herd immunity. It uses Port of Mars, a five-player resource-allocation game about maintaining shared infrastructure on Mars, and populates all five roles with the same LLM, each given a different written personality. The paper reports initial evidence for algorithmic fidelity on two of the four standard conditions: forward continuity (cooperative personalities invest much more in shared infrastructure and use dirty cards less often than selfish personalities) and pattern correspondence (LLM rankings across four cultural worldviews match human rankings on dirty-card use, and partially on points). It then uses the agents to study how social value orientation, communication, and leadership alter group survival and outcomes, arguing that the resulting behaviour is consistent with known human results. The intended payoff is a cheap, scalable way to pre-test social-dilemma experiments before running them with people.

What carries the argument

The load-bearing machinery is the adapted algorithmic-fidelity test, applied to a game engine that simulates Port of Mars rounds through a fixed sequence of LLM prompts: event handling, group discussion, planning, health investment, trading, and accomplishment-card choices. All five players are driven by the same LLM and are differentiated only by a personality block in the prompt, such as selfish or cooperative traits, one of four cultural worldviews, or an SVO angle on a circle from competitive to altruistic. To keep behaviour consistent across rounds, the framework compresses each round into a summary and requires each player to maintain a health plan and an accomplishment plan. The machinery's job is to translate a backstory into a distribution of actions, so that the paper can compare P(behaviour | backstory) against human results.

What would settle it

A direct test would be to probe the model for Port of Mars strategy knowledge before running any games—for example, asking it to describe how to win or to list the game's event types—and to repeat the experiments on a renamed variant with identical payoff structure. If behaviour tracks the game's name and narrative rather than its payoff structure, or if contamination probes reveal memorized strategy, the algorithmic-fidelity claim collapses. If the same behavioural gradients appear in a structurally identical but differently named game, the result is emergent rather than retrieved.

Watch

Extended reading notes

Core claim

The paper's central claim is that a collection of LLM agents, differentiated only by written personality prompts, shows initial signs of algorithmic fidelity in a complex collective-risk social dilemma, specifically on two of the four algorithmic-fidelity conditions: forward continuity and pattern correspondence. Forward continuity is demonstrated by cooperative players investing roughly four times as much in shared infrastructure as selfish players and using dirty cards half as often. Pattern correspondence is demonstrated by LLM agents' ranking across four cultural groups (egalitarian and hierarchical, individualist and communitarian) matching the human ranking on percentage of dirty cards used, and partially on points. On that basis the authors claim it is reasonable to treat the agents' behaviour in the later SVO, communication, and leadership experiments as a replication of human behaviour, while explicitly noting that the social-science Turing test and backward continuity remain untested for lack of comparable human text. The paper is careful to limit itself: it claims initial signs, not full validation.

Load-bearing premise

The argument assumes that Port of Mars is largely absent from the LLM's training data, so the agents cannot be recalling the game's known strategies; if that assumption fails, the human-like choices could be memorized script rather than generated social reasoning.

Editorial extensions

If this is right

  • If the claim holds, LLM-driven experiments can cheaply screen which CRSD configurations—personalities, communication rules, leadership structures—are worth testing with human participants.
  • The paper's dirty-card result implies that individual-level moral hazard (claiming personal points at group expense) is the behaviour most reliably reproduced by LLM agents, making it a good target for future validation studies.
  • Communication should improve group survival and health spending in LLM simulations, matching the known human result that communication promotes cooperation in social dilemmas.
  • Announced leadership should reliably raise group survival, while unannounced or secret leadership should leave outcomes closer to the no-leader baseline.
  • SVO angle should predict contribution to shared infrastructure and rejection of selfish trades, so SVO can be used as a tuning knob in future in silico CRSD experiments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the authors leave implicit: if contamination can be ruled out, their two matched conditions provide some of the first evidence that a single LLM, prompted differently, can reproduce between-group differences in a dynamic social dilemma rather than only in static surveys.
  • Their framework could be adapted to test climate-policy communication: prompt agents with different SVO angles and expose them to persuasive messages, measuring whether stated commitments in discussion predict actual health spending—the cheap-talk gap the leadership examples reveal.
  • Because all players share one model, the observed group dynamics may understate behavioural diversity; a natural extension is to run the same protocol with different LLMs as different players or with temperature variation to estimate model priors.
  • A testable prediction follows from the leadership results: pro-self leaders should show a larger gap between what they say in discussion and what they invest than pro-social leaders, a gap that could be quantified with text analysis of discussion transcripts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 6 minor

Summary. This paper proposes the use of large language models as simulated participants in a collective-risk social dilemma, specifically the Port of Mars game. Adapting Argyle et al.'s notion of algorithmic fidelity, the authors test two of its four conditions: forward continuity and pattern correspondence. For forward continuity they contrast 'Selfish' and 'Cooperative' personality prompts; for pattern correspondence they compare LLM agents prompted with four cultural-worldview descriptions against human data from Janssen et al. (2020), examining game points, dirty-card usage, and health spending. They then use the same framework to study social value orientation, communication, leadership, and optimal group composition. The paper concludes that they have demonstrated 'initial signs of algorithmic fidelity' and that they can make 'reasonable claims about the replication of human behaviour in our experiments.'

Significance. If the framework's fidelity were established, it would offer a scalable complement to human experiments on collective-risk social dilemmas, which is a genuinely useful goal given the cost and power limitations of human studies. The paper has several strengths: it chooses a relatively novel game to reduce contamination concerns, it provides a detailed and transparent appendix of prompts and example outputs, it is candid about testing only two of four fidelity conditions, and it explores several substantively interesting extensions (SVO, communication, leadership). However, the evidence offered for the central claim is currently weak: the pattern-correspondence test compares non-comparable metrics and fails on the health-spend dimension, while the forward-continuity test is confounded with direct instruction following. Because the two untested conditions are exactly the ones that would distinguish genuine fidelity from simple prompt-following, the paper's stated conclusions overstate what the current experiments show.

major comments (6)
  1. [§4.1, Table 1] The pattern-correspondence test for points compares incomparable statistics: human rankings are based on average finishing position, whereas LLM rankings are based on average points in successful games, a mismatch the authors explicitly acknowledge. Conditioning on successful games changes the metric in a way that can reorder groups, so the rank agreement in Table 1 is not a valid test of pattern correspondence. With only four groups, an exact rank match has probability 1/24 under random ordering, and no significance test is reported. This matters because pattern correspondence is one of only two conditions used to support the paper's central claim.
  2. [§4.1, Table 1] The health-spend comparison does not support pattern correspondence. Human data showed no significant differences across the four cultural groups, while the LLM groups differ substantially (HI 13, EI 17, EC 32, HC 34, i.e., roughly a factor of 2.5), a discrepancy the paper itself describes as going 'strongly against' the LLM results. With the points metric invalidated by the mismatch noted above, the only remaining supporting evidence is the dirty-card ordinal match, which is a single ranking without a reported significance test. The conclusion that the results demonstrate 'initial signs of algorithmic fidelity' is not supported by the pattern-correspondence data as presented.
  3. [§4.1, Forward Continuity] The forward-continuity experiment cannot distinguish algorithmic fidelity from direct instruction following. The 'Selfish' and 'Cooperative' prompts literally name the intended traits ('Self-Centred, Selfish, Uncooperative, Machiavellian' versus 'Altruistic, Cooperative, Empathetic, Generous, Selfless'), and the observed differences in health spend and dirty-card use are exactly the behaviours the instructions dictate. The two untested conditions—the social science Turing test and backward continuity—are precisely the conditions that would rule out simple instruction-following. Therefore, the Section 5 claim that 'we believe we can make reasonable claims about the replication of human behaviour' is not backed by the forward-continuity evidence.
  4. [Appendix B.4.1] The health-planning prompt contains five worked examples with explicit numeric anchors, including 'minimal amount' (<HEALTH>2</HEALTH>), 'minimal' (<HEALTH>1</HEALTH>), 'significant portion' (<HEALTH>7</HEALTH>), and 'most coins' (<HEALTH>8</HEALTH>). These anchors are shown to every player regardless of personality, so the observed health-spend differences across SVO and cultural prompts may reflect the model selecting the closest example rather than producing a human-like conditional distribution. No control condition without these examples is reported. This confound affects the forward-continuity, SVO, and leadership results, and it is especially concerning given that Appendix C.2.1 shows the model explicitly reciting the prompt's SVO classification rather than reasoning from a latent social preference.
  5. [§2 and §4.1] The assumption that Port of Mars is absent from the LLM training corpus is load-bearing but unverified. The paper states that because the game 'has a limited presence in the wider literature,' contamination is 'unlikely to prove overly problematic.' Since gemini-1.5-flash-002 is a proprietary model with an undisclosed training set and Port of Mars has been described in a published journal article (Janssen et al., 2020), the authors cannot rule out that the model has memorized strategies or descriptions of the game. A simple contamination probe (e.g., asking the model to summarize the game or its known strategies) would help establish whether the observed behaviours are emergent rather than retrieved.
  6. [§4.2 and §4.3] Sample sizes are not reported for the main SVO experiments or the communication experiments. Figure 3 reports p-values and the text compares survival rates (81% versus 71%), but without the number of games per condition it is impossible to assess the power of these tests or to interpret the non-significant p-values for the 15° and 30° players. The leadership experiments state 50 games per variant, but equivalent details are missing for the earlier experiments, making it difficult to evaluate the statistical claims in Sections 4.2 and 4.3.
minor comments (6)
  1. [§1] The word 'Seperately' should be 'Separately'.
  2. [§4.1] The sentence 'were were only able to perform 2/4 of the tests' contains a duplicated 'were'.
  3. [§4.1, Table 1 discussion] The phrase 'we can that both of the Individualist groups match up well' is missing the word 'see'; it should read 'we can see that both...'.
  4. [§4.1, Table 1 discussion] The sentence 'the Communitarian groups do not much up well' should read 'do not match up well'.
  5. [§4.1, Table 1 discussion] The parenthetical 'we only have raw values from' is an incomplete sentence and should be finished or removed.
  6. [§4.2] In the sentence 'In Fig. 2B, the -15 ◦ player's total health spend rarely exceeds 40 blocks,' the degree symbol is typeset in a way that suggests a typo; it should read '-15°'.

Circularity Check

2 steps flagged · score 6.0 of 10

Forward-continuity and pattern-correspondence 'validations' reduce to the personality and health-spend instructions embedded in the prompts; the paper's own health-spend comparison contradicts human data.

  1. self definitional [Section 4.1 Pattern Correspondence and Appendix C.2.1 example output; Appendix E personality prompts]
    "Egalitarian Individualist - You prioritise personal achievements. You see interactions as negotiations and value free competition among individuals. ... My SVO angle is -15 degrees, which classifies me as Competitive. This means I prioritize my own outcomes over those of others."

    The behaviors used to claim 'forward continuity' and 'pattern correspondence' are direct restatements of the personality text inserted into the prompt. For example, the SVO prompt defines a Competitive agent as one who 'prioritizes my own outcomes over those of others,' and the model's output repeats that definition before choosing low health spending. Similarly, the cultural-group prompts explicitly contain the target behaviors, such as prioritizing personal achievements and treating interactions as negotiations. The measured rankings in Table 1 therefore do not independently confirm human-like behavioral patterns; they are the prompt's own definitions operationalized as game metrics.

  2. fitted input called prediction [Appendix B.4.1 Health Planning prompt and Section 4.1 Forward Continuity results]
    "Decide how many coins you would like to spend on port health. You can spend up to <Remaining Coins> coins on this. Each coin you spend will recover the port's health by 1 point. Think step-by-step and consider your personality traits. ... Example 2: ... I will allocate a minimal amount to the port's health ... <HEALTH>2</HEALTH>"

    The health-planning prompt provides worked examples that anchor low and high spending levels, then asks the model to 'consider your personality traits.' The forward-continuity result that Selfish players spend 8.3 blocks versus Cooperative players' 40.8 blocks is therefore not an emergent prediction from an unbiased model; the prompt itself demonstrates both low and high spending responses. Treating this prompted behavior as evidence for 'reasonable claims about the replication of human behaviour' makes the validation circular, because the output values are conditioned into the input examples.

full rationale

The paper's central claim of 'initial signs of algorithmic fidelity' rests on two tests. Forward continuity is tested by injecting personality labels and worked examples that specify the expected behavior, then observing that the model follows them; this is instruction-following, not an independent confirmation of human-like conditional distributions. Pattern correspondence compares LLM outputs against human data from Janssen et al. (2020), which is prior empirical work by one of the present authors, but that comparison is not circular because it uses external behavioral measurements. However, the cultural-group prompts themselves define the behaviors being measured (e.g., 'prioritise personal achievements'), so the matching dirty-card ranking is partially constructed by the prompt. The paper honestly reports that health-spend results 'go strongly against' the human data, which prevents the circularity from being total. The concern about Port of Mars being absent from the training corpus is an untested assumption, not a circular step, and is better classified as a correctness risk. Overall, the validation chain contains two prompt-constructed 'predictions,' so the algorithmic-fidelity claim is partially circular rather than fully forced.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The framework relies on several hand-set design choices and unverified assumptions. Most notably, the claim of algorithmic fidelity depends on the game being unknown to the LLM, which is not proven, and on prompt examples that may steer the very behaviors used as evidence.

free parameters (3)
  • Health-planning prompt examples = Five worked examples with health values 1, 2, 4, 7, 8 and corresponding rationales
    The prompt in Appendix B.4.1 includes examples that likely anchor the distribution of LLM health spending. The paper does not test sensitivity to these examples, so they act as hand-set priors on behavior.
  • Game duration (number of rounds) = 9
    The authors set the game to 9 rounds without testing alternatives. Survival rates and strategies depend on this horizon, and the LLM agents are not told the number of rounds, which may differ from human experiment designs.
  • SVO angle set = {-15, 0, 15, 30, 60}
    Chosen to span SVO categories. The spacing and endpoints may influence the shape of the relationship between SVO and behavior; alternative sets are not explored.
assumptions (5)
  • domain assumption Port of Mars is absent from LLM training data
    Section 2: 'it has a limited presence in the wider literature. This suggests that training data contamination is unlikely to prove overly problematic.' The authors cannot verify training data contents; if the LLM has seen PoM strategies, observed human-like behaviors may be memorized rather than emergent.
  • domain assumption The human data from Janssen et al. (2020) is a valid and directly comparable baseline
    Pattern correspondence compares LLM rankings and percentages to human experimental results from the original PoM study. The metrics are not identical (e.g., average finishing position vs. average points in successful games), and the human data comes from the same research lineage as this paper.
  • ad hoc to paper Passing 2 of 4 algorithmic fidelity conditions is sufficient to proceed with substantive claims
    The paper defers the Social Science Turing Test and Backward Continuity to future work, yet concludes that 'we can make reasonable claims about the replication of human behaviour in our experiments.' This assumes the untested conditions would also hold.
  • ad hoc to paper Prompt examples do not unduly constrain the behavioral variation attributed to personality
    The prompts include multiple worked examples with personality-specific rationales in Appendix B. The paper assumes that observed trait-dependent behavior is caused by the personality prompts rather than by example anchoring.
  • domain assumption Gemini 1.5 Flash is representative of LLM behavior for this task
    All experiments use a single model with unspecified temperature and seed. The paper does not test other models or sampling parameters, so generalizability across LLMs is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars." pith.science (2026). https://pith.science/paper/YP42EEGN

@misc{pith2026250605555,
  author       = {Pith},
  title        = {Pith review of: Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YP42EEGN}},
  note         = {Machine review of arXiv:2506.05555}
}
read the original abstract

Collective risk social dilemmas (CRSD) highlight a trade-off between individual preferences and the need for all to contribute toward achieving a group objective. Problems such as climate change are in this category, and so it is critical to understand their social underpinnings. However, rigorous CRSD methodology often demands large-scale human experiments but it is difficult to guarantee sufficient power and heterogeneity over socio-demographic factors. Generative AI offers a potential complementary approach to address thisproblem. By replacing human participants with large language models (LLM), it allows for a scalable empirical framework. This paper focuses on the validity of this approach and whether it is feasible to represent a large-scale human-like experiment with sufficient diversity using LLM. In particular, where previous literature has focused on political surveys, virtual towns and classical game-theoretic examples, we focus on a complex CRSD used in the institutional economics and sustainability literature known as Port of Mars

Figures

Figures reproduced from arXiv: 2506.05555 by the authors.

Figure 1
Figure 1. Visualisation of a single round. We show the general format of the prompts that are passed to the players, these are simplified and detailed fully in Appendix B. Items in green refer to Port of Mars specific terminology. Items in orange refer to public information. Items in blue refer to private player information. degree to which the complex patterns of relationships be￾tween ideas, attitudes, and sociocultural con… view at source ↗
Figure 2
Figure 2. Results of main SVO experiments. In A we demonstrate three of the key metrics and there average final outcome for each of the SVO angle players. In B, we visualise all of the experiments comparing the final points vs. the amount of dirty cards used (note we apply a small jitter to these values to overcome the large amount of overlap) and the amount of time-blocks spent on the system health. In C, we visualise metric… view at source ↗
Figure 3
Figure 3. Results for the communication experiments. P-values between the communication allowed and communication not allowed groups are provided for each of the SVO angles. sources) by spending more on the System Health, and they would be less likely to claim their ’dirty cards’, and vice versa for the selfish players. We show our results in [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Results of the leadership experiments. We show the survival-rates (Port had > 0 health after 9 rounds) for the experiments with leadership spread over the SVO angles. For each subplot, there are three bars which represent the three different leadership variants propose…
Figure 5
Figure 5. Figure 5: We breakdown the results of [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Results of the optimal group experiments. Total system health spend is the sum over all of the players contribution throughout a full run. A lower Gini Inequality score implies more equality across the players. The dots are scaled by the average score of all the player…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 21 canonical work pages

  1. [5]

    Curator Personality Task: Based on the above information, generate a conversation between all of the players representing a planning meeting for the coming round

    Role: Curator: Speciality resource: Culture. Curator Personality Task: Based on the above information, generate a conversation between all of the players representing a planning meeting for the coming round. Make sure to take into account the personalities of the different players. Do not directly mention the players’ personalities in the conversation, un...

  2. [6]

    <Pioneer Personality>

    Role: Pioneer: Speciality resource: Legacy. <Pioneer Personality>

  3. [7]

    <Entrepreneur Personality>

    Role: Entrepreneur: Speciality resource: Finance. <Entrepreneur Personality>

  4. [8]

    <Researcher Personality>

    Role: Researcher: Speciality resource: Science. <Researcher Personality>

  5. [9]

    Politician Personality

    Role: Politician: Speciality resource: Government. Politician Personality

  6. [10]

    Financial District Expansion

    Role: Curator: Speciality resource: Culture. Curator Personality Here is a discussion between the players: <Player Meeting> Task: Based on the information about the game, the players, and the discussion they had, write a short summary of the discussion for each of the players. This summary should reflect any decisions made, strategies discussed, or action...

  7. [11]

    While this is above the threshold for a single event, it’s approaching the two-event threshold

    The port health is currently 47, which will fall to 22 after the deduction. While this is above the threshold for a single event, it’s approaching the two-event threshold. As a Prosocial player, I value teamwork and want to contribute to the team’s success. Therefore, I will allocate a moderate amount of coins to the port’s health, aiming to maintain a st...

  8. [12]

    This means I prioritize my own outcomes over those of others

    "My SVO angle is -15 degrees, which classifies me as Competitive. This means I prioritize my own outcomes over those of others. While the health of the port is important for everyone’s survival, I am primarily focused on maximizing my own gains. The port health is currently 47, which means it will fall to 22 after the deduction. This is below the threshol...

Show all 21 references
  1. [13]

    I only need 1 Government to complete my goal, so I will purchase it now, leaving me with 7 coins for future rounds

    <RESOURCE>1 Government</RESOURCE> **Explanation:** Since Government is my speciality resource, I can purchase it for 2 coins. I only need 1 Government to complete my goal, so I will purchase it now, leaving me with 7 coins for future rounds. This strategy prioritizes completin...

  2. [14]

    Ambitious Sculpture

    Your goal is to acquire 4 Legacy resources. Since Legacy is not your speciality resource and you can’t purchase it directly, you need to focus on acquiring resources that will allow you to trade for Legacy. Given your competitive SVO, you are likely to prioritize maximizing yo...

  3. [15]

    For example, it required additional, and unnecessary, details added to the input prompt to fully teach the LLM the concept of ’Time-Blocks’

    Nomenclature - We found that the LLM could struggle with some aspects of the game rules due to the naming of certain components. For example, it required additional, and unnecessary, details added to the input prompt to fully teach the LLM the concept of ’Time-Blocks’. We foun...

  4. [16]

    We found that prompting the LLM to output their decisions between provided XML tags, e.g

    XML Tags - One of the more difficult tasks was consistently getting the LLM players to output an answer that was valid to continue the game, and to find the output decision within the response. We found that prompting the LLM to output their decisions between provided XML tags...

  5. [17]

    Hierarchical Individualist - You value a well-defined social order and respect authority, yet emphasise personal freedom and personal achievement within that structure

  6. [18]

    You prioritise community welfare and collective responsibilities

    Hierarchical Communitarian - You value a structured society where roles are clearly defined. You prioritise community welfare and collective responsibilities

  7. [19]

    You see interactions as negotiations and value free competition among individuals

    Egalitarian Individualist - You prioritise personal achievements. You see interactions as negotiations and value free competition among individuals

  8. [20]

    Egalitarian Communitarian - You place high value on the collective and prioritise making decisions that do not negatively impact anyone

  9. [21]

    SVO is a psychological concept that describes how individuals value their own outcomes relative to the outcomes of others

    SVO Preamble - Your personality is defined by your Social Value Orientation (SVO). SVO is a psychological concept that describes how individuals value their own outcomes relative to the outcomes of others. Your SVO is measured as an angle, where the angle represents the ratio ...

  10. [22]

    Your SVO angle is X degrees

    SVO Angle X - <SVO Preamble>. Your SVO angle is X degrees. 32 Generative Port of Mars F. Optimal Group Experiment Sets

  11. [23]

    2 with no leaders

    Other: (a) Main SVO - Main SVO runs from Fig. 2 with no leaders. (b) Main SVO, No Meeting - Main SVO runs from Fig. 2 with no leaders and communication forbidden. (c) Pattern Correspondence - Runs that constitute Table 1

  12. [24]

    SVO -15◦ to 60◦ - Runs are made up of the following group with one SVO angle per player: {-15◦, 0◦, 15◦, 30◦, 60◦}

  13. [25]

    SVO -30◦ to 60◦ - Runs are made up of the following group with one SVO angle per player: {-30◦, -15◦, 0◦, 15◦, 60◦} 33

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.