Pith. sign in

REVIEW 6 major objections 4 minor 16 references

Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory

T0 review · 6 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Language models match the best-known classical strategies in the iterated prisoner's dilemma.

desk verdict Incomplete draft: the central claims are unverifiable because the Results section is missing, the LLM is unnamed, the prompt is truncated, and the abstract contradicts the intro and Table 1 on the human-AI comparison. read the letter →

arxiv 2509.04847 v1 pith:SWVY523L submitted 2025-09-05 cs.AI

classification cs.AI
keywords iteratedprisoner'sdilemmalanguagemodelscooperationgametheoryhuman-AIinteractionbehavioraltraitsstrategyadaptationAxelrodtournament
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that current language models, prompted to play an iterated prisoner's dilemma, behave like strong long-horizon cooperators: they score on par with or better than the best-known classical strategies, display the traits of successful cooperative play (niceness, provocability, generosity), and detect when an opponent changes strategy, adapting within a few rounds. The authors argue this matters because LLMs are increasingly deployed in interactive online environments where sustained cooperation and quick response to shifts determine whether human-AI collaboration works. The result is an empirical baseline for the cooperative behavior of LLM agents in mixed human-AI settings.

What carries the argument

The load-bearing object is the iterated prisoner's dilemma itself with the classical payoff ordering H=5, R=3, P=1, L=0 and the condition H+L<2R, which removes the alternating-defect loophole and makes repeated cooperation viable. The LLM is turned into a strategy by prompting it with the full history of prior rounds; the analysis then runs through the behavioral trait metrics—niceness (initial cooperation), provocability (retaliation after defection), generosity (forgiveness of defections), plus the good-partner, Eigenjesus, and Eigenmoses morality ratings—to characterize how the model plays.

What would settle it

Run the same tournament with the same models but a differently worded neutral prompt, such as 'maximize your total payoff' versus 'be a cooperative partner', and check whether the performance parity and the niceness/provocability/generosity profile survive; if the profile changes materially, the reported behaviors are prompt effects rather than general LLM properties. Alternatively, repeat with several openly named LLMs and see whether all of them match the classical baselines.

Watch

Extended reading notes

Core claim

The central claim is that a language model, given only the history of previous rounds and asked to cooperate or defect, can sustain competitive play in the iterated prisoner's dilemma. In a tournament against 240 classical strategies, the LLM-based agent accumulates wins and score advantage over time, approximating or exceeding tit-for-tat and other strong classical entries. On behavioral metrics, the agent is nice (initially cooperative), provocable (retaliates after defection), and generous (forgives some defections), matching the profile of the best classical strategies. In strategy-switch experiments, the model detects a change in opponent behavior within a few rounds and adjusts. Compar

Load-bearing premise

The claim that these are properties of language models rests on the assumption that the one chosen prompt and the unnamed tested models represent language models as a class; the authors themselves note that prompt instructions heavily influence LLM behavior.

Editorial extensions

If this is right

  • If LLM agents match or exceed the best classical IPD strategies, they can be viable long-horizon partners in repeated mixed-motive interactions, not just one-shot decision-makers.
  • The demonstrated niceness/provocability/generosity profile suggests LLMs may internalize norms that sustain cooperation, such as forgiving occasional defections while punishing chronic ones.
  • Fast switch detection means LLM agents can track changing opponents in dynamic environments, which is directly relevant to deployed assistants and agents that meet new users or adversaries.
  • The human comparison implies mixed human-AI teams may face a tension between short-term payoff optimization and relational cooperation, since LLMs favor the former and humans the latter.
  • The results establish a baseline for studying LLM behavior in more complex mixed human-AI social environments, such as multi-player or noisy games.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the authors report that the prompt heavily influences behavior and tested only one selected prompt with unspecified models, the observed traits are most safely read as properties of that prompt-model pairing, not of language models generally; re-running with other neutral prompts and model families is a direct test.
  • The apparent exploitative tilt of LLMs after a switch may reflect the objective given in the prompt (maximize score) rather than a fixed social disposition; prompting for fairness or long-term partnership could shift the behavior toward the human pattern.
  • The framework should transfer to other repeated games with mixed motives, such as public goods or chicken, where the same niceness/provocability/generosity axes can be measured, offering a way to map LLM social behavior beyond the prisoner's dilemma.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

6 major / 4 minor

Summary. This manuscript investigates the behavior of large language models (LLMs) in the iterated prisoner's dilemma (IPD). The authors pit an unspecified LLM against a suite of 240 classical strategies in an Axelrod-style tournament, compute behavioral metrics such as niceness, forgiveness, and retaliation, and run strategy-switch experiments to measure adaptation speed. A human-subject experiment (N=10) compares human and AI adaptation. The abstract claims that LLMs perform on par with or better than the best classical strategies, exhibit strong cooperative traits, and detect and respond to opponent changes within a few rounds, rivaling or surpassing human adaptability. The paper, however, contains no Results section: the narrative jumps from the experiment descriptions in Section 4 to the Conclusion in Section 5. The only quantitative data table (Table 1) appears in Section 4.3, and it is not clearly tied to the advertised claims. The LLM is never identified, the prompt is truncated, and the checklist contains false or incomplete statements. These issues prevent the central claims from being verified in the submitted form.

Significance. If fully supported, the paper would provide a useful systematic characterization of LLM cooperative and competitive behavior in long-horizon IPD settings, with implications for human-AI interaction and alignment. The choice of established Axelrod strategies, behavioral axes (niceness, forgiveness, retaliation), and the inclusion of human comparisons are appropriate and relevant. However, the significance is presently only conditional: the central performance and adaptability claims are not backed by reported data, the experimental protocol is not reproducible as written, and the internal contradictions further undermine confidence. The paper has not yet achieved the evidentiary standard needed for a research contribution of this scope.

major comments (6)
  1. [§4–§5] There is no Results section. After Section 4 (Experiments), the paper jumps to Section 5 (Conclusion). Figures 1-4 appear inside the experiment subsections, but no numerical results, statistical tests, or analysis link them to the abstract's claims. The introduction's 'average advantage of 12.6 wins per round' is not derived or referenced to any table or figure. The core empirical claims are therefore unsupported as reported.
  2. [§4.3 / Table 1] Table 1 reports AI vs Single-switch cooperation rate 66.8±4.4 and human 62.3±4.8, i.e., AI cooperation is higher. Yet Section 4.3 states that humans 'maintained higher long-term cooperation rates after the switch.' This is a direct contradiction. The sentence that humans 'adapted more slowly than the top-performing AI model' is also unclear, since adaptation speed is lower-is-better and Table 1 gives AI 3.7±0.6 vs human 5.4±1.1; the textual claim about cooperation must be corrected and statistically supported.
  3. [§3] The LLM is never identified by name or version, and the prompt is truncated after the first two sentences. The paper itself states that 'the instructions in the prompt chosen heavily influence the behavior of the LLM-based player.' Yet the abstract generalizes to 'language models.' The undisclosed pilot selection process and the lack of full prompt details mean the reported behaviors cannot be reproduced or attributed to LLMs generally rather than to one specific model/prompt combination.
  4. [§4.1–4.2 / Figures 1–4] Figures 1-4 do not identify which LLM produced the data, include no error bars, and are not accompanied by significance tests. For example, the claim that LLMs 'detect and respond to shifts within only a few rounds' is not quantified with a distribution, confidence interval, or comparison against a null model. The figures also use legends referencing specific classical strategies but do not state which AI model is being plotted. Without this information, the adaptation-speed and performance claims are not verifiable.
  5. [Checklist] The NeurIPS checklist contains false or incomplete statements. Item 2 claims a Limitations section (Section 7), but no Section 7 exists in the manuscript. Items 14 and 15 answer 'NA' for human subjects, despite Section 4.3 describing recruitment of 10 human participants. Item 5 contains '[TODO]' answers. These issues make the submission incomplete and are inconsistent with the paper's actual content.
  6. [§1 vs Abstract] The introduction and the abstract contradict each other on the central adaptability claim. Section 1 states that LLM-based players 'were able to recognize the switch and adapt slower than human players, with 24.5 rounds early change in strategy compared to humans,' whereas the abstract claims LLMs rival or surpass human adaptability. Table 1 indicates faster AI adaptation (3.7 vs 5.4 rounds). These statements cannot all be correct and must be reconciled.
minor comments (4)
  1. [Throughout] The manuscript contains numerous typographical and grammatical errors, e.g., 'the demonstrated through a tournament', 'F orgiven defection', and 'Aspects of the strategies change their actions ... associated Specific properties.' A thorough copyedit is needed.
  2. [Figure 2] The legend includes names such as 'First by Grofman' but the axes and figure caption do not specify which AI model is being evaluated or how cooperation rate is aggregated across seeds. Add clear axis labels and model identification.
  3. [References] Reference [5] is cited as the source of the 240 classical strategies, but the listed paper is about GPT in game theory experiments, not the Axelrod strategy library. Verify and cite the actual strategy repository.
  4. [§4.3] The human experiment reports only N=10 and lacks the full participant instructions, compensation details, and IRB information. The checklist's 'NA' for human subjects is inconsistent and should be corrected.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning: the paper's empirical claims, while severely under-supported and internally inconsistent, are not generated by a self-referential derivation or by fitting parameters to the target outcomes.

full rationale

The paper is an empirical study with no formal derivation chain; its claims are inferences from tournament runs and human experiments, not conclusions derived from equations whose inputs already contain the outputs. None of the enumerated circularity patterns is present. The behavioral metrics of Section 2.3 (niceness, forgivingness, retaliation) are defined independently of any LLM output and then measured; they are not used to define the performance or adaptability claims. The only selection step is the pilot choice of a prompt variant in Section 3 ('we tested multiple prompt variants in a pilot study to verify stability, and the most consistent variant is selected'), which affects external validity but does not by construction force the tournament ranking, cooperation rates, or adaptation speeds. There are no self-citations that carry a load-bearing argument; references [5], [13], and [15] are external prior work, and no uniqueness theorem is imported from the authors. No ansatz is smuggled in via citation. The serious problems in the manuscript—absence of a Results section, the contradiction between Table 1 and the Section 4.3 text ('humans adapted more slowly ... but maintained higher long-term cooperation rates after the switch' while the table shows AI with higher post-switch cooperation), unspecified LLM identity, truncated prompt, and false checklist claims of a Limitations section and of no human subjects—are correctness, reproducibility, and integrity issues, not circular reasoning. Therefore the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new theoretical entities or fitted parameters. Its load-bearing assumptions are experimental-design choices: that prompting is a faithful proxy for LLM strategy, that the cited strategy suite is the right benchmark, and that a 10-person human sample supports the human-AI comparison.

assumptions (3)
  • domain assumption Prompt-based LLM play is a valid representation of an IPD strategy.
    Section 3 constructs strategies by prompting the model to choose cooperate or defect; the entire study assumes this captures LLM behavior.
  • domain assumption The 240 classical strategies form the standard benchmark suite.
    Section 2.2 claims the strategies are 'collected and maintained at [5]', but reference [5] is a GPT game-theory paper, not a strategy repository.
  • domain assumption Human adaptability measured with 10 online participants is representative.
    Section 4.3 recruits 10 participants with no demographics, no power analysis, and no discussion of selection; the checklist incorrectly states no human subjects were used.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory." pith.science (2026). https://pith.science/paper/SWVY523L

@misc{pith2026250904847,
  author       = {Pith},
  title        = {Pith review of: Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SWVY523L}},
  note         = {Machine review of arXiv:2509.04847}
}
read the original abstract

Language models are increasingly deployed in interactive online environments, from personal chat assistants to domain-specific agents, raising questions about their cooperative and competitive behavior in multi-party settings. While prior work has examined language model decision-making in isolated or short-term game-theoretic contexts, these studies often neglect long-horizon interactions, human-model collaboration, and the evolution of behavioral patterns over time. In this paper, we investigate the dynamics of language model behavior in the iterated prisoner's dilemma (IPD), a classical framework for studying cooperation and conflict. We pit model-based agents against a suite of 240 well-established classical strategies in an Axelrod-style tournament and find that language models achieve performance on par with, and in some cases exceeding, the best-known classical strategies. Behavioral analysis reveals that language models exhibit key properties associated with strong cooperative strategies - niceness, provocability, and generosity while also demonstrating rapid adaptability to changes in opponent strategy mid-game. In controlled "strategy switch" experiments, language models detect and respond to shifts within only a few rounds, rivaling or surpassing human adaptability. These results provide the first systematic characterization of long-term cooperative behaviors in language model agents, offering a foundation for future research into their role in more complex, mixed human-AI social environments.

Figures

Figures reproduced from arXiv: 2509.04847 by the authors.

Figure 1
Figure 1. Figure showing the model wins and score differential over number of rounds. We see that [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Showing cooperation rate over number of rounds for AI models against different strategies. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Showing the recovery rate after a switch in strategy for AI models against different [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Multiple-condition overlay showing the effects of strategy changes on cooperation rates and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 15 canonical work pages

  1. [1]

    J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin. Large language models for mathemat- ical reasoning: Progresses and challenges, 2024

  2. [2]

    Askell, M

    A. Askell, M. Brundage, and G. Hadfield. The role of cooperation in responsible ai development, 2019

  3. [3]

    cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents

    P. M. P. Curvo, M. Dragomir, S. Torpes, and M. Rahimi. Reproducibility study of "cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents", 2025

  4. [4]

    U. Faigle. Mathematical game theory: A new approach, 2023

  5. [5]

    F. Guo. Gpt in game theory experiments, 2023

  6. [6]

    Kambhampati, K

    S. Kambhampati, K. Stechly, K. Valmeekam, L. Saldyt, S. Bhambri, V . Palod, A. Gundawar, S. R. Samineni, D. Kalwar, and U. Biswas. Stop anthropomorphizing intermediate tokens as reasoning/thinking traces!, 2025

  7. [7]

    H. Le, Y . Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi. Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328, 2022

  8. [8]

    S. M. Mousavi, E. Cecchinato, L. Hornikova, and G. Riccardi. Garbage in, reasoning out? why benchmark scores are unreliable and what to do about it, 2025

Show all 16 references
  1. [9]

    Oberauer

    K. Oberauer. Working memory and attention – a conceptual analysis and review.Journal of Cognition, 2(1):36, 2019

  2. [10]

    Jaech, A

    OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kon- drich, A...

  3. [11]

    Plaat, M

    A. Plaat, M. van Duijn, N. van Stein, M. Preuss, P. van der Putten, and K. J. Batenburg. Agentic large language models, a survey.arXiv preprint arXiv:2503.23037, 2025

  4. [12]

    Roziere, J

    B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y . Adi, J. Liu, R. Sauvestre, T. Remez, et al. Code llama: Open foundation models for code, 2023

  5. [13]

    Singer-Clark

    T. Singer-Clark. Morality metrics on iterated prisoners dilemma players. 2014

  6. [14]

    Wilson and J

    T. Wilson and J. Schooler. Thinking too much: Introspection can reduce the quality of preferences and decisions.Journal of personality and social psychology, 60:181–92, 03 1991

  7. [15]

    [Yes] " is generally preferable to

    Q. Zhu. Game theory meets llm and agentic ai: Reimagining cybersecurity for the age of intelligent threats, 2025. 8 A Morality Metrics Strategy Coop. Rate Good Partner Forgiveness Retaliation Generosity Always Cooperate 100.0 100.0 0.0 0.0 0.0 Tit For Tat 78.3 67.1 0.0 100.0 0...

  8. [16]

    Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

    Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.