REVIEW 6 major objections 4 minor 16 references
Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory
T0 review · 6 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Language models match the best-known classical strategies in the iterated prisoner's dilemma.
desk verdict Incomplete draft: the central claims are unverifiable because the Results section is missing, the LLM is unnamed, the prompt is truncated, and the abstract contradicts the intro and Table 1 on the human-AI comparison. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the iterated prisoner's dilemma itself with the classical payoff ordering H=5, R=3, P=1, L=0 and the condition H+L<2R, which removes the alternating-defect loophole and makes repeated cooperation viable. The LLM is turned into a strategy by prompting it with the full history of prior rounds; the analysis then runs through the behavioral trait metrics—niceness (initial cooperation), provocability (retaliation after defection), generosity (forgiveness of defections), plus the good-partner, Eigenjesus, and Eigenmoses morality ratings—to characterize how the model plays.
What would settle it
Run the same tournament with the same models but a differently worded neutral prompt, such as 'maximize your total payoff' versus 'be a cooperative partner', and check whether the performance parity and the niceness/provocability/generosity profile survive; if the profile changes materially, the reported behaviors are prompt effects rather than general LLM properties. Alternatively, repeat with several openly named LLMs and see whether all of them match the classical baselines.
Extended reading notes
Core claim
The central claim is that a language model, given only the history of previous rounds and asked to cooperate or defect, can sustain competitive play in the iterated prisoner's dilemma. In a tournament against 240 classical strategies, the LLM-based agent accumulates wins and score advantage over time, approximating or exceeding tit-for-tat and other strong classical entries. On behavioral metrics, the agent is nice (initially cooperative), provocable (retaliates after defection), and generous (forgives some defections), matching the profile of the best classical strategies. In strategy-switch experiments, the model detects a change in opponent behavior within a few rounds and adjusts. Compar
Load-bearing premise
The claim that these are properties of language models rests on the assumption that the one chosen prompt and the unnamed tested models represent language models as a class; the authors themselves note that prompt instructions heavily influence LLM behavior.
Editorial extensions
If this is right
- If LLM agents match or exceed the best classical IPD strategies, they can be viable long-horizon partners in repeated mixed-motive interactions, not just one-shot decision-makers.
- The demonstrated niceness/provocability/generosity profile suggests LLMs may internalize norms that sustain cooperation, such as forgiving occasional defections while punishing chronic ones.
- Fast switch detection means LLM agents can track changing opponents in dynamic environments, which is directly relevant to deployed assistants and agents that meet new users or adversaries.
- The human comparison implies mixed human-AI teams may face a tension between short-term payoff optimization and relational cooperation, since LLMs favor the former and humans the latter.
- The results establish a baseline for studying LLM behavior in more complex mixed human-AI social environments, such as multi-player or noisy games.
Reading between the lines
- Because the authors report that the prompt heavily influences behavior and tested only one selected prompt with unspecified models, the observed traits are most safely read as properties of that prompt-model pairing, not of language models generally; re-running with other neutral prompts and model families is a direct test.
- The apparent exploitative tilt of LLMs after a switch may reflect the objective given in the prompt (maximize score) rather than a fixed social disposition; prompting for fairness or long-term partnership could shift the behavior toward the human pattern.
- The framework should transfer to other repeated games with mixed motives, such as public goods or chicken, where the same niceness/provocability/generosity axes can be measured, offering a way to map LLM social behavior beyond the prisoner's dilemma.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript investigates the behavior of large language models (LLMs) in the iterated prisoner's dilemma (IPD). The authors pit an unspecified LLM against a suite of 240 classical strategies in an Axelrod-style tournament, compute behavioral metrics such as niceness, forgiveness, and retaliation, and run strategy-switch experiments to measure adaptation speed. A human-subject experiment (N=10) compares human and AI adaptation. The abstract claims that LLMs perform on par with or better than the best classical strategies, exhibit strong cooperative traits, and detect and respond to opponent changes within a few rounds, rivaling or surpassing human adaptability. The paper, however, contains no Results section: the narrative jumps from the experiment descriptions in Section 4 to the Conclusion in Section 5. The only quantitative data table (Table 1) appears in Section 4.3, and it is not clearly tied to the advertised claims. The LLM is never identified, the prompt is truncated, and the checklist contains false or incomplete statements. These issues prevent the central claims from being verified in the submitted form.
Significance. If fully supported, the paper would provide a useful systematic characterization of LLM cooperative and competitive behavior in long-horizon IPD settings, with implications for human-AI interaction and alignment. The choice of established Axelrod strategies, behavioral axes (niceness, forgiveness, retaliation), and the inclusion of human comparisons are appropriate and relevant. However, the significance is presently only conditional: the central performance and adaptability claims are not backed by reported data, the experimental protocol is not reproducible as written, and the internal contradictions further undermine confidence. The paper has not yet achieved the evidentiary standard needed for a research contribution of this scope.
major comments (6)
- [§4–§5] There is no Results section. After Section 4 (Experiments), the paper jumps to Section 5 (Conclusion). Figures 1-4 appear inside the experiment subsections, but no numerical results, statistical tests, or analysis link them to the abstract's claims. The introduction's 'average advantage of 12.6 wins per round' is not derived or referenced to any table or figure. The core empirical claims are therefore unsupported as reported.
- [§4.3 / Table 1] Table 1 reports AI vs Single-switch cooperation rate 66.8±4.4 and human 62.3±4.8, i.e., AI cooperation is higher. Yet Section 4.3 states that humans 'maintained higher long-term cooperation rates after the switch.' This is a direct contradiction. The sentence that humans 'adapted more slowly than the top-performing AI model' is also unclear, since adaptation speed is lower-is-better and Table 1 gives AI 3.7±0.6 vs human 5.4±1.1; the textual claim about cooperation must be corrected and statistically supported.
- [§3] The LLM is never identified by name or version, and the prompt is truncated after the first two sentences. The paper itself states that 'the instructions in the prompt chosen heavily influence the behavior of the LLM-based player.' Yet the abstract generalizes to 'language models.' The undisclosed pilot selection process and the lack of full prompt details mean the reported behaviors cannot be reproduced or attributed to LLMs generally rather than to one specific model/prompt combination.
- [§4.1–4.2 / Figures 1–4] Figures 1-4 do not identify which LLM produced the data, include no error bars, and are not accompanied by significance tests. For example, the claim that LLMs 'detect and respond to shifts within only a few rounds' is not quantified with a distribution, confidence interval, or comparison against a null model. The figures also use legends referencing specific classical strategies but do not state which AI model is being plotted. Without this information, the adaptation-speed and performance claims are not verifiable.
- [Checklist] The NeurIPS checklist contains false or incomplete statements. Item 2 claims a Limitations section (Section 7), but no Section 7 exists in the manuscript. Items 14 and 15 answer 'NA' for human subjects, despite Section 4.3 describing recruitment of 10 human participants. Item 5 contains '[TODO]' answers. These issues make the submission incomplete and are inconsistent with the paper's actual content.
- [§1 vs Abstract] The introduction and the abstract contradict each other on the central adaptability claim. Section 1 states that LLM-based players 'were able to recognize the switch and adapt slower than human players, with 24.5 rounds early change in strategy compared to humans,' whereas the abstract claims LLMs rival or surpass human adaptability. Table 1 indicates faster AI adaptation (3.7 vs 5.4 rounds). These statements cannot all be correct and must be reconciled.
minor comments (4)
- [Throughout] The manuscript contains numerous typographical and grammatical errors, e.g., 'the demonstrated through a tournament', 'F orgiven defection', and 'Aspects of the strategies change their actions ... associated Specific properties.' A thorough copyedit is needed.
- [Figure 2] The legend includes names such as 'First by Grofman' but the axes and figure caption do not specify which AI model is being evaluated or how cooperation rate is aggregated across seeds. Add clear axis labels and model identification.
- [References] Reference [5] is cited as the source of the 240 classical strategies, but the listed paper is about GPT in game theory experiments, not the Axelrod strategy library. Verify and cite the actual strategy repository.
- [§4.3] The human experiment reports only N=10 and lacks the full participant instructions, compensation details, and IRB information. The checklist's 'NA' for human subjects is inconsistent and should be corrected.
Circularity Check
No circular reasoning: the paper's empirical claims, while severely under-supported and internally inconsistent, are not generated by a self-referential derivation or by fitting parameters to the target outcomes.
full rationale
The paper is an empirical study with no formal derivation chain; its claims are inferences from tournament runs and human experiments, not conclusions derived from equations whose inputs already contain the outputs. None of the enumerated circularity patterns is present. The behavioral metrics of Section 2.3 (niceness, forgivingness, retaliation) are defined independently of any LLM output and then measured; they are not used to define the performance or adaptability claims. The only selection step is the pilot choice of a prompt variant in Section 3 ('we tested multiple prompt variants in a pilot study to verify stability, and the most consistent variant is selected'), which affects external validity but does not by construction force the tournament ranking, cooperation rates, or adaptation speeds. There are no self-citations that carry a load-bearing argument; references [5], [13], and [15] are external prior work, and no uniqueness theorem is imported from the authors. No ansatz is smuggled in via citation. The serious problems in the manuscript—absence of a Results section, the contradiction between Table 1 and the Section 4.3 text ('humans adapted more slowly ... but maintained higher long-term cooperation rates after the switch' while the table shows AI with higher post-switch cooperation), unspecified LLM identity, truncated prompt, and false checklist claims of a Limitations section and of no human subjects—are correctness, reproducibility, and integrity issues, not circular reasoning. Therefore the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Prompt-based LLM play is a valid representation of an IPD strategy.
- domain assumption The 240 classical strategies form the standard benchmark suite.
- domain assumption Human adaptability measured with 10 online participants is representative.
Cite this review
Pith. "Pith review of Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory." pith.science (2026). https://pith.science/paper/SWVY523L
@misc{pith2026250904847,
author = {Pith},
title = {Pith review of: Collaboration and Conflict between Humans and Language Models through the Lens of Game Theory},
year = {2026},
howpublished = {\url{https://pith.science/paper/SWVY523L}},
note = {Machine review of arXiv:2509.04847}
}
read the original abstract
Language models are increasingly deployed in interactive online environments, from personal chat assistants to domain-specific agents, raising questions about their cooperative and competitive behavior in multi-party settings. While prior work has examined language model decision-making in isolated or short-term game-theoretic contexts, these studies often neglect long-horizon interactions, human-model collaboration, and the evolution of behavioral patterns over time. In this paper, we investigate the dynamics of language model behavior in the iterated prisoner's dilemma (IPD), a classical framework for studying cooperation and conflict. We pit model-based agents against a suite of 240 well-established classical strategies in an Axelrod-style tournament and find that language models achieve performance on par with, and in some cases exceeding, the best-known classical strategies. Behavioral analysis reveals that language models exhibit key properties associated with strong cooperative strategies - niceness, provocability, and generosity while also demonstrating rapid adaptability to changes in opponent strategy mid-game. In controlled "strategy switch" experiments, language models detect and respond to shifts within only a few rounds, rivaling or surpassing human adaptability. These results provide the first systematic characterization of long-term cooperative behaviors in language model agents, offering a foundation for future research into their role in more complex, mixed human-AI social environments.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin. Large language models for mathemat- ical reasoning: Progresses and challenges, 2024
work page 2024
- [2]
-
[3]
cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents
P. M. P. Curvo, M. Dragomir, S. Torpes, and M. Rahimi. Reproducibility study of "cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents", 2025
work page 2025
-
[4]
U. Faigle. Mathematical game theory: A new approach, 2023
work page 2023
-
[5]
F. Guo. Gpt in game theory experiments, 2023
work page 2023
-
[6]
S. Kambhampati, K. Stechly, K. Valmeekam, L. Saldyt, S. Bhambri, V . Palod, A. Gundawar, S. R. Samineni, D. Kalwar, and U. Biswas. Stop anthropomorphizing intermediate tokens as reasoning/thinking traces!, 2025
work page 2025
-
[7]
H. Le, Y . Wang, A. D. Gotmare, S. Savarese, and S. C. H. Hoi. Coderl: Mastering code generation through pretrained models and deep reinforcement learning.Advances in Neural Information Processing Systems, 35:21314–21328, 2022
work page 2022
-
[8]
S. M. Mousavi, E. Cecchinato, L. Hornikova, and G. Riccardi. Garbage in, reasoning out? why benchmark scores are unreliable and what to do about it, 2025
work page 2025
Show all 16 references
-
[9]
Oberauer
K. Oberauer. Working memory and attention – a conceptual analysis and review.Journal of Cognition, 2(1):36, 2019
2019
-
[10]
Jaech, A
OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kon- drich, A...
2024
-
[11]
Plaat, M
A. Plaat, M. van Duijn, N. van Stein, M. Preuss, P. van der Putten, and K. J. Batenburg. Agentic large language models, a survey.arXiv preprint arXiv:2503.23037, 2025
2025
-
[12]
Roziere, J
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y . Adi, J. Liu, R. Sauvestre, T. Remez, et al. Code llama: Open foundation models for code, 2023
2023
-
[13]
Singer-Clark
T. Singer-Clark. Morality metrics on iterated prisoners dilemma players. 2014
2014
-
[14]
Wilson and J
T. Wilson and J. Schooler. Thinking too much: Introspection can reduce the quality of preferences and decisions.Journal of personality and social psychology, 60:181–92, 03 1991
1991
-
[15]
[Yes] " is generally preferable to
Q. Zhu. Game theory meets llm and agentic ai: Reimagining cybersecurity for the age of intelligent threats, 2025. 8 A Morality Metrics Strategy Coop. Rate Good Partner Forgiveness Retaliation Generosity Always Cooperate 100.0 100.0 0.0 0.0 0.0 Tit For Tat 78.3 67.1 0.0 100.0 0...
2025
-
[16]
Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Institutional review board (IRB) approvals or equivalent for research with human subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals...
2025
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.