REVIEW 5 major objections 6 minor 37 references
Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"
T0 review · 5 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A low-budget replication of the GovSim cooperation benchmark reproduces the original pass/fail results and extends them to DeepSeek-V3, a harmful-resource framing, and mixed-model teams.
desk verdict A useful, honest reproduction study with real new results, but its confirmation rests on three runs at one seed and the MultiGov claim runs ahead of the data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is GovSim, a five-agent simulation of a shared resource over twelve monthly rounds in which each agent harvests privately, sees all harvests, then talks freely before the resource grows by a factor of two up to a cap of 100, collapsing if it falls below $C=5$; the sustainability threshold $f(t)$ is the maximum harvest that leaves the resource able to recover. The two treatment levers are the default instructions and the universalization prompt, which tells agents to ask what happens if everyone takes more than the sustainable harvest. The argument runs by comparing survival time, total gain, efficiency, equality, and over-usage statistics, and the extensions swap in new models, a Japanese translation, a negative resource, and mixed-model agent rosters.
What would settle it
Rerun the GovSim Fishery default scenario for GPT-4o-mini, Qwen2.5-7B, and DeepSeek-V3 across ten seeds and two temperatures; if a model that passed at seed 42 fails in most other seeds, or a classified-as-failing model survives twelve months in most runs, then the reproduced pass/fail alignment does not establish the models' true behavior.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the original GovSim findings reproduce: in the Fishery scenario with five homogeneous agents, GPT-4-turbo and GPT-4o survive all 12 months while GPT-3.5, Mistral-7B, and the Llama and Qwen models collapse in the first one or two months, and the universalization principle lifts survival times substantially for several small models. Among the extensions, DeepSeek-V3 achieves the same perfect sustainability metrics as GPT-4-turbo, GPT-4o-mini jumps from one month to twelve months under universalization, and a mathematically equivalent trash scenario flips the usual ranking because almost every model keeps the shared environment livable for 12 months. In heterogeneous teams, one high-performing agent among four low-performing agents is not enough to prevent collapse, but four high-performing agents can persuade a lone low-performing agent to cut its harvest after the first discussion round. The Japanese-translated prompts produced no meaningful behavioral shift, which the paper reads as evidence that language alone does not trigger collectivist cooperation.
Load-bearing premise
The results rest on the assumption that three runs with one fixed random seed capture how each model actually behaves, so the pass/fail agreement with the original study is not a coincidence of that particular seed.
Editorial extensions
If this is right
- If the reproduction claim is right, the GovSim pass/fail classification of LLMs is stable enough to serve as a benchmark for future models without rerunning the full original battery.
- The universalization result implies that a single explicit reasoning prompt can convert several sub-12-month models into full-term cooperators, which makes the principle a practical intervention rather than an evaluation detail.
- The trash-scenario flip implies that mathematically identical resource problems are not behaviorally identical: agents treat removal of a harmful resource as a different game, so prompt framing is part of the benchmark.
- The heterogeneous-team result implies that mixed-model deployments can keep cooperative performance while using fewer expensive large agents, as long as enough high-performing agents are present.
- The Japanese result implies that translating instructions into a collectivist-associated language does not by itself change harvesting behavior, at least for the models and neutral narrative tested here.
Reading between the lines
- I would not extend the pass/fail alignment beyond the tested seed: with three runs at seed 42, the confirmation of the original claims could be a lucky alignment if harvest decisions vary strongly across seeds or temperatures.
- The trash-scenario results suggest a testable framing hypothesis: if loss aversion is the mechanism, the size of the survival-time reversal should increase with the severity of the negative framing, which the paper does not vary.
- The influence result invites a probe of mechanism: recording which argument types, such as numeric limits, appeals to fairness, or threats, actually change a low-performing agent's harvest could turn the observation into a design principle for multi-agent coordination.
- The neutral Japanese translation may understate language effects; a culturally embedded version referencing local fishing practices could elicit the collectivist shift the authors hypothesized but did not find.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a reproducibility study of Piatti et al.'s GovSim framework, focusing on the Fishery scenario in both default and universalization settings. The authors reproduce the original pass/fail outcomes for eight previously tested models, evaluate four new models (DeepSeek-V3, GPT-4o-mini, Qwen2.5-0.5B, Qwen2.5-7B), and introduce three extensions: Japanese-language prompts, an inverse "trash" scenario, and heterogeneous multi-agent teams (MultiGov). The paper concludes that the original claims are confirmed, that DeepSeek-V3 behaves similarly to GPT-4-turbo, that most models pass the sustainability test when the resource is framed as harmful trash, and that high-performing agents can steer low-performing agents toward cooperative behavior.
Significance. If the reported results hold, the paper provides a useful independent check of the original GovSim findings and extends the benchmark to new models, a new language, and mixed-model teams. The authors deserve credit for making their code available, reporting computational costs and energy consumption transparently, and framing the extensions as falsifiable hypotheses. The reproduction tables are largely consistent with the original pass/fail outcomes, and the absence of fitted parameters or circular derivations is a strength. However, the thin evidence base (three runs at a fixed seed, no significance tests, and internally inconsistent run-count reporting) means that the central confirmation claim is not yet established at the level the conclusions assert.
major comments (5)
- [§3.2 and Table 2 caption] The manuscript is internally inconsistent about the number of runs: Section 3.2 states that "three runs" were conducted for each reproduction configuration, Table 2's caption attributes metric differences to a "single-run approach," and Table 11's footnote says "Each experiment was run 3 times." This ambiguity is load-bearing because the error bars in Tables 1 and 2, and the claim that results fall "within the original error margin," depend on knowing the actual sample size. Please state the exact run count per configuration and recompute or relabel the statistics accordingly.
- [§4.1, Tables 1–2] The confirmation of Claims 1 and 2 rests on pass/fail alignment with the original paper, but the evidence is three runs (or possibly one, per the caption) at a fixed seed of 42, with no significance tests or seed variation. Some models sit near the sustainability threshold, notably Qwen2.5-7B (survival time 7.7±5.9, survival rate 0.3) and Mistral-7B (survival time 6.7±1.5) in the universalization scenario. For such models, a different seed could flip the pass/fail outcome, which would change the qualitative conclusions in Section 5. Please provide multi-seed results (or a clear statistical justification that pass/fail is insensitive to seed), or substantially soften the statement that the reproduction "confirmed the claims of the original study."
- [§4.2, Table 6, MultiGov] The claim that high-performing agents can guide low-performing agents to sustainable cooperation is only partially supported by Table 6. The 4×GPT-4o-Turbo + 1×GPT-4o-mini configuration has survival rate 0.0 and mean survival time 3.7±0.6, while the 2×DeepSeek-V3 + 3×GPT-4o-mini and 1×DeepSeek-V3 + 4×GPT-4o-mini configurations also fail. Only the 4×DeepSeek-V3 + 1×GPT-4o-mini configuration achieves full survival. The text acknowledges behavioral shifts but the conclusion that influence "often led to the survival of the group" overstates the quantitative results. Please reconcile the narrative with the actual survival rates and discuss the failed configurations explicitly.
- [§4.2, Table 5, inverse (trash) scenario] The text states "Except for Mistral-7B and Qwen2.5-0.5B, all models maintained cooperation for the full 12 months," but Table 5 reports Llama-2-7B with survival rate 0.3 and survival time 11.0±1.0, and Mistral-7B with survival rate 0.0 and survival time 8.0±5.2. The table and prose are therefore in direct conflict. This matters because the "striking contrast" of nearly universal success in the trash scenario is used to support the loss-aversion interpretation. Please correct the inaccuracy and quantify how many runs passed for each model.
- [§4.2, Table 4, Japanese translation] The claim that "We have found no significant differences in the models' behavior when instructed in Japanese" is not supported by any statistical test, and Table 4 reports single values with no error bars. For GPT-4o, survival time decreased from 12 months in English to 11 months in Japanese; for DeepSeek-V3, total gain dropped from 119.4 to 85.8. Without variance estimates or a significance test, the claim of no difference is unjustified. Please provide per-run results or rephrase the conclusion as an absence of large qualitative differences rather than statistical insignificance.
minor comments (6)
- [Throughout] Model naming is inconsistent: the text and some figures use "GPT-4-turbo," "GPT-4o-Turbo," and "GPT-4 Turbo" interchangeably. Please standardize the model names across tables, figures, and text.
- [Table 1] Several entries are concatenated without spacing, for example the GPT-4o row shows "71.3±0.6 59.4±0.5 98.5±0.60.0±0.0." This makes the table difficult to read and should be reformatted.
- [Table 4] Table 4 omits the Survival Rate column that appears in the other metric tables, which complicates direct comparison with Tables 1, 2, and 5. Please add the column for consistency.
- [Figure 2] The captions for panels (c), (d), and (e) are identical even though the plots appear to show different runs or outcomes. Please clarify what distinguishes these panels.
- [Appendix F, Table 8] The fixed-parameter table has formatting issues: the "Observation Strategy" row appears to contain stray text and the table layout is broken. Please fix the table so each parameter and value is clearly separated.
- [References] Several references have malformed author names (e.g., "Mert et al. Cemri", "Qwen and et al.") and the Hardin 1968a/1968b entries are duplicated. Please clean up the reference list.
Circularity Check
No circularity: this is a benchmark reproduction compared against externally published results, with no fitted parameters or self-citation chain.
full rationale
This paper is a reproducibility study rather than a derivation. Its central validation step compares fresh simulation runs against pass/fail outcomes previously published by Piatti et al., an external group; no parameter is fitted to the comparison target, and no quantity is predicted from the very values it claims to confirm. The pass/fail alignment in Tables 1 and 2 is an empirical comparison, not an identity by construction. Reusing the original GovSim repository and configuration is standard benchmark reuse, not a self-citation chain, and citing the original paper as the source of the benchmark does not make the confirmation circular. The extensions (new models, Japanese prompts, trash scenario, MultiGov) are independently executed experiments whose conclusions rest on the new runs. The internal inconsistency between Section 3.2 stating three runs per setup and the Table 2 caption describing a single-run approach, together with the fixed seed of 42, is a methodological robustness concern, not a circularity concern, because no fitted input is renamed as a prediction. The paper is self-contained against an external benchmark and therefore warrants a score of 0.
Assumptions & free parameters
assumptions (5)
- domain assumption GovSim environment and metrics from Piatti et al. are an accepted benchmark for LLM cooperation.
- domain assumption The trash scenario is mathematically equivalent to the fishery scenario.
- domain assumption Three simulation runs (or one run, per Table 2 caption) at seed 42 are enough to classify pass/fail.
- domain assumption A DeepL translation partly reviewed by one Japanese speaker is a valid operationalization of Japanese instructions.
- domain assumption Observed reductions in a weak agent's harvest after discussion are caused by the strong agent's communication.
Cite this review
Pith. "Pith review of Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"." pith.science (2026). https://pith.science/paper/6WLWAHGC
@misc{pith2026250509289,
author = {Pith},
title = {Pith review of: Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"},
year = {2026},
howpublished = {\url{https://pith.science/paper/6WLWAHGC}},
note = {Machine review of arXiv:2505.09289}
}
read the original abstract
This study evaluates and extends the findings made by Piatti et al., who introduced GovSim, a simulation framework designed to assess the cooperative decision-making capabilities of large language models (LLMs) in resource-sharing scenarios. By replicating key experiments, we validate claims regarding the performance of large models, such as GPT-4-turbo, compared to smaller models. The impact of the universalization principle is also examined, with results showing that large models can achieve sustainable cooperation, with or without the principle, while smaller models fail without it. In addition, we provide multiple extensions to explore the applicability of the framework to new settings. We evaluate additional models, such as DeepSeek-V3 and GPT-4o-mini, to test whether cooperative behavior generalizes across different architectures and model sizes. Furthermore, we introduce new settings: we create a heterogeneous multi-agent environment, study a scenario using Japanese instructions, and explore an "inverse environment" where agents must cooperate to mitigate harmful resource distributions. Our results confirm that the benchmark can be applied to new models, scenarios, and languages, offering valuable insights into the adaptability of LLMs in complex cooperative tasks. Moreover, the experiment involving heterogeneous multi-agent systems demonstrates that high-performing models can influence lower-performing ones to adopt similar behaviors. This finding has significant implications for other agent-based applications, potentially enabling more efficient use of computational resources and contributing to the development of more effective cooperative AI systems.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
“Japanese Collectivism”, pp.\ 1–16. Culture and Psychology. Cambridge University Press, Cambridge, 2024. ISBN 978-1-108-83320-2. doi:10.1017/9781108973625.002. URL https://www.cambridge.org/core/books/cultural-stereotype-and-its-hazards/japanese-collectivism/619E4CD39EBC4A2896DA4A0B11A3E105
-
[2]
Loss aversion under prospect theory: A parameter-free measurement
Mohammed Abdellaoui, Han Bleichrodt, and Corina Paraschiv. Loss aversion under prospect theory: A parameter-free measurement. Management Science, 53 0 (10): 0 1659–1674, October 2007. ISSN 0025-1909. doi:10.1287/mnsc.1070.0711
-
[3]
The claude 3 model family: Opus, sonnet, haiku
Anthropic. The claude 3 model family: Opus, sonnet, haiku. 2024. URL https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf
work page 2024
-
[4]
Association of Issuing Bodies . European residual mixes 2023. 2023. URL https://www.aib-net.org/sites/default/files/assets/facts/residual-mix/2023/AIB_2023_Residual_Mix_FINALResults09072024.pdf
work page 2023
-
[5]
Mert et al. Cemri. Why do multi-agent llm systems fail? arXiv preprint arXiv:2503.13657, 2025. doi:10.48550/arXiv.2503.13657. URL http://arxiv.org/abs/2503.13657
-
[6]
Zhibo et al. Chu. Fairness in large language models: A taxonomic survey. arXiv preprint arXiv:2404.01349, 2024. doi:10.48550/arXiv.2404.01349. URL http://arxiv.org/abs/2404.01349
-
[7]
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. 0 (arXiv:2401.06066), January 2024. doi:10.48550/arXiv.24...
- [8]
Show all 37 references
- [9]
-
[10]
Analyzing cultural representations of emotions in llms through mixed emotion survey
Shiran Dudy, Ibrahim Said Ahmad, Ryoko Kitajima, and Agata Lapedriza. Analyzing cultural representations of emotions in llms through mixed emotion survey. 0 (arXiv:2408.02143), August 2024. doi:10.48550/arXiv.2408.02143. URL http://arxiv.org/abs/2408.02143. arXiv:2408.02143 [cs]
-
[11]
Large language models empowered agent-based modeling and simulation: A survey and perspectives
Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Ding, Zhilun Zhou, Fengli Xu, and Yong Li. Large language models empowered agent-based modeling and simulation: A survey and perspectives. 0 (arXiv:2312.11970), December 2023. doi:10.48550/arXiv.2312.11970. URL http://arxiv....
-
[12]
Parameter-efficient mixture-of-experts architecture for pre-trained language models
Ze-Feng Gao, Peiyu Liu, Wayne Xin Zhao, Zhong-Yi Lu, and Ji-Rong Wen. Parameter-efficient mixture-of-experts architecture for pre-trained language models. 0 (arXiv:2203.01104), October 2022. doi:10.48550/arXiv.2203.01104. URL http://arxiv.org/abs/2203.01104. arXiv:2203.01104 [cs]
-
[13]
Scott Gordon
H. Scott Gordon. The economic theory of a common-property resource: The fishery. Journal of Political Economy, 62 0 (2): 0 124--142, 1954. doi:10.1086/257497. URL https://www.journals.uchicago.edu/doi/epdf/10.1086/257497
1954 doi
- [14]
-
[16]
The tragedy of the commons: The population problem has no technical solution; it requires a fundamental extension in morality
Garrett Hardin. The tragedy of the commons: The population problem has no technical solution; it requires a fundamental extension in morality. Science, 162 0 (3859): 0 1243–1248, December 1968 b . ISSN 0036-8075, 1095-9203. doi:10.1126/science.162.3859.1243
1968
-
[17]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. 0 (arXiv:2009.03300), January 2021. doi:10.48550/arXiv.2009.03300. URL http://arxiv.org/abs/2009.03300. arXiv:2009.03300 [cs]
- [18]
-
[19]
Build an influential bot in social media simulations with large language models
Bailu Jin and Weisi Guo. Build an influential bot in social media simulations with large language models. arXiv preprint arXiv:2411.19635, 2024. doi:10.48550/arXiv.2411.19635. URL http://arxiv.org/abs/2411.19635
- [20]
- [21]
-
[22]
Kochenderfer, Christopher Amato, Girish Chowdhary, Jonathan P
Mykel J. Kochenderfer, Christopher Amato, Girish Chowdhary, Jonathan P. How, Hayley J. Davison Reynolds, Jason R. Thornton, Pedro A. Torres-Carrasquillo, N. Kemal Üre, and John Vian. Decision Making Under Uncertainty: Theory and Application. The MIT Press, July 2015. ISBN 978-...
2015 doi
-
[23]
Charles D. Kolstad. Environmental economics. New York: Oxford University Press, 2011. ISBN 978-0-19-973264-7. URL http://archive.org/details/environmentaleco0000kols_i2l8
2011
-
[24]
Comparing biases and the impact of multilingual training across multiple languages
Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. Comparing biases and the impact of multilingual training across multiple languages. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceeding...
2023 doi
- [25]
-
[26]
Gpt-4o-mini: Advancing cost-efficient intelligence, 2024
OpenAI. Gpt-4o-mini: Advancing cost-efficient intelligence, 2024. URL https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/
2024
- [27]
-
[28]
Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents
Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Sch \"o lkopf, Mrinmaya Sachan, and Rada Mihalcea. Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
- [29]
-
[30]
What is loss aversion? Journal of Risk and Uncertainty, 30 0 (2): 0 157–167, March 2005
Ulrich Schmidt and Horst Zank. What is loss aversion? Journal of Risk and Uncertainty, 30 0 (2): 0 157–167, March 2005. ISSN 1573-0476. doi:10.1007/s11166-005-6564-6
2005 doi
- [31]
-
[32]
Greenhouse gas equivalencies calculator - calculations and references, August 2015
OAR US EPA. Greenhouse gas equivalencies calculator - calculations and references, August 2015. URL https://www.epa.gov/energy/greenhouse-gas-equivalencies-calculator-calculations-and-references
2015
- [33]
-
[34]
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks...
- [35]
- [36]
-
[37]
Cultural value differences of llms: Prompt, language, and model size
Qishuai Zhong, Yike Yun, and Aixin Sun. Cultural value differences of llms: Prompt, language, and model size. 0 (arXiv:2407.16891), June 2024. doi:10.48550/arXiv.2407.16891. URL http://arxiv.org/abs/2407.16891. arXiv:2407.16891 [cs]
- [38]
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.