Pith. sign in

REVIEW 5 major objections 6 minor 37 references

Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"

T0 review · 5 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A low-budget replication of the GovSim cooperation benchmark reproduces the original pass/fail results and extends them to DeepSeek-V3, a harmful-resource framing, and mixed-model teams.

desk verdict A useful, honest reproduction study with real new results, but its confirmation rests on three runs at one seed and the MultiGov claim runs ahead of the data. read the letter →

arxiv 2505.09289 v1 pith:6WLWAHGC submitted 2025-05-14 cs.AI

classification cs.AI
keywords reproducibilitystudyGovSimLLMagentssustainablecooperationtragedyofthecommonsuniversalizationprinciplemulti-agentsystemslossaversion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This reproducibility study tries to establish that the cooperative-behavior benchmark from the original GovSim work is stable enough to build on: the same models pass or fail the 12-month fishery sustainability test as in the original study, and the universalization prompt rescues several small models. It then argues the benchmark transfers to new settings, finding that DeepSeek-V3 matches GPT-4-turbo, that most models pass when the shared resource is framed as harmful trash that must be removed, and that four high-performing agents can talk a lone low-performing agent into sustainable harvesting. A sympathetic reader would take away that LLM cooperation is reproducible, scale-dependent, and sensitive to how the shared resource is framed.

What carries the argument

The carrying object is GovSim, a five-agent simulation of a shared resource over twelve monthly rounds in which each agent harvests privately, sees all harvests, then talks freely before the resource grows by a factor of two up to a cap of 100, collapsing if it falls below $C=5$; the sustainability threshold $f(t)$ is the maximum harvest that leaves the resource able to recover. The two treatment levers are the default instructions and the universalization prompt, which tells agents to ask what happens if everyone takes more than the sustainable harvest. The argument runs by comparing survival time, total gain, efficiency, equality, and over-usage statistics, and the extensions swap in new models, a Japanese translation, a negative resource, and mixed-model agent rosters.

What would settle it

Rerun the GovSim Fishery default scenario for GPT-4o-mini, Qwen2.5-7B, and DeepSeek-V3 across ten seeds and two temperatures; if a model that passed at seed 42 fails in most other seeds, or a classified-as-failing model survives twelve months in most runs, then the reproduced pass/fail alignment does not establish the models' true behavior.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that the original GovSim findings reproduce: in the Fishery scenario with five homogeneous agents, GPT-4-turbo and GPT-4o survive all 12 months while GPT-3.5, Mistral-7B, and the Llama and Qwen models collapse in the first one or two months, and the universalization principle lifts survival times substantially for several small models. Among the extensions, DeepSeek-V3 achieves the same perfect sustainability metrics as GPT-4-turbo, GPT-4o-mini jumps from one month to twelve months under universalization, and a mathematically equivalent trash scenario flips the usual ranking because almost every model keeps the shared environment livable for 12 months. In heterogeneous teams, one high-performing agent among four low-performing agents is not enough to prevent collapse, but four high-performing agents can persuade a lone low-performing agent to cut its harvest after the first discussion round. The Japanese-translated prompts produced no meaningful behavioral shift, which the paper reads as evidence that language alone does not trigger collectivist cooperation.

Load-bearing premise

The results rest on the assumption that three runs with one fixed random seed capture how each model actually behaves, so the pass/fail agreement with the original study is not a coincidence of that particular seed.

Editorial extensions

If this is right

  • If the reproduction claim is right, the GovSim pass/fail classification of LLMs is stable enough to serve as a benchmark for future models without rerunning the full original battery.
  • The universalization result implies that a single explicit reasoning prompt can convert several sub-12-month models into full-term cooperators, which makes the principle a practical intervention rather than an evaluation detail.
  • The trash-scenario flip implies that mathematically identical resource problems are not behaviorally identical: agents treat removal of a harmful resource as a different game, so prompt framing is part of the benchmark.
  • The heterogeneous-team result implies that mixed-model deployments can keep cooperative performance while using fewer expensive large agents, as long as enough high-performing agents are present.
  • The Japanese result implies that translating instructions into a collectivist-associated language does not by itself change harvesting behavior, at least for the models and neutral narrative tested here.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I would not extend the pass/fail alignment beyond the tested seed: with three runs at seed 42, the confirmation of the original claims could be a lucky alignment if harvest decisions vary strongly across seeds or temperatures.
  • The trash-scenario results suggest a testable framing hypothesis: if loss aversion is the mechanism, the size of the survival-time reversal should increase with the severity of the negative framing, which the paper does not vary.
  • The influence result invites a probe of mechanism: recording which argument types, such as numeric limits, appeals to fairness, or threats, actually change a low-performing agent's harvest could turn the observation into a design principle for multi-agent coordination.
  • The neutral Japanese translation may understate language effects; a culturally embedded version referencing local fishing practices could elicit the collectivist shift the authors hypothesized but did not find.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript reports a reproducibility study of Piatti et al.'s GovSim framework, focusing on the Fishery scenario in both default and universalization settings. The authors reproduce the original pass/fail outcomes for eight previously tested models, evaluate four new models (DeepSeek-V3, GPT-4o-mini, Qwen2.5-0.5B, Qwen2.5-7B), and introduce three extensions: Japanese-language prompts, an inverse "trash" scenario, and heterogeneous multi-agent teams (MultiGov). The paper concludes that the original claims are confirmed, that DeepSeek-V3 behaves similarly to GPT-4-turbo, that most models pass the sustainability test when the resource is framed as harmful trash, and that high-performing agents can steer low-performing agents toward cooperative behavior.

Significance. If the reported results hold, the paper provides a useful independent check of the original GovSim findings and extends the benchmark to new models, a new language, and mixed-model teams. The authors deserve credit for making their code available, reporting computational costs and energy consumption transparently, and framing the extensions as falsifiable hypotheses. The reproduction tables are largely consistent with the original pass/fail outcomes, and the absence of fitted parameters or circular derivations is a strength. However, the thin evidence base (three runs at a fixed seed, no significance tests, and internally inconsistent run-count reporting) means that the central confirmation claim is not yet established at the level the conclusions assert.

major comments (5)
  1. [§3.2 and Table 2 caption] The manuscript is internally inconsistent about the number of runs: Section 3.2 states that "three runs" were conducted for each reproduction configuration, Table 2's caption attributes metric differences to a "single-run approach," and Table 11's footnote says "Each experiment was run 3 times." This ambiguity is load-bearing because the error bars in Tables 1 and 2, and the claim that results fall "within the original error margin," depend on knowing the actual sample size. Please state the exact run count per configuration and recompute or relabel the statistics accordingly.
  2. [§4.1, Tables 1–2] The confirmation of Claims 1 and 2 rests on pass/fail alignment with the original paper, but the evidence is three runs (or possibly one, per the caption) at a fixed seed of 42, with no significance tests or seed variation. Some models sit near the sustainability threshold, notably Qwen2.5-7B (survival time 7.7±5.9, survival rate 0.3) and Mistral-7B (survival time 6.7±1.5) in the universalization scenario. For such models, a different seed could flip the pass/fail outcome, which would change the qualitative conclusions in Section 5. Please provide multi-seed results (or a clear statistical justification that pass/fail is insensitive to seed), or substantially soften the statement that the reproduction "confirmed the claims of the original study."
  3. [§4.2, Table 6, MultiGov] The claim that high-performing agents can guide low-performing agents to sustainable cooperation is only partially supported by Table 6. The 4×GPT-4o-Turbo + 1×GPT-4o-mini configuration has survival rate 0.0 and mean survival time 3.7±0.6, while the 2×DeepSeek-V3 + 3×GPT-4o-mini and 1×DeepSeek-V3 + 4×GPT-4o-mini configurations also fail. Only the 4×DeepSeek-V3 + 1×GPT-4o-mini configuration achieves full survival. The text acknowledges behavioral shifts but the conclusion that influence "often led to the survival of the group" overstates the quantitative results. Please reconcile the narrative with the actual survival rates and discuss the failed configurations explicitly.
  4. [§4.2, Table 5, inverse (trash) scenario] The text states "Except for Mistral-7B and Qwen2.5-0.5B, all models maintained cooperation for the full 12 months," but Table 5 reports Llama-2-7B with survival rate 0.3 and survival time 11.0±1.0, and Mistral-7B with survival rate 0.0 and survival time 8.0±5.2. The table and prose are therefore in direct conflict. This matters because the "striking contrast" of nearly universal success in the trash scenario is used to support the loss-aversion interpretation. Please correct the inaccuracy and quantify how many runs passed for each model.
  5. [§4.2, Table 4, Japanese translation] The claim that "We have found no significant differences in the models' behavior when instructed in Japanese" is not supported by any statistical test, and Table 4 reports single values with no error bars. For GPT-4o, survival time decreased from 12 months in English to 11 months in Japanese; for DeepSeek-V3, total gain dropped from 119.4 to 85.8. Without variance estimates or a significance test, the claim of no difference is unjustified. Please provide per-run results or rephrase the conclusion as an absence of large qualitative differences rather than statistical insignificance.
minor comments (6)
  1. [Throughout] Model naming is inconsistent: the text and some figures use "GPT-4-turbo," "GPT-4o-Turbo," and "GPT-4 Turbo" interchangeably. Please standardize the model names across tables, figures, and text.
  2. [Table 1] Several entries are concatenated without spacing, for example the GPT-4o row shows "71.3±0.6 59.4±0.5 98.5±0.60.0±0.0." This makes the table difficult to read and should be reformatted.
  3. [Table 4] Table 4 omits the Survival Rate column that appears in the other metric tables, which complicates direct comparison with Tables 1, 2, and 5. Please add the column for consistency.
  4. [Figure 2] The captions for panels (c), (d), and (e) are identical even though the plots appear to show different runs or outcomes. Please clarify what distinguishes these panels.
  5. [Appendix F, Table 8] The fixed-parameter table has formatting issues: the "Observation Strategy" row appears to contain stray text and the table layout is broken. Please fix the table so each parameter and value is clearly separated.
  6. [References] Several references have malformed author names (e.g., "Mert et al. Cemri", "Qwen and et al.") and the Hardin 1968a/1968b entries are duplicated. Please clean up the reference list.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: this is a benchmark reproduction compared against externally published results, with no fitted parameters or self-citation chain.

full rationale

This paper is a reproducibility study rather than a derivation. Its central validation step compares fresh simulation runs against pass/fail outcomes previously published by Piatti et al., an external group; no parameter is fitted to the comparison target, and no quantity is predicted from the very values it claims to confirm. The pass/fail alignment in Tables 1 and 2 is an empirical comparison, not an identity by construction. Reusing the original GovSim repository and configuration is standard benchmark reuse, not a self-citation chain, and citing the original paper as the source of the benchmark does not make the confirmation circular. The extensions (new models, Japanese prompts, trash scenario, MultiGov) are independently executed experiments whose conclusions rest on the new runs. The internal inconsistency between Section 3.2 stating three runs per setup and the Table 2 caption describing a single-run approach, together with the fixed seed of 42, is a methodological robustness concern, not a circularity concern, because no fitted input is renamed as a prediction. The paper is self-contained against an external benchmark and therefore warrants a score of 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new mathematical or physical entities and fits no parameters; its claims rest on the original GovSim implementation, several experimental-design assumptions, and unverified statistical sufficiency. Those assumptions are listed above.

assumptions (5)
  • domain assumption GovSim environment and metrics from Piatti et al. are an accepted benchmark for LLM cooperation.
    The paper adopts the original framework without re-deriving its dynamics; all conclusions are relative to this benchmark.
  • domain assumption The trash scenario is mathematically equivalent to the fishery scenario.
    Section 4.2 asserts equivalence and uses it to attribute outcome differences to framing and loss aversion; the equivalence is stated, not shown.
  • domain assumption Three simulation runs (or one run, per Table 2 caption) at seed 42 are enough to classify pass/fail.
    Section 3.2 calls three runs sufficient for validation, while the Table 2 caption calls the results a single-run approach; no variance or power analysis is provided.
  • domain assumption A DeepL translation partly reviewed by one Japanese speaker is a valid operationalization of Japanese instructions.
    Section 3.2 admits the translation was not validated by a native speaker, which limits the cross-lingual conclusion.
  • domain assumption Observed reductions in a weak agent's harvest after discussion are caused by the strong agent's communication.
    Section 4.2 infers influence from transcripts and small-N results without a controlled baseline or statistical test.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"." pith.science (2026). https://pith.science/paper/6WLWAHGC

@misc{pith2026250509289,
  author       = {Pith},
  title        = {Pith review of: Reproducibility Study of "Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents"},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6WLWAHGC}},
  note         = {Machine review of arXiv:2505.09289}
}
read the original abstract

This study evaluates and extends the findings made by Piatti et al., who introduced GovSim, a simulation framework designed to assess the cooperative decision-making capabilities of large language models (LLMs) in resource-sharing scenarios. By replicating key experiments, we validate claims regarding the performance of large models, such as GPT-4-turbo, compared to smaller models. The impact of the universalization principle is also examined, with results showing that large models can achieve sustainable cooperation, with or without the principle, while smaller models fail without it. In addition, we provide multiple extensions to explore the applicability of the framework to new settings. We evaluate additional models, such as DeepSeek-V3 and GPT-4o-mini, to test whether cooperative behavior generalizes across different architectures and model sizes. Furthermore, we introduce new settings: we create a heterogeneous multi-agent environment, study a scenario using Japanese instructions, and explore an "inverse environment" where agents must cooperate to mitigate harmful resource distributions. Our results confirm that the benchmark can be applied to new models, scenarios, and languages, offering valuable insights into the adaptability of LLMs in complex cooperative tasks. Moreover, the experiment involving heterogeneous multi-agent systems demonstrates that high-performing models can influence lower-performing ones to adopt similar behaviors. This finding has significant implications for other agent-based applications, potentially enabling more efficient use of computational resources and contributing to the development of more effective cooperative AI systems.

Figures

Figures reproduced from arXiv: 2505.09289 by the authors.

Figure 1
Figure 1. Example of a conversation between two agents in the MultiGov scenario. John ( [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Sustainability test results for the multi-agent fishery scenario with multi-agent [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 3
Figure 3. Sustainability test results for the homogeneous-agent fishery [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Sustainability test results for the homogeneous-agent fishery [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Sustainability test results for the homogeneous-agent fishery [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: Sustainability test results for the homogeneous-agent trash [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Sustainability test results for homogeneous-agent trash scenario with [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: The sixth communication phase of one run of the inverse (trash) scenario with the [PITH_FULL_IMAGE:figures/full_fig_p022_8.png]
Figure 9
Figure 9. Figure 9: Prompts from the first communication in the 1- [PITH_FULL_IMAGE:figures/full_fig_p023_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 13 canonical work pages

  1. [1]

    Japanese Collectivism

    “Japanese Collectivism”, pp.\ 1–16. Culture and Psychology. Cambridge University Press, Cambridge, 2024. ISBN 978-1-108-83320-2. doi:10.1017/9781108973625.002. URL https://www.cambridge.org/core/books/cultural-stereotype-and-its-hazards/japanese-collectivism/619E4CD39EBC4A2896DA4A0B11A3E105

  2. [2]

    Loss aversion under prospect theory: A parameter-free measurement

    Mohammed Abdellaoui, Han Bleichrodt, and Corina Paraschiv. Loss aversion under prospect theory: A parameter-free measurement. Management Science, 53 0 (10): 0 1659–1674, October 2007. ISSN 0025-1909. doi:10.1287/mnsc.1070.0711

  3. [3]

    The claude 3 model family: Opus, sonnet, haiku

    Anthropic. The claude 3 model family: Opus, sonnet, haiku. 2024. URL https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf

  4. [4]

    European residual mixes 2023

    Association of Issuing Bodies . European residual mixes 2023. 2023. URL https://www.aib-net.org/sites/default/files/assets/facts/residual-mix/2023/AIB_2023_Residual_Mix_FINALResults09072024.pdf

  5. [5]

    Mert et al. Cemri. Why do multi-agent llm systems fail? arXiv preprint arXiv:2503.13657, 2025. doi:10.48550/arXiv.2503.13657. URL http://arxiv.org/abs/2503.13657

  6. [6]

    Zhibo et al. Chu. Fairness in large language models: A taxonomic survey. arXiv preprint arXiv:2404.01349, 2024. doi:10.48550/arXiv.2404.01349. URL http://arxiv.org/abs/2404.01349

  7. [7]

    Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models. 0 (arXiv:2401.06066), January 2024. doi:10.48550/arXiv.24...

  8. [8]

    URL https://www.deepl.com/translator

    DeepL. URL https://www.deepl.com/translator

Show all 37 references
  1. [9]

    Deepseek-v3 technical report

    DeepSeek-AI and et al. Deepseek-v3 technical report. 0 (arXiv:2412.19437), December 2024. doi:10.48550/arXiv.2412.19437. URL http://arxiv.org/abs/2412.19437. arXiv:2412.19437 [cs]

  2. [10]

    Analyzing cultural representations of emotions in llms through mixed emotion survey

    Shiran Dudy, Ibrahim Said Ahmad, Ryoko Kitajima, and Agata Lapedriza. Analyzing cultural representations of emotions in llms through mixed emotion survey. 0 (arXiv:2408.02143), August 2024. doi:10.48550/arXiv.2408.02143. URL http://arxiv.org/abs/2408.02143. arXiv:2408.02143 [cs]

  3. [11]

    Large language models empowered agent-based modeling and simulation: A survey and perspectives

    Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Ding, Zhilun Zhou, Fengli Xu, and Yong Li. Large language models empowered agent-based modeling and simulation: A survey and perspectives. 0 (arXiv:2312.11970), December 2023. doi:10.48550/arXiv.2312.11970. URL http://arxiv....

  4. [12]

    Parameter-efficient mixture-of-experts architecture for pre-trained language models

    Ze-Feng Gao, Peiyu Liu, Wayne Xin Zhao, Zhong-Yi Lu, and Ji-Rong Wen. Parameter-efficient mixture-of-experts architecture for pre-trained language models. 0 (arXiv:2203.01104), October 2022. doi:10.48550/arXiv.2203.01104. URL http://arxiv.org/abs/2203.01104. arXiv:2203.01104 [cs]

  5. [13]

    Scott Gordon

    H. Scott Gordon. The economic theory of a common-property resource: The fishery. Journal of Political Economy, 62 0 (2): 0 124--142, 1954. doi:10.1086/257497. URL https://www.journals.uchicago.edu/doi/epdf/10.1086/257497

  6. [14]

    Aaron Grattafiori and et. al. The llama 3 herd of models. 0 (arXiv:2407.21783), November 2024. doi:10.48550/arXiv.2407.21783. URL http://arxiv.org/abs/2407.21783. arXiv:2407.21783 [cs]

  7. [16]

    The tragedy of the commons: The population problem has no technical solution; it requires a fundamental extension in morality

    Garrett Hardin. The tragedy of the commons: The population problem has no technical solution; it requires a fundamental extension in morality. Science, 162 0 (3859): 0 1243–1248, December 1968 b . ISSN 0036-8075, 1095-9203. doi:10.1126/science.162.3859.1243

  8. [17]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. 0 (arXiv:2009.03300), January 2021. doi:10.48550/arXiv.2009.03300. URL http://arxiv.org/abs/2009.03300. arXiv:2009.03300 [cs]

  9. [18]

    Tiancheng et al. Hu. Generative language models exhibit social identity biases. arXiv preprint arXiv:2310.15819, 2024. doi:10.48550/arXiv.2310.15819. URL http://arxiv.org/abs/2310.15819

  10. [19]

    Build an influential bot in social media simulations with large language models

    Bailu Jin and Weisi Guo. Build an influential bot in social media simulations with large language models. arXiv preprint arXiv:2411.19635, 2024. doi:10.48550/arXiv.2411.19635. URL http://arxiv.org/abs/2411.19635

  11. [20]

    Ariba et al. Khan. Randomness, not representation: The unreliability of evaluating cultural alignment in llms. arXiv preprint arXiv:2503.08688, 2025. doi:10.48550/arXiv.2503.08688. URL http://arxiv.org/abs/2503.08688

  12. [21]

    Kharchenko

    Julia et al. Kharchenko. How well do llms represent values across cultures? arXiv preprint arXiv:2406.14805, 2024. doi:10.48550/arXiv.2406.14805. URL http://arxiv.org/abs/2406.14805

  13. [22]

    Kochenderfer, Christopher Amato, Girish Chowdhary, Jonathan P

    Mykel J. Kochenderfer, Christopher Amato, Girish Chowdhary, Jonathan P. How, Hayley J. Davison Reynolds, Jason R. Thornton, Pedro A. Torres-Carrasquillo, N. Kemal Üre, and John Vian. Decision Making Under Uncertainty: Theory and Application. The MIT Press, July 2015. ISBN 978-...

  14. [23]

    Charles D. Kolstad. Environmental economics. New York: Oxford University Press, 2011. ISBN 978-0-19-973264-7. URL http://archive.org/details/environmentaleco0000kols_i2l8

  15. [24]

    Comparing biases and the impact of multilingual training across multiple languages

    Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. Comparing biases and the impact of multilingual training across multiple languages. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Proceeding...

  16. [25]

    Luccioni

    Alexandra Sasha et al. Luccioni. Bridging the gap: Integrating ethics and environmental sustainability in ai. arXiv preprint arXiv:2504.00797, 2025. doi:10.48550/arXiv.2504.00797. URL http://arxiv.org/abs/2504.00797

  17. [26]

    Gpt-4o-mini: Advancing cost-efficient intelligence, 2024

    OpenAI. Gpt-4o-mini: Advancing cost-efficient intelligence, 2024. URL https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/

  18. [27]

    Jinghua et al. Piao. Emergence of human-like polarization among llm agents. arXiv preprint arXiv:2501.05171, 2025. doi:10.48550/arXiv.2501.05171. URL http://arxiv.org/abs/2501.05171

  19. [28]

    Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents

    Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Sch \"o lkopf, Mrinmaya Sachan, and Rada Mihalcea. Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  20. [29]

    Qwen2.5 technical report

    Qwen and et al. Qwen2.5 technical report. 0 (arXiv:2412.15115), January 2025. doi:10.48550/arXiv.2412.15115. URL http://arxiv.org/abs/2412.15115. arXiv:2412.15115 [cs]

  21. [30]

    What is loss aversion? Journal of Risk and Uncertainty, 30 0 (2): 0 157–167, March 2005

    Ulrich Schmidt and Horst Zank. What is loss aversion? Journal of Risk and Uncertainty, 30 0 (2): 0 157–167, March 2005. ISSN 1573-0476. doi:10.1007/s11166-005-6564-6

  22. [31]

    Schramowski

    Patrick et al. Schramowski. Llms contain human-like biases of right and wrong. arXiv preprint arXiv:2103.11790, 2022. doi:10.48550/arXiv.2103.11790. URL http://arxiv.org/abs/2103.11790

  23. [32]

    Greenhouse gas equivalencies calculator - calculations and references, August 2015

    OAR US EPA. Greenhouse gas equivalencies calculator - calculations and references, August 2015. URL https://www.epa.gov/energy/greenhouse-gas-equivalencies-calculator-calculations-and-references

  24. [33]

    Angelina et al. Wang. Llms that replace human participants can misportray identity groups. arXiv preprint arXiv:2402.01908, 2025. doi:10.48550/arXiv.2402.01908. URL http://arxiv.org/abs/2402.01908

  25. [34]

    Ethical and social risks of harm from language models

    Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zac Kenton, Sasha Brown, Will Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks...

  26. [35]

    Xiaolin et al. Xing. Evaluating knowledge-based cross-lingual inconsistency in llms. arXiv preprint arXiv:2407.01358, 2024. doi:10.48550/arXiv.2407.01358. URL http://arxiv.org/abs/2407.01358

  27. [36]

    Ye et al. Yuan. Measuring social norms of large language models. arXiv preprint arXiv:2404.02491, 2024. doi:10.48550/arXiv.2404.02491. URL http://arxiv.org/abs/2404.02491

  28. [37]

    Cultural value differences of llms: Prompt, language, and model size

    Qishuai Zhong, Yike Yun, and Aixin Sun. Cultural value differences of llms: Prompt, language, and model size. 0 (arXiv:2407.16891), June 2024. doi:10.48550/arXiv.2407.16891. URL http://arxiv.org/abs/2407.16891. arXiv:2407.16891 [cs]

  29. [38]

    Yiming et al. Zhu. Characterizing llm-driven social network: The chirper.ai case. arXiv preprint arXiv:2504.10286, 2025. doi:10.48550/arXiv.2504.10286. URL http://arxiv.org/abs/2504.10286

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.