REVIEW 3 major objections 2 minor 38 references
Group selection on transmitted prompts promotes and stabilizes prosocial behavior in populations of LLM agents playing social dilemmas.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Group selection on transmitted prompts in LLM agent populations promotes and stabilizes cooperation in social dilemmas while individual selection produces collective defection.
T0 review reviewed 2026-06-26 challenge →
load-bearing objection Group selection on prompt copying stabilizes cooperation in these LLM simulations while individual selection does not, but the mechanism may not be specifically prosocial content. the 3 major comments →
Group Selection Promotes Prosocial Prompts in Populations of LLM Agents
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is that in populations of LLM agents playing repeated social dilemma games, transmitting natural-language prompts from high-performing groups to the next generation promotes prosociality and stabilizes cooperation, whereas selection at the individual level allows self-interested prompts to dominate and drive populations to collective defection. This pattern is robust and can be reproduced theoretically with a replicator-mutator model that predicts a phase transition.
What carries the argument
Evolutionary transmission of natural-language prompts under group-level versus individual-level selection in a multi-agent simulation of a social dilemma game.
Load-bearing premise
Copying prompts from successful groups will transmit the prosocial elements of those prompts rather than unrelated features that merely correlate with group performance.
What would settle it
Running the simulation under group selection but finding that average cooperation levels do not rise over generations or that prosocial prompts do not increase in frequency.
If this is right
- Cooperation persists across generations only when selection acts on groups.
- Selfish prompts spread rapidly under individual selection, collapsing cooperation.
- The effect is independent of specific prompt wording or model choice within tested ranges.
- A critical threshold exists in the transmission process beyond which cooperation stabilizes.
Where Pith is reading between the lines
- This suggests that multi-agent AI systems may require group-based evaluation to avoid defection equilibria.
- Agents might develop the ability to anticipate and adapt to known selection pressures, as seen in one model.
- Extending the framework to other games or real-world tasks could test the generality of prompt evolution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a multi-agent simulation in which LLM agents play repeated social dilemma games and transmit natural-language prompts across generations under either individual or group selection. It claims that group selection transmits prompts from high-performing groups, thereby promoting prosociality and stabilizing cooperation, while individual selection favors self-interested prompts and leads to collective defection. The gap is reported as robust across prompt ablations, game framings, and model swaps. A replicator-mutator model with an empirically estimated transmission kernel reproduces key results and predicts a phase transition; preliminary observations also note anticipatory donation adjustment in one model (GPT-5.4).
Significance. If the central claim holds after addressing the transmission mechanism, the work would demonstrate that unguided group-level selection on natural-language prompts can evolve and stabilize cooperation in LLM populations without human-specified individual rewards. The combination of agent-based simulation, cross-model robustness checks, and a matching replicator-mutator model supplies a falsifiable framework that could inform the design of multi-agent LLM systems.
major comments (3)
- [Results (behavioral outcomes and prompt transmission)] The central claim that group selection specifically transmits and amplifies prosocial prompt content (rather than other correlates of group success) is load-bearing for the interpretation. The manuscript reports behavioral outcomes and robustness across ablations but does not supply a direct content analysis (e.g., keyword frequency, sentiment, or embedding comparison) of the transmitted prompts under the two regimes; without this, the observed cooperation could arise from non-prosocial prompt traits that happen to correlate with performance in the chosen game.
- [§4] §4 (replicator-mutator model): The empirical transmission kernel is stated to predict a phase transition at a critical threshold, yet the manuscript does not detail how the kernel is estimated from the LLM simulation runs (sample size, mutation rate, or how prosocial vs. non-prosocial prompt features are encoded in the kernel). This leaves open whether the theoretical reproduction assumes the very prosocial transmission that the simulations are meant to demonstrate.
- [Methods and Results (robustness checks)] The abstract asserts robustness across prompt ablations, alternative game framings, and model swaps, but the methods section supplies no statistical tests, error bars, or pre-registered analysis plan for these checks. It is therefore impossible to evaluate whether the reported gap between selection regimes survives multiple-comparison correction or is driven by a subset of conditions.
minor comments (2)
- [§2] Notation for the repeated social dilemma payoff matrix is introduced without an explicit equation number; adding Eq. (1) would improve readability when the replicator-mutator model is later compared to the simulation.
- [Figures 2-4] Figure captions for the generational trajectories do not state the number of independent runs or the precise definition of 'high-performing group' used for prompt transmission.
Simulated Author's Rebuttal
We thank the referee for their constructive comments, which identify key areas where additional detail and analysis would strengthen the manuscript. We respond to each major comment below and indicate planned revisions.
read point-by-point responses
-
Referee: [Results (behavioral outcomes and prompt transmission)] The central claim that group selection specifically transmits and amplifies prosocial prompt content (rather than other correlates of group success) is load-bearing for the interpretation. The manuscript reports behavioral outcomes and robustness across ablations but does not supply a direct content analysis (e.g., keyword frequency, sentiment, or embedding comparison) of the transmitted prompts under the two regimes; without this, the observed cooperation could arise from non-prosocial prompt traits that happen to correlate with performance in the chosen game.
Authors: We agree that a direct content analysis of transmitted prompts would provide stronger support for the claim that prosocial content is specifically selected and amplified. The current manuscript relies on behavioral outcomes and robustness across conditions to support the interpretation. We will add a new analysis subsection comparing prompt content under the two regimes, using embedding similarity and keyword frequencies for prosocial versus self-interested language. revision: yes
-
Referee: [§4] §4 (replicator-mutator model): The empirical transmission kernel is stated to predict a phase transition at a critical threshold, yet the manuscript does not detail how the kernel is estimated from the LLM simulation runs (sample size, mutation rate, or how prosocial vs. non-prosocial prompt features are encoded in the kernel). This leaves open whether the theoretical reproduction assumes the very prosocial transmission that the simulations are meant to demonstrate.
Authors: The kernel is fitted directly to observed transmission events from the LLM simulations. We will expand §4 with explicit details on the estimation procedure, including the number of simulation runs, the mutation rate used, and the encoding of prompt features based on associated behavioral outcomes, to clarify that no assumption of prosocial transmission is built into the model construction. revision: yes
-
Referee: [Methods and Results (robustness checks)] The abstract asserts robustness across prompt ablations, alternative game framings, and model swaps, but the methods section supplies no statistical tests, error bars, or pre-registered analysis plan for these checks. It is therefore impossible to evaluate whether the reported gap between selection regimes survives multiple-comparison correction or is driven by a subset of conditions.
Authors: The robustness checks are presented as consistent patterns across conditions. We will add error bars to relevant figures and include statistical comparisons (with multiple-comparison adjustments noted) between selection regimes. A pre-registered plan cannot be added retrospectively, but the analysis approach will be documented transparently in the methods. revision: partial
Circularity Check
No circularity: results from direct simulation and empirical kernel reproduction
full rationale
The paper reports outcomes from a multi-agent LLM simulation under individual vs. group selection, with prompt transmission observed to affect cooperation levels. A replicator-mutator model is then used to reproduce those outcomes via an empirical transmission kernel derived from the same simulations. No equations, fitted parameters, or self-citations are shown to create a definitional loop or to rename simulation outputs as independent predictions. The derivation chain remains self-contained in the simulation framework and its direct measurements.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Group Selection Promotes Prosocial Prompts in Populations of LLM Agents." pith.science (2026). https://pith.science/paper/MTGR2A5E
@misc{pith2026260623343,
author = {Pith},
title = {Pith review of: Group Selection Promotes Prosocial Prompts in Populations of LLM Agents},
year = {2026},
howpublished = {\url{https://pith.science/paper/MTGR2A5E}},
note = {Machine review of arXiv:2606.23343}
}
read the original abstract
Current approaches to instill prosociality in large language model (LLM) agents often rely on humans specifying desired behaviors at the individual level, which does not guarantee cooperation within LLM populations. As frontier training shifts toward individual rewards for verifiable tasks, such as mathematics and coding, this outcome-based focus may further undermine cooperation in multi-agent settings. Large-scale cooperation in human populations emerged via unguided evolutionary mechanisms, not a central architect. Group selection, in which cooperative groups within a population outcompete less cooperative ones, has been argued to be essential. In this study, we explore whether group selection can promote cooperation in populations of LLM agents. We introduce a multi-agent simulation framework in which LLM agents play a repeated social dilemma game and transmit their natural-language prompts across generations under either individual- or group-level selection. Under group selection, prompts from high-performing groups are transmitted, thereby promoting prosociality and stabilizing cooperation. Under individual selection, self-interested prompts dominate, causing populations to collapse into collective defection. This gap is robust across prompt ablations, alternative game framings, and model swaps. We theoretically reproduce key results using a replicator-mutator model, whose empirical transmission kernel predicts a phase transition at a critical threshold. Preliminary findings show that, when informed about the selection mechanism, GPT-5.4 preemptively and gradually adjusts first-generation donations. This demonstrates strong anticipatory behavior that was not observed in the other tested models. These results demonstrate that prosocial prompts and cooperative behaviors evolve in LLM agent populations under group selection.
Figures
Reference graph
Works this paper leans on
-
[1]
Laine, Rudolf and Chughtai, Bilal and Betley, Jan and Hariharan, Kaivalya and Balesni, Mikita and Scheurer, J. Me,. 38th Conference on Neural Information Processing Systems (NeurIPS 2024) , year =
2024
-
[2]
, publisher =
Boyd, Robert and Richerson, Peter J. , publisher =. Culture and the. 1985 , address =
1985
-
[3]
, title =
Boyd, Robert and Richerson, Peter J. , title =. Philosophical Transactions of the Royal Society B: Biological Sciences , volume =. 2009 , month =
2009
-
[4]
and Nussberger, Anne-Marie and Czaplicka, Agnieszka and Acerbi, Alberto and Griffiths, Thomas L
Brinkmann, Levin and Baumann, Fabian and Bonnefon, Jean-François and Derex, Maxime and Müller, Thomas F. and Nussberger, Anne-Marie and Czaplicka, Agnieszka and Acerbi, Alberto and Griffiths, Thomas L. and Henrich, Joseph and Leibo, Joel Z. and McElreath, Richard and Oudeyer, Pierre-Yves and Stray, Jonathan and Rahwan, Iyad , title =. Nature Human Behavio...
2023
-
[5]
Economics Bulletin , volume =
Brookins, Philip and DeBacker, Jason , title =. Economics Bulletin , volume =. 2024 , month =
2024
-
[6]
, title =
Cooney, Daniel B. , title =. Bulletin of Mathematical Biology , volume =. 2022 , month =
2022
-
[7]
and Szathmáry, Eörs , title =
Czégel, Dániel and Giaffar, Hamza and Tenenbaum, Joshua B. and Szathmáry, Eörs , title =. BioEssays , volume =. 2022 , url =
2022
-
[8]
Nature , volume =
Efferson, Charles and Bernhard, Helen and Fischbacher, Urs and Fehr, Ernst , title =. Nature , volume =. 2024 , url =
2024
-
[9]
Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution , booktitle =
Fernando, Chrisantha and Banarse, Dylan Sunil and Michalewski, Henryk and Osindero, Simon and Rockt. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution , booktitle =. 2024 , url =
2024
-
[10]
Nicer than Humans: How Do Large Language Models Behave in the Prisoner's Dilemma? , journal =
Fontana, Nicol. Nicer than Humans: How Do Large Language Models Behave in the Prisoner's Dilemma? , journal =. 2025 , url =
2025
-
[11]
SIAM Journal on Applied Mathematics , volume =
Hadeler, Karl Peter , title =. SIAM Journal on Applied Mathematics , volume =. 1981 , month =
1981
-
[12]
arXiv preprint arXiv:2510.04860 , year =
Han, Siwei and Xiong, Kaiwen and Liu, Jiaqi and Ye, Xinyu and Su, Yaofeng and Duan, Wenbo and Liu, Xinyuan and Xie, Cihang and Bansal, Mohit and Ding, Mingyu and Zhang, Linjun and Yao, Huaxiu , title =. arXiv preprint arXiv:2510.04860 , year =
-
[13]
Proceedings of the Royal Society B: Biological Sciences , volume =
Hauert, Christoph and Holmes, Miranda and Doebeli, Michael , title =. Proceedings of the Royal Society B: Biological Sciences , volume =. 2006 , month =
2006
-
[14]
Journal of Economic Behavior & Organization , volume =
Henrich, Joseph , title =. Journal of Economic Behavior & Organization , volume =. 2004 , url =
2004
-
[15]
Annual Review of Psychology , volume =
Henrich, Joseph and Muthukrishna, Michael , title =. Annual Review of Psychology , volume =. 2021 , url =
2021
-
[16]
arXiv preprint arXiv:2409.00993 , year =
Horiguchi, Ilya and Yoshida, Takahide and Ikegami, Takashi , title =. arXiv preprint arXiv:2409.00993 , year =
-
[17]
AI Alignment: A Comprehensive Survey
Ji, Jiaming and Qiu, Tianyi and Chen, Boyuan and Zhang, Borong and Lou, Hantao and Wang, Kaile and Duan, Yawen and He, Zhonghao and Vierling, Lukas and Hong, Donghai and Zhou, Jiayi and Zhang, Zhaowei and Zeng, Fanzhi and Dai, Juntao and Pan, Xuehai and Ng, Kwan Yee and O'Gara, Aidan and Xu, Hua and Tse, Brian and Fu, Jie and McAleer, Stephen and Yang, Ya...
work page internal anchor Pith review Pith/arXiv arXiv
- [18]
-
[19]
and Zambaldi, Vinicius and Lanctot, Marc and Marecki, Janusz and Graepel, Thore , title =
Leibo, Joel Z. and Zambaldi, Vinicius and Lanctot, Marc and Marecki, Janusz and Graepel, Thore , title =. Proceedings of the 16th Conference on Autonomous Agents and MultiAgent Systems , pages =. 2017 , url =
2017
-
[20]
arXiv preprint arXiv:2401.04620 , year =
Li, Shimin and Sun, Tianxiang and Cheng, Qinyuan and Qiu, Xipeng , title =. arXiv preprint arXiv:2401.04620 , year =
-
[21]
arXiv preprint arXiv:2602.14471 , year =
Mumcu, Furkan and Yilmaz, Yasin , title =. arXiv preprint arXiv:2602.14471 , year =
-
[22]
and Sigmund, Karl , title =
Nowak, Martin A. and Sigmund, Karl , title =. Journal of Theoretical Biology , volume =. 1998 , month =
1998
-
[23]
and Komarova, Natalia L
Nowak, Martin A. and Komarova, Natalia L. and Niyogi, Partha , title =. Science , volume =. 2001 , month =
2001
-
[24]
, title =
Nowak, Martin A. , title =. Science , volume =. 2006 , month =
2006
-
[25]
, publisher =
Nowak, Martin A. , publisher =. Evolutionary. 2006 , address =
2006
-
[26]
and Nowak, Martin A
Page, Karen M. and Nowak, Martin A. , title =. Journal of Theoretical Biology , volume =. 2002 , month =
2002
-
[27]
and Cai, Carrie J
Park, Joon Sung and O'Brien, Joseph C. and Cai, Carrie J. and Morris, Meredith Ringel and Liang, Percy and Bernstein, Michael S. , title =. Proceedings of the 36th Annual. 2023 , month =
2023
-
[28]
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of
Piatti, Giorgio and Jin, Zhijing and Kleiman-Weiner, Max and Sch. Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of. Advances in Neural Information Processing Systems , volume =. 2024 , url =
2024
-
[29]
arXiv preprint arXiv:2505.09388 , year =
work page internal anchor Pith review Pith/arXiv arXiv
-
[30]
2026 , url =
GPT-5.4 , author =. 2026 , url =
2026
-
[31]
and Demps, Kathryn and Frost, Karl and Hillis, Vicken and Mathew, Sarah and Newton, Emily K
Richerson, Peter and Baldini, Ryan and Bell, Adrian V. and Demps, Kathryn and Frost, Karl and Hillis, Vicken and Mathew, Sarah and Newton, Emily K. and Naar, Nicole and Newson, Lesley and Ross, Cody and Smaldino, Paul E. and Waring, Timothy M. and Zefferman, Matthew , title =. Behavioral and Brain Sciences , volume =. 2016 , url =
2016
-
[32]
Sigmund, Karl , title =
-
[33]
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
Tewolde, Emanuel and Zhang, Xiao and Guzman Piedrahita, David and Conitzer, Vincent and Jin, Zhijing , title =. arXiv preprint arXiv:2604.15267 , year =
work page internal anchor Pith review Pith/arXiv arXiv
-
[34]
arXiv preprint arXiv:2412.10270 , year =
Vallinder, Aron and Hughes, Edward , title =. arXiv preprint arXiv:2412.10270 , year =
-
[35]
Hammond, Lewis and Chan, Alan and Clifton, Jesse and Hoelscher-Obermaier, Jason and Khan, Akbir and McLean, Euan and Smith, Chandler and Barfuss, Wolfram and Foerster, Jakob and Gaven. Multi-. 2025 , journal =
2025
-
[36]
2025 , url =
Nature , author =. 2025 , url =
2025
-
[37]
AlphaEvolve: A coding agent for scientific and algorithmic discovery
Novikov, Alexander and V. arXiv preprint arXiv:2506.13131 , year =
work page internal anchor Pith review Pith/arXiv arXiv
-
[38]
2024 , url =
Llama 3 Model Card , author =. 2024 , url =
2024
This paper was first reviewed by grok-4.3 on June 26, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.