REVIEW 2 major objections 6 minor 2 cited by
Using LLMs to Advance the Cognitive Science of Collectives
T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that LLMs, used as participants, interviewers, environments, routers, or analysts, can make collective cognition experimentally tractable.
desk verdict A useful organizing roadmap for LLMs in collective cognition, honest about its limits, but the participant role remains a promise awaiting validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a two-part taxonomy. The first part divides the difficulty of studying collectives into three axes: structural complexity at the network level (who is connected to whom), interactional complexity at the edge level (how relationships unfold over time and across modalities), and individual complexity at the node level (which agents are heterogeneous in beliefs and culture). The second part assigns LLMs five roles—participant, interviewer, environment, router, and data analyst—that can be deployed along these axes. The framework does the argument's work by turning an unwieldy problem into a grid: each axis names a bottleneck that has made collective cognition hard to scale, and each role names a concrete way an LLM can relieve that bottleneck, with existing studies mapped onto the grid as evidence that the roles are feasible.
What would settle it
Run the same collective task—for example, story transmission through a 50-node network or a deliberation-to-consensus protocol—with matched human and LLM-agent groups, varying network topology and cultural prompts; if the LLM groups fail to reproduce the qualitative distribution of outcomes, such as converging to consensus too readily or failing to restore cultural variation under persona prompting, the central premise collapses.
Extended reading notes
Core claim
The paper's central claim is that LLMs can be made into load-bearing instruments for the cognitive science of collectives, not merely for individual cognition. It identifies three axes of complexity that have blocked large-scale studies of group behavior: structural complexity (the size and topology of social networks), interactional complexity (the content and temporality of the connections between agents), and individual complexity (the beliefs, backgrounds, and cultural contexts of the nodes). For each axis, LLMs can contribute through at least one of five roles: as a participant that behaves like a human agent, as an interviewer that elicits data, as an environment that generates a shared world, as a router that summarizes and relays information among humans, and as an analyst that converts unstructured interaction data into quantified form. The paper draws on existing studies of deliberation, story-diffusion networks, and multilingual text analysis as proof-of-concept examples, and it argues that the remaining obstacles are identifiable research problems rather than reasons to abandon the program.
Load-bearing premise
The paper's program depends on LLM outputs being able to stand in for human cognitive behavior in collective settings while still preserving the heterogeneity and cultural diversity from which emergent group phenomena arise.
Editorial extensions
If this is right
- Scaling along the structural axis becomes feasible: researchers could run hundreds of interacting agents, human and simulated, in controlled network topologies that would be prohibitively expensive to recruit and orchestrate by hand.
- The same LLM can be reused across roles—as interviewer and analyst, or participant and router—so a single study can jointly address interactional and individual complexity rather than one axis at a time.
- Collective phenomena such as the flow of cultural knowledge or the emergence of communicative norms could be studied over many simulated generations, a timescale that is inaccessible with human participants alone.
- The risks named in the paper set the research agenda: improving representational alignment, measuring and counteracting homogenization, broadening cultural representation, ensuring reproducibility across model versions, and managing the computational cost of multi-agent simulations.
Reading between the lines
- A natural next test is to use LLM-as-router and LLM-as-participant in the same study, manipulating network topology while holding agent diversity fixed, to isolate the causal contribution of structure versus heterogeneity in emergent outcomes.
- If LLM agents turn out to be over-homogeneous even under persona prompting, simulation-first pipelines may systematically underestimate polarization and overestimate consensus; a benchmark of collective-behavior distributions, not just mean outputs, would reveal this.
- The framework generalizes beyond human cognitive science: hybrid human-AI collectives are themselves becoming real-world objects of study, so the same axes and roles could be used to forecast how AI presence reshapes deliberation, culture, and knowledge accumulation in actual societies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective paper argues that large language models (LLMs) are underused in the cognitive science of collectives, and it proposes a framework with three axes of complexity—structural, interactional, and individual—and five roles for LLMs (participant, interviewer, environment, router, and data analyst). It illustrates the framework with existing studies (Tessler et al. 2024; Shiiku et al. 2025) and reviews risks such as homogenization, cultural misrepresentation, reproducibility, and compute costs. The paper makes no new empirical claims; it is a research agenda and a call for collaboration.
Significance. The paper's main contribution is organizational: it gives the community a vocabulary for where LLMs can enter collective-cognition research and it candidly lists capability gaps. If the heterogeneity and validity concerns can be resolved, the framework could provide a useful roadmap for scaling studies of consensus, cultural evolution, and networked decision-making. As a perspective, it does not require machine-checked proofs or code; its value is in framing. The authors should be credited for naming explicit risks rather than presenting an uncritical vision, although the current text does not offer a validation strategy for the strongest version of its own proposal.
major comments (2)
- [Missing capabilities and usage risks, 'Homogenization' and 'Cultural representation gaps'] These two subsections directly undermine the 'participant' role in Table 2, because the paper states that emergent group phenomena depend critically on heterogeneity and that LLMs produce narrower response distributions and reflect dominant cultural norms. No calibration, prompting strategy, or validation protocol is offered to show that LLM-agent collectives preserve the heterogeneity required for the collective phenomena the framework aims to study. This is load-bearing: the participant role is the mechanism for scaling individual complexity, and the only multi-agent example (Shiiku et al. 2025) demonstrates story-selection diversity, not that LLM agents reproduce known human collective outcomes. The paper should either reframe the central claim as a conditional agenda requiring validation or specify falsifiable benchmarks for heterogeneity preservation.
- [Taking on the axes of complexity with LLMs, paragraphs on Tessler et al. and Shiiku et al.] The two illustrative studies are presented as beginning to 'demonstrate the potential power' of LLMs, but neither study jointly addresses multiple axes, and neither validates LLM-agent collectives against human behavioral benchmarks. Given that the conclusion urges scaling along all three axes jointly, the manuscript should include a short validation agenda—for example, comparing LLM-agent collectives to human collectives on established tasks such as cultural transmission, consensus reaching, or convention formation—so that the proposal is actionable rather than merely suggestive.
minor comments (6)
- [Introduction, first paragraph] The term 'collective cognition' is never explicitly defined; a one-sentence definition near the start would help readers who are not specialists in this subfield.
- [Table 2, row 'Participant'] The cited example (Marjieh et al. 2024) concerns individual sensory judgments, not collective behavior; using an individual-level example to illustrate the participant role in collectives weakens the table's message.
- [Missing capabilities and usage risks, 'Homogenization'] The sentence citing Park et al. 2024 for 'personas' or character prompting is imprecise, since that work constructs personas from qualitative interviews rather than from simple character prompting.
- [Missing capabilities and usage risks, 'Compute costs'] The O(n^2) estimate for pairwise interactions assumes that all pairs of agents interact, which is not the case in many networked experimental designs; a brief qualification would avoid overstating the computational burden.
- [Looking ahead, first paragraph] The sentence beginning 'Of course, we note that LLMs are just one tool' is redundant after the preceding paragraphs and could be trimmed or merged with the following sentence.
- [Full text header] The running header 'A P REPRINT' appears to be a formatting artifact from the preprint template; it should be removed or corrected in a revised version.
Circularity Check
No circular derivation: this is a programmatic framework paper with illustrative self-citations, not a chain of claims that reduces to its own inputs.
full rationale
The paper does not present a derivation chain, fitted parameters, or predictions; it proposes a taxonomy (three axes of complexity, five LLM roles) and illustrates each role with existing studies. Several examples come from the authors' own prior work (e.g., Marjieh et al. 2024, Rathje et al. 2024, Collins et al. 2024), but these citations function as empirical illustrations, not as load-bearing premises that force the framework's conclusion. The central proposal is explicitly conditional ('we lay out how LLMs may be able to address...') and the paper lists open capability gaps rather than claiming to have closed them. Its own risk section acknowledges that LLMs homogenize responses and underrepresent cultural diversity, which is a substantive limitation of the participant role but not a circular move: the paper does not derive the success of the framework from the assumption that homogenization is absent. There is no equation equating an output to an input, no fitted value renamed as a prediction, and no uniqueness theorem imported from the authors' prior work. The self-citations are normal scholarly attribution in a perspective piece and do not create a circularity score above 1.
Assumptions & free parameters
assumptions (3)
- domain assumption LLMs can act as valid stand-ins for human participants or interaction partners in collective cognition studies.
- ad hoc to paper The three axes (structural, interactional, individual complexity) adequately capture the main barriers to studying collectives.
- domain assumption Scaling along these axes with LLMs will produce scientifically valid insights about human collective behavior.
Cite this review
Pith. "Pith review of Using LLMs to Advance the Cognitive Science of Collectives." pith.science (2026). https://pith.science/paper/O6QOBXH2
@misc{pith2026250600052,
author = {Pith},
title = {Pith review of: Using LLMs to Advance the Cognitive Science of Collectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/O6QOBXH2}},
note = {Machine review of arXiv:2506.00052}
}
read the original abstract
LLMs are already transforming the study of individual cognition, but their application to studying collective cognition has been underexplored. We lay out how LLMs may be able to address the complexity that has hindered the study of collectives and raise possible risks that warrant new methods.
Figures
Forward citations
Cited by 2 Pith papers
-
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.
-
Generation and Evaluation in the Human Invention Process through the Lens of Game Design
A two-stage model adding simulated-play funness to a language-model proposal prior best fits novice-invented games, but the model comparison is undermined by including the observed games in the normalization set and b...
Reference graph
Works this paper leans on
-
[1]
Centaur: a foundation model of human cognition
Marcel Binz, Elif Akata, Matthias Bethge, Franziska Br \"a ndle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, No \'e mi \'E ltet o , et al. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268, 2024
-
[2]
Levin Brinkmann, Fabian Baumann, Jean-Fran c ois Bonnefon, Maxime Derex, Thomas F M \"u ller, Anne-Marie Nussberger, Agnieszka Czaplicka, Alberto Acerbi, Thomas L Griffiths, Joseph Henrich, et al. Machine culture. Nature Human Behaviour, 7 0 (11): 0 1855--1868, 2023
work page 2023
-
[3]
Building machines that learn and think with people
Katherine M Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, et al. Building machines that learn and think with people. Nature Human Behaviour, 8 0 (10): 0 1851--1863, 2024
work page 2024
-
[4]
Large language models and games: A survey and roadmap
Roberto Gallotta, Graham Todd, Marvin Zammit, Sam Earle, Antonios Liapis, Julian Togelius, and Georgios N Yannakakis. Large language models and games: A survey and roadmap. IEEE Transactions on Games, 2024
2024
-
[5]
Hawkins, Michael Franke, Michael C
Robert D. Hawkins, Michael Franke, Michael C. Frank, Adele E. Goldberg, Kenny Smith, Thomas L. Griffiths, and Noah D. Goodman. From partners to populations: A hierarchical bayesian account of coordination and convention. Psychological Review, 130 0 (4): 0 977--1016, 2023. doi:10.1037/rev0000348
-
[6]
Large language models predict human sensory judgments across six modalities
Raja Marjieh, Ilia Sucholutsky, Pol van Rijn, Nori Jacoby, and Thomas L Griffiths. Large language models predict human sensory judgments across six modalities. Scientific Reports, 14 0 (1): 0 21445, 2024
work page 2024
-
[7]
Benchmarking distributional alignment of large language models
Nicole Meister, Carlos Guestrin, and Tatsunori Hashimoto. Benchmarking distributional alignment of large language models. arXiv preprint arXiv:2411.05403, 2024
arXiv 2024
-
[8]
Generative agent simulations of 1,000 people
Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109, 2024
arXiv 2024
Show all 16 references
-
[9]
Robertson, and Jay J
Steve Rathje, Dan-Mircea Mirea, Ilia Sucholutsky, Raja Marjieh, Claire E. Robertson, and Jay J. Van Bavel. GPT is an effective tool for multilingual psychological text analysis. Proceedings of the National Academy of Sciences, 121 0 (34): 0 e2308950121, 2024
2024
-
[10]
Unintended impacts of LLM alignment on global representation
Michael Ryan, William Held, and Diyi Yang. Unintended impacts of LLM alignment on global representation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16121--16140, 2024
2024
-
[11]
The dynamics of collective creativity in human-ai social networks
Shota Shiiku, Raja Marjieh, Manuel Anglada-Tort, and Nori Jacoby. The dynamics of collective creativity in human-ai social networks. arXiv preprint arXiv:2502.17962, 2025
2025 arXiv
-
[12]
Getting aligned on representational alignment
Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C Love, Erin Grant, Iris Groen, Jascha Achterberg, et al. Getting aligned on representational alignment. arXiv preprint arXiv:2310.13018, 2023
-
[13]
Bakker, Daniel Jarrett, Hannah Sheahan, Martin J
Michael Henry Tessler, Michiel A. Bakker, Daniel Jarrett, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. AI can help humans find common ground in dem...
2024
-
[14]
Complex cognitive algorithms preserved by selective social learning in experimental populations
Bill Thompson, B Van Opheusden, T Sumers, and TL Griffiths. Complex cognitive algorithms preserved by selective social learning in experimental populations. Science, 376 0 (6588): 0 95--98, 2022
2022
-
[15]
From word models to world models: Translating from natural language to the probabilistic language of thought
Lionel Wong, Gabriel Grand, Alexander K Lew, Noah D Goodman, Vikash K Mansinghka, Jacob Andreas, and Joshua B Tenenbaum. From word models to world models: Translating from natural language to the probabilistic language of thought. arXiv preprint arXiv:2306.12672, 2023
2023 arXiv
-
[16]
On benchmarking human-like intelligence in machines
Lance Ying, Katherine M Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L Griffiths, and Joshua B Tenenbaum. On benchmarking human-like intelligence in machines. arXiv preprint arXiv:2502.20502, 2025
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.