Pith. sign in

REVIEW 2 major objections 6 minor 2 cited by

Using LLMs to Advance the Cognitive Science of Collectives

T0 review · 2 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper argues that LLMs, used as participants, interviewers, environments, routers, or analysts, can make collective cognition experimentally tractable.

desk verdict A useful organizing roadmap for LLMs in collective cognition, honest about its limits, but the participant role remains a promise awaiting validation. read the letter →

arxiv 2506.00052 v1 pith:O6QOBXH2 submitted 2025-05-28 q-bio.NC cs.AIcs.HCcs.MAcs.SI

classification q-bio.NCcs.AIcs.HCcs.MAcs.SI
keywords collectivecognitionlargelanguagemodelsmulti-agentsimulationsocialnetworksculturaldiversityLLMrolescognitivesciencehomogenization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This position paper argues that large language models are not just tools for studying individual cognition but a promising way to study cognition at the level of groups, organizations, and societies. It proposes that the field has struggled with three kinds of complexity: structural complexity of network topology, interactional complexity of ongoing relationships, and individual complexity of cultural and cognitive diversity. The paper maps five roles LLMs can play—participant, interviewer, environment, router, and data analyst—and shows how recent studies already instantiate some of them. It also argues that realizing the vision depends on solving alignment, cultural representation, homogenization, reproducibility, and compute-cost problems.

What carries the argument

The organizing device is a two-part taxonomy. The first part divides the difficulty of studying collectives into three axes: structural complexity at the network level (who is connected to whom), interactional complexity at the edge level (how relationships unfold over time and across modalities), and individual complexity at the node level (which agents are heterogeneous in beliefs and culture). The second part assigns LLMs five roles—participant, interviewer, environment, router, and data analyst—that can be deployed along these axes. The framework does the argument's work by turning an unwieldy problem into a grid: each axis names a bottleneck that has made collective cognition hard to scale, and each role names a concrete way an LLM can relieve that bottleneck, with existing studies mapped onto the grid as evidence that the roles are feasible.

What would settle it

Run the same collective task—for example, story transmission through a 50-node network or a deliberation-to-consensus protocol—with matched human and LLM-agent groups, varying network topology and cultural prompts; if the LLM groups fail to reproduce the qualitative distribution of outcomes, such as converging to consensus too readily or failing to restore cultural variation under persona prompting, the central premise collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLMs can be made into load-bearing instruments for the cognitive science of collectives, not merely for individual cognition. It identifies three axes of complexity that have blocked large-scale studies of group behavior: structural complexity (the size and topology of social networks), interactional complexity (the content and temporality of the connections between agents), and individual complexity (the beliefs, backgrounds, and cultural contexts of the nodes). For each axis, LLMs can contribute through at least one of five roles: as a participant that behaves like a human agent, as an interviewer that elicits data, as an environment that generates a shared world, as a router that summarizes and relays information among humans, and as an analyst that converts unstructured interaction data into quantified form. The paper draws on existing studies of deliberation, story-diffusion networks, and multilingual text analysis as proof-of-concept examples, and it argues that the remaining obstacles are identifiable research problems rather than reasons to abandon the program.

Load-bearing premise

The paper's program depends on LLM outputs being able to stand in for human cognitive behavior in collective settings while still preserving the heterogeneity and cultural diversity from which emergent group phenomena arise.

Editorial extensions

If this is right

  • Scaling along the structural axis becomes feasible: researchers could run hundreds of interacting agents, human and simulated, in controlled network topologies that would be prohibitively expensive to recruit and orchestrate by hand.
  • The same LLM can be reused across roles—as interviewer and analyst, or participant and router—so a single study can jointly address interactional and individual complexity rather than one axis at a time.
  • Collective phenomena such as the flow of cultural knowledge or the emergence of communicative norms could be studied over many simulated generations, a timescale that is inaccessible with human participants alone.
  • The risks named in the paper set the research agenda: improving representational alignment, measuring and counteracting homogenization, broadening cultural representation, ensuring reproducibility across model versions, and managing the computational cost of multi-agent simulations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test is to use LLM-as-router and LLM-as-participant in the same study, manipulating network topology while holding agent diversity fixed, to isolate the causal contribution of structure versus heterogeneity in emergent outcomes.
  • If LLM agents turn out to be over-homogeneous even under persona prompting, simulation-first pipelines may systematically underestimate polarization and overestimate consensus; a benchmark of collective-behavior distributions, not just mean outputs, would reveal this.
  • The framework generalizes beyond human cognitive science: hybrid human-AI collectives are themselves becoming real-world objects of study, so the same axes and roles could be used to forecast how AI presence reshapes deliberation, culture, and knowledge accumulation in actual societies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This perspective paper argues that large language models (LLMs) are underused in the cognitive science of collectives, and it proposes a framework with three axes of complexity—structural, interactional, and individual—and five roles for LLMs (participant, interviewer, environment, router, and data analyst). It illustrates the framework with existing studies (Tessler et al. 2024; Shiiku et al. 2025) and reviews risks such as homogenization, cultural misrepresentation, reproducibility, and compute costs. The paper makes no new empirical claims; it is a research agenda and a call for collaboration.

Significance. The paper's main contribution is organizational: it gives the community a vocabulary for where LLMs can enter collective-cognition research and it candidly lists capability gaps. If the heterogeneity and validity concerns can be resolved, the framework could provide a useful roadmap for scaling studies of consensus, cultural evolution, and networked decision-making. As a perspective, it does not require machine-checked proofs or code; its value is in framing. The authors should be credited for naming explicit risks rather than presenting an uncritical vision, although the current text does not offer a validation strategy for the strongest version of its own proposal.

major comments (2)
  1. [Missing capabilities and usage risks, 'Homogenization' and 'Cultural representation gaps'] These two subsections directly undermine the 'participant' role in Table 2, because the paper states that emergent group phenomena depend critically on heterogeneity and that LLMs produce narrower response distributions and reflect dominant cultural norms. No calibration, prompting strategy, or validation protocol is offered to show that LLM-agent collectives preserve the heterogeneity required for the collective phenomena the framework aims to study. This is load-bearing: the participant role is the mechanism for scaling individual complexity, and the only multi-agent example (Shiiku et al. 2025) demonstrates story-selection diversity, not that LLM agents reproduce known human collective outcomes. The paper should either reframe the central claim as a conditional agenda requiring validation or specify falsifiable benchmarks for heterogeneity preservation.
  2. [Taking on the axes of complexity with LLMs, paragraphs on Tessler et al. and Shiiku et al.] The two illustrative studies are presented as beginning to 'demonstrate the potential power' of LLMs, but neither study jointly addresses multiple axes, and neither validates LLM-agent collectives against human behavioral benchmarks. Given that the conclusion urges scaling along all three axes jointly, the manuscript should include a short validation agenda—for example, comparing LLM-agent collectives to human collectives on established tasks such as cultural transmission, consensus reaching, or convention formation—so that the proposal is actionable rather than merely suggestive.
minor comments (6)
  1. [Introduction, first paragraph] The term 'collective cognition' is never explicitly defined; a one-sentence definition near the start would help readers who are not specialists in this subfield.
  2. [Table 2, row 'Participant'] The cited example (Marjieh et al. 2024) concerns individual sensory judgments, not collective behavior; using an individual-level example to illustrate the participant role in collectives weakens the table's message.
  3. [Missing capabilities and usage risks, 'Homogenization'] The sentence citing Park et al. 2024 for 'personas' or character prompting is imprecise, since that work constructs personas from qualitative interviews rather than from simple character prompting.
  4. [Missing capabilities and usage risks, 'Compute costs'] The O(n^2) estimate for pairwise interactions assumes that all pairs of agents interact, which is not the case in many networked experimental designs; a brief qualification would avoid overstating the computational burden.
  5. [Looking ahead, first paragraph] The sentence beginning 'Of course, we note that LLMs are just one tool' is redundant after the preceding paragraphs and could be trimmed or merged with the following sentence.
  6. [Full text header] The running header 'A P REPRINT' appears to be a formatting artifact from the preprint template; it should be removed or corrected in a revised version.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation: this is a programmatic framework paper with illustrative self-citations, not a chain of claims that reduces to its own inputs.

full rationale

The paper does not present a derivation chain, fitted parameters, or predictions; it proposes a taxonomy (three axes of complexity, five LLM roles) and illustrates each role with existing studies. Several examples come from the authors' own prior work (e.g., Marjieh et al. 2024, Rathje et al. 2024, Collins et al. 2024), but these citations function as empirical illustrations, not as load-bearing premises that force the framework's conclusion. The central proposal is explicitly conditional ('we lay out how LLMs may be able to address...') and the paper lists open capability gaps rather than claiming to have closed them. Its own risk section acknowledges that LLMs homogenize responses and underrepresent cultural diversity, which is a substantive limitation of the participant role but not a circular move: the paper does not derive the success of the framework from the assumption that homogenization is absent. There is no equation equating an output to an input, no fitted value renamed as a prediction, and no uniqueness theorem imported from the authors' prior work. The self-citations are normal scholarly attribution in a perspective piece and do not create a circularity score above 1.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The framework rests on two unproved premises: that the axes describe the real barriers in collective cognition research, and that LLMs can preserve human diversity while acting as stand-ins. No free parameters or invented entities are introduced.

assumptions (3)
  • domain assumption LLMs can act as valid stand-ins for human participants or interaction partners in collective cognition studies.
    The entire proposal depends on this premise. It is cited to Binz et al. and Marjieh et al. but not established here, and the paper's own risk section partially undermines it.
  • ad hoc to paper The three axes (structural, interactional, individual complexity) adequately capture the main barriers to studying collectives.
    These axes are introduced by the authors as a taxonomy; no empirical or formal justification is provided for choosing exactly these three dimensions.
  • domain assumption Scaling along these axes with LLMs will produce scientifically valid insights about human collective behavior.
    The paper frames LLM-based simulations as enabling new cognitive science, but it does not validate that results from hybrid or AI-only collectives transfer to human groups.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Using LLMs to Advance the Cognitive Science of Collectives." pith.science (2026). https://pith.science/paper/O6QOBXH2

@misc{pith2026250600052,
  author       = {Pith},
  title        = {Pith review of: Using LLMs to Advance the Cognitive Science of Collectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O6QOBXH2}},
  note         = {Machine review of arXiv:2506.00052}
}
read the original abstract

LLMs are already transforming the study of individual cognition, but their application to studying collective cognition has been underexplored. We lay out how LLMs may be able to address the complexity that has hindered the study of collectives and raise possible risks that warrant new methods.

Figures

Figures reproduced from arXiv: 2506.00052 by the authors.

Figure 1
Figure 1. (Left) Axes of complexity in cognitive science. (i) Collective behavior, reflecting complexity at the network level (group structure and topology); (ii) Interactions, representing complexity at the edge level (connections within the social network); and (iii) Individual differences, capturing complexity at the node level (diversity among individuals). (Right) Roles of LLMs in cognitive science research. LLMs can ser… view at source ↗
Figure 2
Figure 2. Future experiments can explore integrating the three axes of complexity to raise and address [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.

  2. Generation and Evaluation in the Human Invention Process through the Lens of Game Design

    cs.HC 2025-08 reject novelty 5.0 of 10

    A two-stage model adding simulated-play funness to a language-model proposal prior best fits novice-invented games, but the model comparison is undermined by including the observed games in the normalization set and b...

Reference graph

Works this paper leans on

16 extracted references · 8 canonical work pages · cited by 2 Pith papers

  1. [1]

    Centaur: a foundation model of human cognition

    Marcel Binz, Elif Akata, Matthias Bethge, Franziska Br \"a ndle, Fred Callaway, Julian Coda-Forno, Peter Dayan, Can Demircan, Maria K Eckstein, No \'e mi \'E ltet o , et al. Centaur: a foundation model of human cognition. arXiv preprint arXiv:2410.20268, 2024

  2. [2]

    Machine culture

    Levin Brinkmann, Fabian Baumann, Jean-Fran c ois Bonnefon, Maxime Derex, Thomas F M \"u ller, Anne-Marie Nussberger, Agnieszka Czaplicka, Alberto Acerbi, Thomas L Griffiths, Joseph Henrich, et al. Machine culture. Nature Human Behaviour, 7 0 (11): 0 1855--1868, 2023

  3. [3]

    Building machines that learn and think with people

    Katherine M Collins, Ilia Sucholutsky, Umang Bhatt, Kartik Chandra, Lionel Wong, Mina Lee, Cedegao E Zhang, Tan Zhi-Xuan, Mark Ho, Vikash Mansinghka, et al. Building machines that learn and think with people. Nature Human Behaviour, 8 0 (10): 0 1851--1863, 2024

  4. [4]

    Large language models and games: A survey and roadmap

    Roberto Gallotta, Graham Todd, Marvin Zammit, Sam Earle, Antonios Liapis, Julian Togelius, and Georgios N Yannakakis. Large language models and games: A survey and roadmap. IEEE Transactions on Games, 2024

  5. [5]

    Hawkins, Michael Franke, Michael C

    Robert D. Hawkins, Michael Franke, Michael C. Frank, Adele E. Goldberg, Kenny Smith, Thomas L. Griffiths, and Noah D. Goodman. From partners to populations: A hierarchical bayesian account of coordination and convention. Psychological Review, 130 0 (4): 0 977--1016, 2023. doi:10.1037/rev0000348

  6. [6]

    Large language models predict human sensory judgments across six modalities

    Raja Marjieh, Ilia Sucholutsky, Pol van Rijn, Nori Jacoby, and Thomas L Griffiths. Large language models predict human sensory judgments across six modalities. Scientific Reports, 14 0 (1): 0 21445, 2024

  7. [7]

    Benchmarking distributional alignment of large language models

    Nicole Meister, Carlos Guestrin, and Tatsunori Hashimoto. Benchmarking distributional alignment of large language models. arXiv preprint arXiv:2411.05403, 2024

  8. [8]

    Generative agent simulations of 1,000 people

    Joon Sung Park, Carolyn Q Zou, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Robb Willer, Percy Liang, and Michael S Bernstein. Generative agent simulations of 1,000 people. arXiv preprint arXiv:2411.10109, 2024

Show all 16 references
  1. [9]

    Robertson, and Jay J

    Steve Rathje, Dan-Mircea Mirea, Ilia Sucholutsky, Raja Marjieh, Claire E. Robertson, and Jay J. Van Bavel. GPT is an effective tool for multilingual psychological text analysis. Proceedings of the National Academy of Sciences, 121 0 (34): 0 e2308950121, 2024

  2. [10]

    Unintended impacts of LLM alignment on global representation

    Michael Ryan, William Held, and Diyi Yang. Unintended impacts of LLM alignment on global representation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16121--16140, 2024

  3. [11]

    The dynamics of collective creativity in human-ai social networks

    Shota Shiiku, Raja Marjieh, Manuel Anglada-Tort, and Nori Jacoby. The dynamics of collective creativity in human-ai social networks. arXiv preprint arXiv:2502.17962, 2025

  4. [12]

    Getting aligned on representational alignment

    Ilia Sucholutsky, Lukas Muttenthaler, Adrian Weller, Andi Peng, Andreea Bobu, Been Kim, Bradley C Love, Erin Grant, Iris Groen, Jascha Achterberg, et al. Getting aligned on representational alignment. arXiv preprint arXiv:2310.13018, 2023

  5. [13]

    Bakker, Daniel Jarrett, Hannah Sheahan, Martin J

    Michael Henry Tessler, Michiel A. Bakker, Daniel Jarrett, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. AI can help humans find common ground in dem...

  6. [14]

    Complex cognitive algorithms preserved by selective social learning in experimental populations

    Bill Thompson, B Van Opheusden, T Sumers, and TL Griffiths. Complex cognitive algorithms preserved by selective social learning in experimental populations. Science, 376 0 (6588): 0 95--98, 2022

  7. [15]

    From word models to world models: Translating from natural language to the probabilistic language of thought

    Lionel Wong, Gabriel Grand, Alexander K Lew, Noah D Goodman, Vikash K Mansinghka, Jacob Andreas, and Joshua B Tenenbaum. From word models to world models: Translating from natural language to the probabilistic language of thought. arXiv preprint arXiv:2306.12672, 2023

  8. [16]

    On benchmarking human-like intelligence in machines

    Lance Ying, Katherine M Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L Griffiths, and Joshua B Tenenbaum. On benchmarking human-like intelligence in machines. arXiv preprint arXiv:2502.20502, 2025

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.