Pith. sign in

REVIEW 4 major objections 6 minor 4 references

Large-scale Group Brainstorming using Conversational Swarm Intelligence (CSI) versus Traditional Chat

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read In a head-to-head test with 147 survey responses, groups of 75 people significantly preferred AI-woven small-group 'swarm' brainstorming over one large text chat room on every measure.

desk verdict Plausible preference result for CSI over chat, but the missing surrogate-fidelity check keeps the central claim from being fully isolated. read the letter →

arxiv 2412.14205 v1 pith:C4QHI3MS submitted 2024-12-16 cs.HC cs.AIcs.SI

classification cs.HCcs.AIcs.SI
keywords CollaborationDeliberationCollectiveIntelligenceGenerativeAIConversationalSwarmLargeLanguageModelsBrainstormingAlternativeUseTasks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports a direct comparison of two ways for a large online group to brainstorm in real time: a single large chat room versus a structure called Conversational Swarm Intelligence (CSI), in which participants are split into five-person subgroups and LLM-powered conversational surrogates pass distilled ideas between subgroups. It tries to establish that CSI is the better experience for brainstorming and prioritization at scale. Across 147 survey responses from two 75-person groups, a significant majority preferred CSI on all seven subjective questions, including feeling more collaborative, more productive, hearing better answers, feeling more heard, and having more ownership and buy-in. Overall preference for CSI was 75%, with per-question support between 66% and 88%.

What carries the argument

The load-bearing mechanism is the Conversational Surrogate: an LLM-powered agent placed in each 4–7-person subgroup that observes the local conversation, distills the salient ideas and opinions, and passes them to surrogates in other subgroups, which then voice them in their local deliberations. This creates a fully connected network of overlapping conversations that emulates the information propagation of fish schools without requiring any human to follow multiple threads at once. A matchmaking subsystem tracks which subgroups are ready to receive a new insight and which available insights would most challenge the receiving group, so ideas propagate by merit rather than by a few strong voices.

What would settle it

Blind-rate or count the actual brainstorm ideas produced under each structure in a fresh sample: if the CSI condition does not yield more or better alternative uses than the single chat room, then the strong subjective preference does not reflect objectively better brainstorming output.

Watch

Extended reading notes

Core claim

The central discovery is that large networked groups of roughly 75 people can conduct a real-time brainstorming conversation using CSI and that the participants' subjective experience is systematically better than in a traditional text chat room. The paper shows that on every one of seven survey questions the CSI structure was preferred with statistical significance at a Bonferroni-adjusted 1% level, with an overall preference of 75% and question-specific preferences from 66% to 88%. The authors interpret this as evidence that CSI's architecture of overlapping subgroups woven together by AI surrogate agents preserves the benefits of small-group deliberation while allowing the full population to converge on a short list of prioritized answers.

Load-bearing premise

The study assumes the conversational surrogates faithfully represent subgroup views and that the AI agents themselves did not shape participants' preferences, because no fidelity check or objective measure of brainstorm output was collected.

Editorial extensions

If this is right

  • Large-scale real-time brainstorming and prioritization can be run effectively in text with groups of at least 75 people, a scale where a single chat room normally degrades into monologues.
  • Participants report more balanced participation and greater buy-in with CSI, which may reduce the influence of dominant personalities and early talkers.
  • A fully connected surrogate network can propagate insights to any subgroup, making convergence to prioritized answers more efficient than in natural fish-school-style neighbor-only propagation.
  • The same structure could scale to hundreds or thousands of users, though the present data only directly support groups of about 75.
  • CSI may be useful for enterprise feedback, citizen assemblies, and other deliberative tasks where large-group thoughtfulness is difficult to achieve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • (Inference) The preference results say nothing yet about objective output quality, so a fair next test is to blind-rate the actual ideas produced in each condition; the paper collected no such measure.
  • (Inference) If surrogates are faithful, the same architecture should transfer to voice or video deliberation, but the 12-minute text format and the use of two fixed problem types leave modality effects unknown.
  • (Inference) The design did not include a non-AI small-group condition, so it cannot separate the benefit of small subgroups from the benefit of AI weaving; a control with isolated subgroups but no surrogate could isolate the mechanism.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript reports a between-subjects counterbalanced experiment in which two groups of about 75 participants each performed two Alternative Use Task brainstorming exercises, once in a Conversational Swarm Intelligence (CSI) platform (Thinkscape) and once in a single large text-based chat room. After the two sessions, participants answered seven forced-choice questions comparing the two experiences. The authors report that a significant majority preferred CSI on all seven items, with support ranging from 66% to 88% and an overall preference of 75%. They conclude that CSI is a promising method for large-scale group brainstorming and prioritization.

Significance. If the result holds, the paper provides useful evidence that a structured, multi-subgroup, AI-mediated format can feel more collaborative, productive, and fair than a single large chat room for groups of tens of participants. The study has real strengths: it compares against an external control condition, uses a within-subject design, counterbalances condition order across two groups, and applies an appropriate one-proportion z-test with Bonferroni correction. The central claim, however, depends on an unverified assumption that the LLM-powered conversational surrogates only transmitted human subgroup content and did not introduce their own ideas or opinions. Without a manipulation check, the preference result could be driven by AI-generated content or style rather than by the swarm-weaving structure. The paper also reports no objective brainstorming outcomes, so the broader conclusion that groups 'can successfully brainstorm and prioritize' rests entirely on self-report. These issues are fixable, and the subjective-preference result is still interesting and worth reporting once the missing verification and data are supplied.

major comments (4)
  1. [Section 2] The key assumption that separates the CSI condition from ordinary chat is the statement that 'The AI agents did not introduce any AI generated ideas or opinions into the local conversations – they only passed and received conversational ideas and opinions from other subgroups.' No manipulation check is reported to verify this. Because the surrogates are LLM-powered and generate natural-language utterances, the observed preference could be driven by AI-authored ideas, AI tone, or AI pacing rather than by the weaving structure itself. The paper itself notes that every assertion is stored in a real-time taxonomy database, so a log audit or content analysis (for example, the proportion of surrogate utterances traceable to human subgroup messages) is feasible and should be reported. If such verification cannot be provided, the causal interpretation should be weakened and this limitation stated explicitly.
  2. [Sections 3-4] The statistical analysis is not fully reported. The text says a one-proportion z-test with Bonferroni correction (alpha = 0.01/7 = 0.0014) was used, but it gives no exact proportions, confidence intervals, z-statistics, or p-values for any of the seven items. The only numerical summary is a range of 66% to 88% and an overall 75%. Figure 4 lacks numeric labels and axis descriptions. As a result, a reader cannot verify the claim that p<0.0014 on all seven items or reproduce the confidence intervals. A table with item-level counts, proportions, Bonferroni-adjusted confidence intervals, and p-values for each group and overall should be provided, and the one- versus two-sided nature of the tests should be stated.
  3. [Sections 3-4] The counterbalancing is incomplete and is weaker than implied. Group 1 performed chat first and CSI second, while Group 2 performed CSI first and chat second; with only one group per order, any systematic difference between the two participant batches is fully confounded with order. Additionally, the first task is always traffic cones and the second always toilet plungers, so task content is not counterbalanced independently of order. Reporting results separately by group and by task would help, and future work should randomize task order at the session level. This does not invalidate the aggregate preference result, but it should be acknowledged when interpreting the causal claim.
  4. [Section 5] The conclusion that 'groups of 75 individuals can successfully brainstorm and prioritize' overreaches the data. The only outcome measures are seven subjective preference items; no objective metrics of brainstorming output (idea count, idea quality, novelty, or convergence quality) are reported. If the paper intends to claim successful brainstorming rather than merely preferred experience, it should include objective content measures or restrict the conclusion to subjective preference.
minor comments (6)
  1. [References] The in-text citation for Cooney et al. gives the year 2020, but the reference list entry states 2023; please align these.
  2. [Section 2] 'Woven into a single conversion' appears to be a typo for 'single conversation.'
  3. [Figure 4] The segmented bar chart needs axis labels, value labels, and a clear legend so that the proportions and confidence intervals are readable without the surrounding text.
  4. [Section 4] The statement that 'None of the confidence intervals overlap the 50% dotted line' should be supported by the numeric confidence intervals in a table, since the figure alone does not provide exact values.
  5. [Section 3] Please clarify what is meant by '99% confidence.' If the authors used Bonferroni-corrected per-test alpha = 0.0014, the confidence level for individual intervals should be about 99.86%, not 99%; if they used 99% intervals for each item, the familywise confidence is not the stated 99%.
  6. [Sections 2-3] The paper does not report sample source details, inclusion/exclusion criteria, informed consent, or an institutional review board statement; these should be included in a methods section or appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the preference result is an empirical comparison against an external control condition (traditional chat), not a quantity derived from its own inputs.

full rationale

The paper's central claim is that participants significantly preferred the CSI structure over a single large chat room on seven subjective survey measures. This is an empirical, externally anchored comparison: the control condition is a standard chat room, and the outcome is self-reported preference obtained after counterbalancing task order across two groups. No parameter is fitted to the outcome and then renamed a prediction; the 66–88% preference proportions are survey measurements, and the statistical tests (Bonferroni-adjusted one-proportion z-tests) test those proportions against 50%, not against any value assumed by the authors. The paper cites prior CSI studies for background and motivation, but the present result does not reduce to those citations; it stands on independently collected survey data. The only potentially load-bearing assumption is the Section 2 assertion that the AI agents 'only passed and received conversational ideas and opinions from other subgroups' and did not introduce AI-generated content. No manipulation check is reported, so this is a legitimate threat to construct validity, but it is not circularity: the claim is not defined in terms of the outcome, and the result would not be true by construction even if the assertion were false. Likewise, the lack of an objective measure of brainstorming output is a limitation about what the self-reported preference can establish, not a circular derivation. No equation, fitted parameter, self-citation chain, or definitional equivalence connects the inputs to the conclusions, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central result (subjective preference for CSI) rests on assumptions about ideal group size, surrogate fidelity, and the validity of self-report as a proxy for brainstorming quality. No free parameters are fitted.

assumptions (3)
  • domain assumption Deliberative conversations are most effective in groups of 4 to 7 members and rapidly lose effectiveness with increasing size.
    Invoked in Section 1 to justify the 5-person subgroup size in the CSI condition; not established within this paper and is taken from external literature.
  • domain assumption LLM-powered Conversational Surrogates can faithfully distill and convey ideas across subgroups without introducing bias or distortion.
    The entire CSI architecture depends on surrogates passing information accurately; no fidelity check or content analysis is provided in this study.
  • domain assumption Subjective preference measures are an adequate indicator of brainstorming quality and productivity.
    The study's only outcome measures are self-reported comparative judgments (Section 3); no objective idea quantity or quality metrics are collected.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large-scale Group Brainstorming using Conversational Swarm Intelligence (CSI) versus Traditional Chat." pith.science (2026). https://pith.science/paper/C4QHI3MS

@misc{pith2026241214205,
  author       = {Pith},
  title        = {Pith review of: Large-scale Group Brainstorming using Conversational Swarm Intelligence (CSI) versus Traditional Chat},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C4QHI3MS}},
  note         = {Machine review of arXiv:2412.14205}
}
read the original abstract

Conversational Swarm Intelligence (CSI) is an AI-facilitated method for enabling real-time conversational deliberations and prioritizations among networked human groups of potentially unlimited size. Based on the biological principle of Swarm Intelligence and modelled on the decision-making dynamics of fish schools, CSI has been shown in prior studies to amplify group intelligence, increase group participation, and facilitate productive collaboration among hundreds of participants at once. It works by dividing a large population into a set of small subgroups that are woven together by real-time AI agents called Conversational Surrogates. The present study focuses on the use of a CSI platform called Thinkscape to enable real-time brainstorming and prioritization among groups of 75 networked users. The study employed a variant of a common brainstorming intervention called an Alternative Use Task (AUT) and was designed to compare through subjective feedback, the experience of participants brainstorming using a CSI structure vs brainstorming in a single large chat room. This comparison revealed that participants significantly preferred brainstorming with the CSI structure and reported that it felt (i) more collaborative, (ii) more productive, and (iii) was better at surfacing quality answers. In addition, participants using the CSI structure reported (iv) feeling more ownership and more buy-in in the final answers the group converged on and (v) reported feeling more heard as compared to brainstorming in a traditional text chat environment. Overall, the results suggest that CSI is a very promising AI-facilitated method for brainstorming and prioritization among large-scale, networked human groups.

Figures

Figures reproduced from arXiv: 2412.14205 by the authors.

Figure 2
Figure 2. Swarm Intelligence enables optimized decisions. CSI technology takes this natural process and emulates the dynamics by breaking large human groups into a network of overlapping subgroups, each with 4 to 7 members, as that size enables optimal conversational deliberation. Unfortunately, there is one more barrier that must be overcome – unlike fish, humans cannot participate effectively in overlapping subgroups (i.e. … view at source ↗
Figure 1
Figure 1. Fish School facing simultaneous threats. In the figure above, three predators approach the school, creating a complex problem of life-or-death significance. Like many human organizations, the members of the school all have limited information. In fact, only three small pockets of fish are aware of any predators (the circled areas above). In fact, the vast majority of fish are unaware of any predator and those in the… view at source ↗
Figure 3
Figure 3. Conversational Swarm Intelligence Architecture An example CSI architecture is shown in [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Subjective Feedback Results with Error Bars. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 3 canonical work pages

  1. [1]

    Enhancing Group Social Perceptiveness through a Swarm-based Decision -Making Platform

    Askay, D., Metcalf, L., Rosenberg, L., Willcox, G. (2019) “Enhancing Group Social Perceptiveness through a Swarm-based Decision -Making Platform.” Proceedings of the 52nd Hawaii International Conference on System Sciences (HICSS-52), IEEE. Blair, Clancy & Gamson, David & Thorne, Steven & Baker, David. (2005). Rising Mean IQ: Cognitive Demand of Mathematic...

  2. [33]

    The Cocktail Party Phenomenon: A Review on Speech Intelligibility in Multiple-Talker Conditions

    93-106. 10.1016/j.intell.2004.07.008. Bronkhorst, Adelbert W. (2000). "The Cocktail Party Phenomenon: A Review on Speech Intelligibility in Multiple-Talker Conditions". Acta Acustica United with Acustica. 86: 117–128. Retrieved 2020-11-16. Cooney, G., et. al. (2023) "The Many Minds Problem: Disclosure in Dyadic vs. Group Conversation." Special Issue on Pr...

  3. [2023]

    https://doi.org/10.1007/978-981-97-0180-3_1 Rosenberg, L.; Schumann, H.; Dishop, C.; Willcox, G.; Woolley, A.; and Mani, G

    New Delhi. https://doi.org/10.1007/978-981-97-0180-3_1 Rosenberg, L.; Schumann, H.; Dishop, C.; Willcox, G.; Woolley, A.; and Mani, G. Conversational Swarms of Humans and AI Agents enable Hybrid Collaborative Decision-making. IEEE UEMCON

  4. [2024]

    and Mani, G

    DOI: 10.1109/UEMCON62879.2024.10754763 Rosenberg, L.; Willcox, G.; Schumann, H. and Mani, G. (2024). Towards Collective Superintelligence: Amplifying Group IQ Using Conversational Swarms. In Proceedings of the 26th International Conference on Enterprise Information Systems - Volume 1: ICEIS; ISBN 978-989-758-692-7; SciTePress, pages 759-766. DOI: 10.5220/...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.