REVIEW 3 major objections 5 minor 32 references
LLM-Based Bot Broadens the Range of Arguments in Online Discussions, Even When Transparently Disclosed as AI
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read An LLM bot that injects missing arguments into online chats increases the number of distinct arguments participants voice, and disclosing it as AI does not remove the effect.
desk verdict Solid pre-registered experiments, but the main outcome partly measures echoing of bot-injected arguments rather than just broadening of participants' own argumentation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is ArgumentBot, an LLM-driven moderator that at 2, 5, and 8 minutes compares the ongoing chat log against a list of 40 arguments about AI in healthcare (compiled from 66 expert responses and coded by two co-authors), picks an argument not yet mentioned, and posts it as "Have you considered [argument]?" with a one-line explanation, without asking anyone to reply. The outcome measure is symmetric with the intervention: after stripping bot messages, GPT-4o scans each participant's comments and counts how many of the same 40 listed arguments they mention, with a validation of 100 comments reaching 80 percent agreement with human coders. The bot's role label (regular participant, moderator, AI participant, AI moderator) is the experimental manipulation around which all comparisons are organized.
What would settle it
Re-run the outcome annotation with all comments double-coded by human annotators blind to condition, and compute the unique-argument count separately for arguments that appear before any bot message and arguments that first appear after the bot's prompt. If the treatment effect collapses when only pre-bot or spontaneous mentions are counted, the claim that the bot broadens participants' own argumentation is falsified; it would instead show that participants echo the bot.
Extended reading notes
Core claim
The central claim is that a transparently disclosed LLM bot can expand the range of arguments expressed by human participants in a live online discussion. The paper operationalizes "range" as the number of unique arguments, from a fixed 40-item expert-compiled list, that each participant mentions after bot messages are removed. The effect is robust in both experiments when the bot is cast as a moderator (Study 1: 0.332, p<0.001; Study 2: 0.231, p<0.001) and in the pooled treatment effect (Study 1: 0.244, p<0.001; Study 2: 0.107, p=0.026). The authors also report that labeling the bot as AI, as an "AI moderator" or "AI participant," did not erase these gains, and that participants in all treatment conditions reported encountering more new arguments. They interpret this as evidence that LLM-based moderation can counter the homogeneity of online political discussion without being undermined by transparency requirements.
Load-bearing premise
The headline result rests on counting as "participant-mentioned" any argument from the bot's own 40-item list that GPT-4o finds in a participant's comment after bot messages are removed; if participants mainly echo or agree with bot-posted arguments, the measured broadening may reflect the bot's own prompts rather than participants spontaneously adopting new perspectives.
Editorial extensions
If this is right
- When ArgumentBot acted as moderator, groups produced roughly 8 to 13 percent more distinct arguments than control groups (Study 1: 14.7 vs. 16.2; Study 2: 15.4 vs. 17.4 in the moderator condition).
- Adding the word "AI" to the bot's label produced no statistically significant difference in the number of unique arguments, so transparency requirements need not cancel the intervention's effect.
- Participants in every treatment condition reported seeing new arguments more often than controls did, even when objective counts in a given condition, such as bot as participant in Study 2, did not rise.
- The bot did not change how evenly comments were distributed across participants, and it reduced perceived representativeness in several conditions, so broadening the argument pool and improving deliberative experience are separable.
- The strongest and most consistent effects came from the moderator role, not the participant role, and a bot presented as a plain participant failed to replicate in Study 2.
Reading between the lines
- The measurement design cannot cleanly separate adoption from echoing: a participant who responds "exactly, AI has larger databases" to a bot's prompt is counted as mentioning that argument, so the H1 effect may partly reflect the bot seeding its own uptake rather than participants independently generating new perspectives.
- Because the argument list was compiled from mainstream experts and posted by an authority-labeled bot, the same intervention in a polarized or adversarial setting might trigger reactance instead of broadening; this is a testable extension beyond the open-topic chatroom used here.
- A natural next experiment would compare an echo count (arguments mentioned only after the bot raised them) against a spontaneous count (arguments raised before any bot prompt) to estimate how much of the 8 to 13 percent gain is genuine expansion of participants' own repertoire.
- If the goal is deliberative quality, these results suggest argument-count metrics and perceived-legitimacy metrics can move in opposite directions, so platform designers should not treat "more arguments" as equivalent to "better discussion."
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports two preregistered randomized experiments (Study 1, N=1,786; Study 2, N=2,611) in which small groups discussed whether AI should be used in healthcare while an LLM-based bot (ArgumentBot) introduced previously missing arguments from a fixed 40-item list at the 2-, 5-, and 8-minute marks. The authors test three preregistered hypotheses: H1 that the bot increases the number of unique arguments mentioned by participants, H2 that it balances participation, and H3 that it improves perceived representativeness. They find a significant pooled effect on the objective H1 measure in both studies (0.244, p<0.001; 0.107, p=0.026), no effect on participation balance, a null or negative effect on perceived representativeness, and no significant difference between AI-labeled and non-AI conditions. The abstract claims the bot 'significantly expands the range of arguments, as measured by both objective and subjective metrics.' The main methodological risk is that the primary outcome is measured by GPT-4o annotations of participant comments against the same argument list the bot draws from, which may count brief acknowledgments of bot-injected arguments as participant mentions. The paper's own supplementary snippet illustrates this pattern, and the 'new arguments seen' subjective item is explicitly described as a manipulation check rather than an independent outcome.
Significance. If the H1 finding is valid, the paper is a valuable empirical contribution: it is one of the first large-scale behavioral experiments showing that an LLM-based moderation tool can increase the diversity of arguments in online discussions, and that transparent disclosure of the bot's AI identity does not eliminate the effect. The study has notable strengths: two preregistered experiments, a realistic synchronous chatroom setting, pre-specified regression analyses, robustness checks with negative binomial and multilevel models, open data and code in a Dataverse repository, and transparent reporting of null and negative results for H2 and H3. However, the central claim rests on the construct validity of the objective H1 measure. If the observed increase in 'unique arguments mentioned' largely reflects participants' brief acknowledgments of bot-injected arguments rather than their spontaneous articulation of a broader argumentative repertoire, the headline conclusion—that the bot broadens the range of arguments participants themselves express—is overstated.
major comments (3)
- [Materials and Methods, Outcome measures; Results; Supplementary Snippet 1] The primary outcome for H1 is operationalized by using GPT-4o to detect, in each participant's comments, arguments from the predefined 40-item list—the same list from which ArgumentBot draws (Materials and Methods). The paper's own potential-mechanisms section acknowledges that 'participants directly adopting the arguments provided by the bot' is one possible explanation, and Supplementary Snippet 1 shows Baldwin replying to the bot's 'Have you considered identification of rare symptoms?' with 'exactly ai will have much larger databases,' a comment that is plausibly annotated as a mention of that argument. Because a participant who merely agrees with or acknowledges a bot-injected argument is counted as 'mentioning' it, the H1 effect may partly be a mechanical consequence of conversational alignment rather than evidence that participants broadened their own argumentative repertoires. The validation of 100 comments at 80% agreement does not rule this out, since human coders may apply the same lenient criterion. This concern targets the central claim of the paper, so it is load-bearing. Please provide a re-analysis that either (a) separates spontaneous first mentions from responses to bot messages (e.g., by excluding comments that directly follow or refer to bot posts, or by annotating whether the participant articulated the argument rather than only affirming it) and reports whether the H1 effect survives for spontaneous mentions, or (b) revises the claims so that they are explicitly about 'engaging with' rather than 'expressing' a broader range of arguments.
- [Abstract; Results, 'ArgumentBot increases the number of arguments'] The abstract claims that the bot 'significantly expands the range of arguments, as measured by both objective and subjective metrics.' However, the only subjective measure that is significant is 'New arguments seen,' which the authors themselves describe as functioning 'as a form of manipulation check' (Results). The direct subjective measure of the range of arguments, 'Range of viewpoints seen,' is null in both studies (Study 1 pooled effect 0.037, p=0.46; Study 2 pooled effect 0.026, p=0.59, per Figure S1 and Tables S3/S10). Using an item that is explicitly framed as a manipulation check as evidence for the substantive broadening claim is circular, especially when the more face-valid subjective range measure showed no effect. The abstract and Discussion should be revised to restrict the subjective-evidence claim to the 'new arguments seen' item with the caveat that it is a manipulation check, or to acknowledge that the subjective range measure did not replicate.
- [Results, 'ArgumentBot increases the number of arguments'] The H1 effect is inconsistent across roles: in Study 2, the 'Participant' condition showed no significant effect (-0.011, p=0.859) and the 'AI Participant' condition showed a non-significant positive trend (0.060, p=0.328). The pooled Study 2 effect (0.107) is therefore driven by the moderator conditions, particularly 'Moderator' (0.231, p=0.0002). While the paper reports these estimates, the general framing—in the title, abstract, and discussion—treats the bot as effective regardless of role. The moderation by role is not a fatal flaw, but it should be integrated into the headline message: the bot broadened argumentation when framed as a moderator, whereas the participant-role effect did not replicate. This qualification matters for the practical recommendation that such bots be deployed as moderators.
minor comments (5)
- [Introduction] There is a typo 'Mooreover' in the paragraph describing the hypotheses; it should be 'Moreover.'
- [Materials and Methods, 'Outcome measures'] The description of the GPT-4o annotation says it identifies arguments 'explicitly mentioned in a comment,' but the main text later says H1 measures arguments participants 'engage with, either by introducing them or responding to them.' These are different standards; please clarify which one is used and how the annotation prompt operationalizes 'explicitly mentioned.'
- [Data and Materials Availability] The preregistrations are mentioned as 'pre-registered,' but no preregistration identifiers or OSF links are provided in the main text or the data availability statement. Please add links for both studies so readers can verify that the reported analyses match the pre-analysis plans.
- [Supplementary Tables S20 and S18] There are minor presentation issues: Table S20 has the typo 'Ep. Political Discsussions,' and in Table S18 (Study 2, Part 2) the R2 Cond. value for 'Different backgrounds' appears to be missing. Please fix these in the supplementary materials.
- [Figure 3 caption] The caption says 'standardized effect sizes' but does not state that the dependent variables were z-scored; please add a sentence noting that the outcome variables were standardized before regression, so the coefficients are interpretable in standard deviation units.
Circularity Check
No circularity: the H1 claim is an empirical causal estimate, and shared use of the argument list is a construct-validity issue, not a reduction of the prediction to its inputs.
full rationale
The paper's central claim (H1) is an empirical causal estimate from two preregistered randomized experiments: ArgumentBot posts messages containing missing arguments from a precompiled list, and the outcome counts how many of those arguments participants subsequently mention in their own comments after bot messages are removed. This is not circular: the effect is not equated to the bot's messages by definition, the outcome is not a fitted parameter, and the paper does not invoke any load-bearing self-citation. The fact that the intervention and outcome share the same argument list is a design feature that could inflate the effect through echoing or direct adoption, a construct-validity threat the authors themselves acknowledge when they write that the effect 'may be due to participants directly adopting the arguments provided by the bot.' But this is not a logical reduction of the prediction to its inputs. The subjective 'new arguments seen' measure is explicitly described as a manipulation check, not as a predicted outcome derived from the manipulation. No circular step can be quoted: experts independently compiled the argument list, randomization identifies the causal effect, and GPT-4o annotation was validated against human coding with 80 percent agreement. The derivation chain is therefore self-contained, and the concern raised in the reader's take is best framed as measurement validity, not circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption The 40-item argument list compiled from 66 expert responses is a comprehensive and unbiased representation of relevant arguments on AI in healthcare for the participant population.
- domain assumption GPT-4o accurately maps participant comments to arguments from the list after bot messages are removed.
- domain assumption Random assignment at the group level, with controls for group size and demographics, yields exchangeable treatment and control groups.
- domain assumption The three-item survey composite measures perceived representativeness as intended.
- domain assumption Behavior in short, paid, synchronous chatrooms generalizes to online political discussions.
Cite this review
Pith. "Pith review of LLM-Based Bot Broadens the Range of Arguments in Online Discussions, Even When Transparently Disclosed as AI." pith.science (2026). https://pith.science/paper/UITPBH2O
@misc{pith2026250617073,
author = {Pith},
title = {Pith review of: LLM-Based Bot Broadens the Range of Arguments in Online Discussions, Even When Transparently Disclosed as AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/UITPBH2O}},
note = {Machine review of arXiv:2506.17073}
}
read the original abstract
A wide range of participation is essential for democracy, as it helps prevent the dominance of extreme views, erosion of legitimacy, and political polarization. However, engagement in online political discussions often features a limited spectrum of views due to high levels of self-selection and the tendency of online platforms to facilitate exchanges primarily among like-minded individuals. This study examines whether an LLM-based bot can widen the scope of perspectives expressed by participants in online discussions through two pre-registered randomized experiments conducted in a chatroom. We evaluate the impact of a bot that actively monitors discussions, identifies missing arguments, and introduces them into the conversation. The results indicate that our bot significantly expands the range of arguments, as measured by both objective and subjective metrics. Furthermore, disclosure of the bot as AI does not significantly alter these effects. These findings suggest that LLM-based moderation tools can positively influence online political discourse.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
H. Werner, S. Marien, Process vs. Outcome? How to Evaluate the Effects of Participatory Processes on Legitimacy Perceptions.British Journal of Political Science52(1), 429–436 (2022), doi:10.1017/S0007123420000459
-
[3]
S. Goldberg, Just Advisory and Maximally Representative: A Conjoint Experiment on Non- Participants’ Legitimacy Perceptions of Deliberative Forums.Journal of Deliberative Democ- racy17(1) (2021), doi:10.16997/jdd.973
- [4]
-
[5]
E. Hargittai, K. Jennrich, The Online Participation Divide, inThe Communication Crisis in America, And How to Fix It, M. Lloyd, L. A. Friedland, Eds. (Palgrave Macmillan US), pp. 199–213 (2016), doi:10.1057/978-1-349-94925-0_13
-
[6]
J. Oser, A. Grinson, S. Boulianne, E. Halperin, How Political Efficacy Relates to Online and Offline Political Participation: A Multilevel Meta-analysis.Political Communication39(5), 607–633 (2022), doi:10.1080/10584609.2022.2086329
arXiv 2022
-
[7]
B. Rottinghaus, T. Escher, Mechanisms for inclusion and exclusion through digital political participation: Evidence from a comparative study of online consultations in three German cities. Zeitschrift für Politikwissenschaft30(2), 261–298 (2020), doi:10.1007/s41358-020-00222-7
-
[8]
Y . M. Baek, M. Wojcieszak, M. X. Delli Carpini, Online versus face-to-face deliberation: Who? Why? What? With what effects?New Media & Society14(3), 363–383 (2012), doi:10.1177/1461444811413191
Show all 32 references
-
[9]
Oswald, W
L. Oswald, W. S. Schulz, P. Lorenz-Spreen, A Collective Field Experiment Disentangling Participation in Online Political Discussions (1. April 2025), doi:10.31235/osf.io/p2jaq_v1
2025 doi
-
[10]
K. L. Schlozman, H. E. Brady, S. Verba,Unequal and Unrepresented(Princeton University Press) (2018)
2018
-
[11]
Griffin, T
J. Griffin, T. Abdel-Monem, A. Tomkins, A. Richardson, S. Jorgensen, Understanding Partic- ipant Representativeness in Deliberative Events: A Case Study Comparing Probability and Non-Probability Recruitment Strategies.Journal of Public Deliberation11(1) (2015)
2015
-
[12]
Jacquet, Explaining non-participation in deliberative mini-publics.European Journal of Political Research56(3), 640–659 (2017), doi:10.1111/1475-6765.12195
V . Jacquet, Explaining non-participation in deliberative mini-publics.European Journal of Political Research56(3), 640–659 (2017), doi:10.1111/1475-6765.12195
2017
-
[13]
C. F. Karpowitz, T. Mendelberg,The silent sex(Princeton University Press) (2014)
2014
-
[14]
C. F. Karpowitz, T. Mendelberg, L. Shaker, Gender Inequality in Deliberative Participation. American Political Science Review106(3), 533–547 (2012), doi:10.1017/S0003055412000329
2012 doi
-
[15]
C. R. Sunstein,Republic.com(Princeton University Press) (2001)
2001
-
[16]
C. R. Sunstein,#Republic(Princeton University Press) (2017)
2017
-
[17]
Pariser,The Filter Bubble(Penguin Group) (2011)
E. Pariser,The Filter Bubble(Penguin Group) (2011). 13 LLM-Based Bot Broadens the Range of Arguments in Online Discussions
2011
-
[18]
Yarchi, C
M. Yarchi, C. Baden, N. Kligler-Vilenchik, Political Polarization on the Digital Sphere: A Cross-platform, Over-time Analysis of Interactional, Positional, and Affec- tive Polarization on Social Media.Political Communication38(1-2), 98–139 (2021), doi:10.1080/10584609.2020.1785067
2021
-
[19]
M. Wojcieszak, ‘Don’t talk to me’: effects of ideologically homogeneous online groups and politically dissimilar offline ties on extremism.New Media & Society12(4), 637–655 (2010), doi:10.1177/1461444809342775
2010 doi
-
[20]
S. B. Hobolt, K. Lawall, J. Tilley, The Polarizing Effect of Partisan Echo Chambers.American Political Science Review118(3), 1464–1479 (2024), doi:10.1017/S0003055423001211
2024 doi
-
[21]
D. M. Ryfe, B. Stalsburg, The Participation and Recruitment Challenge, inDemocracy in Motion, T. Nabatchi, J. Gastil, M. Leighninger, G. M. Weiksner, Eds., pp. 43–58 (2012), doi:10.1093/acprof:oso/9780199899265.001.0001
2012
-
[22]
S. Kim, J. Eun, J. Seering, J. Lee, Moderator Chatbot for Deliberative Discussion.Proceedings of the ACM on Human-Computer Interaction5(CSCW1), 1–26 (2021), doi:10.1145/3449161
2021 doi
-
[23]
J. S. Fishkin,et al., Deliberative Democracy with the Online Deliberation Platform (2019)
2019
-
[24]
Goldberg,et al., AI and the Future of Digital Public Squares.arXiv(2024), https://arxiv.org/ abs/2412.09988
B. Goldberg,et al., AI and the Future of Digital Public Squares.arXiv(2024), https://arxiv.org/ abs/2412.09988
2024 arXiv
-
[25]
Hadfi,et al., Conversational agents enhance women’s contribution in online debates.Scientific reports13(1), 14534 (2023), doi:10.1038/s41598-023-41703-3
R. Hadfi,et al., Conversational agents enhance women’s contribution in online debates.Scientific reports13(1), 14534 (2023), doi:10.1038/s41598-023-41703-3
2023 doi
-
[26]
L. P. Argyle,et al., Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale.Proceedings of the National Academy of Sciences of the United States of America120(41), e2311627120 (2023), doi:10.1073/pnas.2311627120
2023 doi
-
[27]
M. H. Tessler,et al., AI can help humans find common ground in democratic deliberation. Science (New York, N.Y.)386(6719), eadq2852 (2024), doi:10.1126/science.adq2852
2024 doi
-
[28]
Noelle-Neumann, The Spiral of Silence a Theory of Public Opinion.Journal of Communica- tion24(2), 43–51 (1974), doi:10.1111/j.1460-2466.1974.tb00367.x
E. Noelle-Neumann, The Spiral of Silence a Theory of Public Opinion.Journal of Communica- tion24(2), 43–51 (1974), doi:10.1111/j.1460-2466.1974.tb00367.x
1974
-
[29]
S. E. Asch, Effects of Group Pressure upon the Modification and Distortion of Judgments, in Groups, Leadership and Men: Research in Human Relations, H. Guetzkow, Ed. (Carnegie Press), pp. 177–190 (1951)
1951
-
[30]
D. E. Broockman, J. L. Kalla, N. Ottone, E. Santoro, A. Weiss, Shared Demographic Char- acteristics Do Not Reliably Facilitate Persuasion in Interpersonal Conversations: Evidence from Eight Experiments.British Journal of Political Science54(4), 1477–1485 (2024), doi:10.1017/S0...
2024 doi
-
[31]
Jungherr, A
A. Jungherr, A. Rauchfleisch, Artificial Intelligence in Deliberation: The AI Penalty and the Emergence of a New Deliberative Divide (10.03.2025), https://arxiv.org/pdf/2503.07690
2025
-
[32]
Have you considered [selected_missing_argument]?
E. Hoes, K. J. Klüser, Tool for Online Discussions (TOD): A New Experimental Re- search Platform to Examine Live User Interactions on Social Media.PsyArXiv(2023), doi:10.31234/osf.io/exk2u. 14 LLM-Based Bot Broadens the Range of Arguments in Online Discussions Acknowledgments ...
2023 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.