{"id":"7ec8f2a7-a0bd-469c-a088-866930a155f6","arxiv_id":"2506.17073","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM chatbot that posts unmentioned arguments into small online chats increased the number of distinct arguments participants raised, and this held when the bot was disclosed as AI.","lead":"Researchers tested a chatbot that injects missing arguments into online political discussions. In two randomized experiments with 4,397 people, the bot increased the variety of arguments participants mentioned, even when labeled as AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The H1 outcome count likely conflates spontaneous participant argumentation with brief acknowledgments of bot-injected arguments; because the outcome is measured against the same list the bot draws from, the reported increases may be partly mechanical echo.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the primary outcome may be inflated by mechanical echoing of the bot's arguments. The paper is a well-designed pair of preregistered experiments with transparent reporting of null and negative results, and the main effect is statistically significant in both studies. However, the construct validity of the H1 measure is the linchpin: if the count mostly captures participants agreeing with or acknowledging bot-injected arguments, the headline claim that participants themselves express a broader range of arguments is overstated, even though the bot does alter conversational content. The proposed re-annotation check would settle this directly and is feasible once the deposited data are released. I therefore agree with the reader's conditional verdict and see no reason to change it.","tokens_in":26744,"tokens_out":5374,"duration_ms":58753,"concrete_test":"Re-annotate the primary outcome with a stricter rule: count an argument only when the participant's comment independently articulates or elaborates the argument, excluding comments that merely agree with, repeat, or directly respond to a bot message introducing that argument. As a complementary check, re-estimate H1 after dropping all participant comments that occur within two turns after a bot message. If the pooled effects in Studies 1 and 2 remain significant with confidence intervals excluding zero, the mechanical-echo concern does not land. If they attenuate to null, the paper should be reframed: the bot increases exposure to and acknowledgment of arguments, not the range participants themselves express.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on H1: ArgumentBot increases the number of unique arguments 'expressed by participants.' The outcome is measured by removing bot messages and using GPT-4o to detect, in each participant's remaining comments, arguments from the same 40-item list the bot draws on (Materials and Methods). A participant who replies to the bot's 'Have you considered X?' with 'exactly' or 'good point' can be counted as mentioning X even if they never articulate the argument themselves. The supplementary snippet illustrates this: after Alex (Moderator) introduces 'identification of rare symptoms,' Baldwin responds 'exactly ai will have much larger databases' and is plausibly counted for that argument. The paper itself acknowledges 'participants directly adopting the arguments provided by the bot' as a possible mechanism. Thus the main effect may reflect conversational alignment or echoing rather than a genuine broadening of participants' own argumentative repertoire. The group-level increases (from 14.7 to 16.2-17.4 distinct arguments out of 40) are consistent with each bot message generating one or two acknowledgments that get annotated as mentions. The validation of 100 comments at 80 percent agreement does not address this, because the human coders may have applied the same lenient criterion. Because this concern targets the primary outcome, it is load-bearing for the headline claim that the bot 'broadens the range of arguments expressed by participants.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two preregistered randomized experiments (Study 1, N=1,786; Study 2, N=2,611) in which small groups discussed whether AI should be used in healthcare while an LLM-based bot (ArgumentBot) introduced previously missing arguments from a fixed 40-item list at the 2-, 5-, and 8-minute marks. The authors test three preregistered hypotheses: H1 that the bot increases the number of unique arguments mentioned by participants, H2 that it balances participation, and H3 that it improves perceived representativeness. They find a significant pooled effect on the objective H1 measure in both studies (0.244, p<0.001; 0.107, p=0.026), no effect on participation balance, a null or negative effect on perceived representativeness, and no significant difference between AI-labeled and non-AI conditions. The abstract claims the bot 'significantly expands the range of arguments, as measured by both objective and subjective metrics.' The main methodological risk is that the primary outcome is measured by GPT-4o annotations of participant comments against the same argument list the bot draws from, which may count brief acknowledgments of bot-injected arguments as participant mentions. The paper's own supplementary snippet illustrates this pattern, and the 'new arguments seen' subjective item is explicitly described as a manipulation check rather than an independent outcome.","tokens_in":26955,"tokens_out":4381,"duration_ms":44682,"significance":"If the H1 finding is valid, the paper is a valuable empirical contribution: it is one of the first large-scale behavioral experiments showing that an LLM-based moderation tool can increase the diversity of arguments in online discussions, and that transparent disclosure of the bot's AI identity does not eliminate the effect. The study has notable strengths: two preregistered experiments, a realistic synchronous chatroom setting, pre-specified regression analyses, robustness checks with negative binomial and multilevel models, open data and code in a Dataverse repository, and transparent reporting of null and negative results for H2 and H3. However, the central claim rests on the construct validity of the objective H1 measure. If the observed increase in 'unique arguments mentioned' largely reflects participants' brief acknowledgments of bot-injected arguments rather than their spontaneous articulation of a broader argumentative repertoire, the headline conclusion—that the bot broadens the range of arguments participants themselves express—is overstated.","major_comments":[{"comment":"The primary outcome for H1 is operationalized by using GPT-4o to detect, in each participant's comments, arguments from the predefined 40-item list—the same list from which ArgumentBot draws (Materials and Methods). The paper's own potential-mechanisms section acknowledges that 'participants directly adopting the arguments provided by the bot' is one possible explanation, and Supplementary Snippet 1 shows Baldwin replying to the bot's 'Have you considered identification of rare symptoms?' with 'exactly ai will have much larger databases,' a comment that is plausibly annotated as a mention of that argument. Because a participant who merely agrees with or acknowledges a bot-injected argument is counted as 'mentioning' it, the H1 effect may partly be a mechanical consequence of conversational alignment rather than evidence that participants broadened their own argumentative repertoires. The validation of 100 comments at 80% agreement does not rule this out, since human coders may apply the same lenient criterion. This concern targets the central claim of the paper, so it is load-bearing. Please provide a re-analysis that either (a) separates spontaneous first mentions from responses to bot messages (e.g., by excluding comments that directly follow or refer to bot posts, or by annotating whether the participant articulated the argument rather than only affirming it) and reports whether the H1 effect survives for spontaneous mentions, or (b) revises the claims so that they are explicitly about 'engaging with' rather than 'expressing' a broader range of arguments.","section":"Materials and Methods, Outcome measures; Results; Supplementary Snippet 1"},{"comment":"The abstract claims that the bot 'significantly expands the range of arguments, as measured by both objective and subjective metrics.' However, the only subjective measure that is significant is 'New arguments seen,' which the authors themselves describe as functioning 'as a form of manipulation check' (Results). The direct subjective measure of the range of arguments, 'Range of viewpoints seen,' is null in both studies (Study 1 pooled effect 0.037, p=0.46; Study 2 pooled effect 0.026, p=0.59, per Figure S1 and Tables S3/S10). Using an item that is explicitly framed as a manipulation check as evidence for the substantive broadening claim is circular, especially when the more face-valid subjective range measure showed no effect. The abstract and Discussion should be revised to restrict the subjective-evidence claim to the 'new arguments seen' item with the caveat that it is a manipulation check, or to acknowledge that the subjective range measure did not replicate.","section":"Abstract; Results, 'ArgumentBot increases the number of arguments'"},{"comment":"The H1 effect is inconsistent across roles: in Study 2, the 'Participant' condition showed no significant effect (-0.011, p=0.859) and the 'AI Participant' condition showed a non-significant positive trend (0.060, p=0.328). The pooled Study 2 effect (0.107) is therefore driven by the moderator conditions, particularly 'Moderator' (0.231, p=0.0002). While the paper reports these estimates, the general framing—in the title, abstract, and discussion—treats the bot as effective regardless of role. The moderation by role is not a fatal flaw, but it should be integrated into the headline message: the bot broadened argumentation when framed as a moderator, whereas the participant-role effect did not replicate. This qualification matters for the practical recommendation that such bots be deployed as moderators.","section":"Results, 'ArgumentBot increases the number of arguments'"}],"minor_comments":[{"comment":"There is a typo 'Mooreover' in the paragraph describing the hypotheses; it should be 'Moreover.'","section":"Introduction"},{"comment":"The description of the GPT-4o annotation says it identifies arguments 'explicitly mentioned in a comment,' but the main text later says H1 measures arguments participants 'engage with, either by introducing them or responding to them.' These are different standards; please clarify which one is used and how the annotation prompt operationalizes 'explicitly mentioned.'","section":"Materials and Methods, 'Outcome measures'"},{"comment":"The preregistrations are mentioned as 'pre-registered,' but no preregistration identifiers or OSF links are provided in the main text or the data availability statement. Please add links for both studies so readers can verify that the reported analyses match the pre-analysis plans.","section":"Data and Materials Availability"},{"comment":"There are minor presentation issues: Table S20 has the typo 'Ep. Political Discsussions,' and in Table S18 (Study 2, Part 2) the R2 Cond. value for 'Different backgrounds' appears to be missing. Please fix these in the supplementary materials.","section":"Supplementary Tables S20 and S18"},{"comment":"The caption says 'standardized effect sizes' but does not state that the dependent variables were z-scored; please add a sentence noting that the outcome variables were standardized before regression, so the coefficients are interpretable in standard deviation units.","section":"Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is well-designed and clearly reported, but the central construct-validity concern about the objective H1 measure is serious and directly affects the headline claim. If the authors can provide a robustness analysis that separates spontaneous mentions from acknowledgments of bot-injected arguments and shows the effect persists, I would view the paper as a strong accept candidate. In the current form, the claims are somewhat broader than the evidence supports, particularly the abstract's 'both objective and subjective metrics' phrasing. I also recommend that the editor request the preregistration links, since they are not currently in the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThis paper is worth taking seriously. As a test of whether an LLM bot that injects missing arguments into small online discussions can widen the set of arguments participants engage with, it is well designed: two pre-registered randomized experiments (N≈1,786 and 2,611), a transparent monitor-and-inject procedure, and honest reporting of null and negative results (no effect on participation balance, a negative effect on perceived representativeness). The pooled H1 effects (0.244 in Study 1, 0.107 in Study 2) are credible, and the group-level increases of roughly 8–13% in distinct arguments are a real behavioral signal. The disclosure-null is consistent with earlier work and useful for the AI Act debate.\n\nThe main soft spot is exactly where the stress-test points. The outcome—number of unique arguments each participant \"mentioned\"—is measured with GPT-4o using the same 40-item list the bot draws on, and the bot messages are removed before annotation. A participant who replies to \"Have you considered X?\" with \"exactly\" or \"good point\" can be counted as mentioning X. The supplementary snippet with Baldwin shows this happening. The paper itself says participants may be \"directly adopting\" bot arguments. So part of the H1 effect is likely mechanical echo, not a broadening of participants' own argumentative repertoire. That matters because the title and abstract claim the bot \"broadens the range of arguments expressed by participants.\" It broadens the range of arguments engaged with, which H1 defines, but the stronger phrasing overstates it.\n\nThis is not fatal: the bot did change conversational content in a measurable way, and the effect persists across two samples. But the authors should report a sensitivity analysis that excludes simple acknowledgments or re-words the claim. Secondary soft spots: the \"participant\" role failed to replicate in Study 2, the subjective \"range of viewpoints\" item was null, and the disclosure finding is an absence of significant difference rather than demonstrated equivalence. Data and code are deposited but not yet available, which is normal.\n\nWho it's for: CSCW, online deliberation, and AI moderation researchers, plus people working on the EU AI Act transparency requirements. It deserves a serious referee. I would send it to review with the measurement issue flagged for revision, not desk-reject it.","headline":"Solid pre-registered experiments, but the main outcome partly measures echoing of bot-injected arguments rather than just broadening of participants' own argumentation.","tokens_in":27532,"tokens_out":2762,"would_cite":false,"duration_ms":26978,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An LLM bot that injects missing arguments into online chats increases the number of distinct arguments participants voice, and disclosing it as AI does not remove the effect.","keywords":["LLM moderation","argument diversity","online deliberation","randomized experiment","AI disclosure","chatbot intervention","political discussion","GPT-4o annotation"],"falsifier":"Re-run the outcome annotation with all comments double-coded by human annotators blind to condition, and compute the unique-argument count separately for arguments that appear before any bot message and arguments that first appear after the bot's prompt. If the treatment effect collapses when only pre-bot or spontaneous mentions are counted, the claim that the bot broadens participants' own argumentation is falsified; it would instead show that participants echo the bot.","tokens_in":1664,"feed_emoji":"🤖","tokens_out":6218,"duration_ms":97859,"temperature":0.7,"pith_summary":"This paper asks whether an LLM-powered bot can broaden the range of arguments in short online political discussions, and whether telling participants the bot is AI destroys the effect. Across two pre-registered randomized chatroom experiments (N=1,786 and N=2,611), a bot that monitors the conversation, finds arguments from a curated list that no one has mentioned, and posts \"Have you considered X?\" increased the number of distinct arguments participants themselves mentioned, with pooled standardized effects of 0.244 in Study 1 and 0.107 in Study 2, and roughly 8 to 13 percent more distinct arguments at the group level. Disclosing the bot as an AI participant or AI moderator did not significantly change this outcome. The same intervention did not make participation more balanced and, in several conditions, lowered perceived representativeness, so the practical message is that AI moderation can enrich the argument pool even though it does not automatically improve the felt quality of discussion.","feed_headline":"Bot that posts missing arguments widens online chats, even labeled AI","feed_subtitle":"In two chat experiments, an LLM moderator added 8-13% more arguments per group; AI labels did not erase the gain.","key_machinery":"The load-bearing device is ArgumentBot, an LLM-driven moderator that at 2, 5, and 8 minutes compares the ongoing chat log against a list of 40 arguments about AI in healthcare (compiled from 66 expert responses and coded by two co-authors), picks an argument not yet mentioned, and posts it as \"Have you considered [argument]?\" with a one-line explanation, without asking anyone to reply. The outcome measure is symmetric with the intervention: after stripping bot messages, GPT-4o scans each participant's comments and counts how many of the same 40 listed arguments they mention, with a validation of 100 comments reaching 80 percent agreement with human coders. The bot's role label (regular participant, moderator, AI participant, AI moderator) is the experimental manipulation around which all comparisons are organized.","core_discovery":"The central claim is that a transparently disclosed LLM bot can expand the range of arguments expressed by human participants in a live online discussion. The paper operationalizes \"range\" as the number of unique arguments, from a fixed 40-item expert-compiled list, that each participant mentions after bot messages are removed. The effect is robust in both experiments when the bot is cast as a moderator (Study 1: 0.332, p<0.001; Study 2: 0.231, p<0.001) and in the pooled treatment effect (Study 1: 0.244, p<0.001; Study 2: 0.107, p=0.026). The authors also report that labeling the bot as AI, as an \"AI moderator\" or \"AI participant,\" did not erase these gains, and that participants in all treatment conditions reported encountering more new arguments. They interpret this as evidence that LLM-based moderation can counter the homogeneity of online political discussion without being undermined by transparency requirements.","pith_inferences":["The measurement design cannot cleanly separate adoption from echoing: a participant who responds \"exactly, AI has larger databases\" to a bot's prompt is counted as mentioning that argument, so the H1 effect may partly reflect the bot seeding its own uptake rather than participants independently generating new perspectives.","Because the argument list was compiled from mainstream experts and posted by an authority-labeled bot, the same intervention in a polarized or adversarial setting might trigger reactance instead of broadening; this is a testable extension beyond the open-topic chatroom used here.","A natural next experiment would compare an echo count (arguments mentioned only after the bot raised them) against a spontaneous count (arguments raised before any bot prompt) to estimate how much of the 8 to 13 percent gain is genuine expansion of participants' own repertoire.","If the goal is deliberative quality, these results suggest argument-count metrics and perceived-legitimacy metrics can move in opposite directions, so platform designers should not treat \"more arguments\" as equivalent to \"better discussion.\""],"forward_implications":["When ArgumentBot acted as moderator, groups produced roughly 8 to 13 percent more distinct arguments than control groups (Study 1: 14.7 vs. 16.2; Study 2: 15.4 vs. 17.4 in the moderator condition).","Adding the word \"AI\" to the bot's label produced no statistically significant difference in the number of unique arguments, so transparency requirements need not cancel the intervention's effect.","Participants in every treatment condition reported seeing new arguments more often than controls did, even when objective counts in a given condition, such as bot as participant in Study 2, did not rise.","The bot did not change how evenly comments were distributed across participants, and it reduced perceived representativeness in several conditions, so broadening the argument pool and improving deliberative experience are separable.","The strongest and most consistent effects came from the moderator role, not the participant role, and a bot presented as a plain participant failed to replicate in Study 2."],"supporting_citations":[{"why":"Prior evidence that conversational agents can change who contributes in online debates; motivates the bot intervention.","marker":"[25]"},{"why":"Shows AI chat assistants in online political conversations can raise perceived quality and reciprocity; the direct precedent this study extends.","marker":"[26]"},{"why":"Establishes that AI tools can help deliberating groups find common ground; used as a springboard for the moderator design.","marker":"[27]"},{"why":"Provides the hypothetical \"AI penalty\" result that the paper's behavioral data directly challenge.","marker":"[31]"},{"why":"Is the chatroom platform on which both experiments were run; supplies the experimental infrastructure.","marker":"[32]"}],"fun_headline_variants":["LLM bot widens online arguments even when labeled AI","Disclosed AI bot broadens discussion range","Bot posts missing arguments, expands chats despite AI tag","AI moderator adds views, transparency doesn't reduce gain","LLM moderator increases argument diversity, even if disclosed"],"cache_read_input_tokens":29568,"weakest_assumption_plain":"The headline result rests on counting as \"participant-mentioned\" any argument from the bot's own 40-item list that GPT-4o finds in a participant's comment after bot messages are removed; if participants mainly echo or agree with bot-posted arguments, the measured broadening may reflect the bot's own prompts rather than participants spontaneously adopting new perspectives.","fun_headline_variants_meta":{"raw":{"variants":["LLM bot widens online arguments even when labeled AI","Disclosed AI bot broadens discussion range","Bot posts missing arguments, expands chats despite AI tag","AI moderator adds views, transparency doesn't reduce gain","LLM moderator increases argument diversity, even if disclosed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1375,"prompt_tokens":915,"completion_tokens":460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":385}},"tokens_in":531,"tokens_out":460,"duration_ms":4957,"temperature":1.0,"reasoning_tokens":385,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:12:51.827306+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the outcome annotation with all comments double-coded by human annotators blind to condition, and compute the unique-argument count separately for arguments that appear before any bot message and arguments that first appear after the bot's prompt. If the treatment effect collapses when only pre-bot or spontaneous mentions are counted, the claim that the bot broadens participants' own argumentation is falsified; it would instead show that participants echo the bot.","supporting_citations":[{"cited_title":"Have you considered [selected_missing_argument]?","cited_arxiv_id":null,"evidence_quote":"Is the chatroom platform on which both experiments were run; supplies the experimental infrastructure."}],"review_version":1}