{"id":"7f023e3e-f9cb-48ef-9664-788434f6a335","arxiv_id":"2506.14295","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Four AI assistance tools in a simulated social media discussion increased engagement and output but reduced perceived quality and authenticity compared to a no-AI control.","lead":"A controlled experiment with 680 US participants found that AI writing aids in online discussion groups increase participation and comment length, but lower perceived quality and authenticity of the conversation. The results highlight a trade-off for social media platforms rather than a clear benefit of AI assistance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'negative spill-over effect on conversations' is not identifiable: whole groups are assigned to a single condition, so there are no non-AI users inside treatment groups whose conversation quality could be compared.","rationale":"I focused on the spillover component because the abstract states it as a finding and the discussion section repeats it as a mechanism ('even among users not using the AI themselves'). Under the paper's own Methods, treatment is randomized at the group level, with no within-group variation in AI access. A treatment-vs-control difference measures the total effect of putting the tool in a group, not an effect that propagates from AI users to non-users. Because no such non-user exists in the control condition's counterpart, the spillover claim requires either a mixed design or a within-thread before/after analysis that the paper does not report. I am not raising this as a generalizability complaint; the external-validity concern identified by the Reader is real but secondary, and the paper is appropriately cautious in its Limitations. The multiple-comparisons/clustering issue is also real and worth an audit, but it affects the strength of the secondary effects, not the identifiability of the headlined claim. The experiment itself is carefully conducted, has a reasonable sample size, and reports many outcome variables transparently; those parts deserve credit. My recommendation is to keep the Reader's CONDITIONAL verdict: the core trade-off findings can stand, but the abstract and discussion should either drop 'negative spill-over' or explicitly downgrade it to a hypothesis pending a suitable analysis. Hence UNCHANGED relative to the Reader's verdict.","tokens_in":29533,"tokens_out":6089,"duration_ms":63738,"concrete_test":"Using the logged interaction data, construct the contrast required by the claim: compare the perceived quality/ratings of comments and replies within treatment groups from participants who never used the AI tool (or during rounds where they did not use it) against the control group, with a model that includes topic, comment depth, time remaining, and participant random effects and clusters standard errors by group. If this contrast is not significant, the 'negative spill-over effect' claim is unsupported and should be removed from the abstract. If the data do not contain a usable non-user contrast, the only definitive test is a follow-up mixed-group experiment with randomly assigned AI access within groups.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim in the abstract—'introduce a negative spill-over effect on conversations'—and the 'How to move forward' section's stronger statement that AI lowers 'the quality of subsequent conversations within threads, even among users not using the AI themselves' cannot be supported by the experimental design. The Methods state that 'Each group was randomly assigned to one of the five experimental conditions' and that participants interacted only within their assigned group. This is cluster randomization at the group level: in every treatment group all five participants have access to the same AI tool, and in the control group nobody does. There is no condition in which some group members have AI access and others do not, so the design cannot estimate a spillover effect onto non-AI users. The only 'non-users' available in the data are treatment participants who voluntarily did not click the tool; they are a self-selected subset, and comparisons involving them are confounded by motivation, engagement, and topic interest. The measured treatment-vs-control contrasts (longer comments, lower perceived quality, more dislikes) are valid average effects of the tool on whole groups, but the spillover sentence is an interpretive leap beyond those contrasts. It should be removed from the abstract, or explicitly framed as a hypothesis, unless a dedicated analysis of conversation quality before/after AI-assisted comments within treatments is provided.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subjects controlled experiment on a custom social-media-like platform in which 680 U.S. participants in 136 five-person groups were randomly assigned to a control condition or one of four GPT-4o-based AI assistance tools (Chat, Conversation Starter, Feedback, Suggestions). Each group held three 10-minute discussions on topics differing in sensitivity. The paper measures producer-side outcomes (comment length, participation equality, self-reported willingness to participate) and consumer-side outcomes (perceived comment quality, ratings of replies, reaction types), together with detailed AI-usage analyses and post-study questionnaires. It concludes that some AI tools increase engagement and content volume while decreasing perceived quality and authenticity, and that they introduce a negative spillover effect on conversations; it then proposes four design principles for AI deployment on social media.","tokens_in":29783,"tokens_out":4061,"duration_ms":46017,"significance":"If the central trade-off claim is correct, this is a valuable and timely contribution: it provides direct experimental evidence on how generative AI writing assistance affects both producers and consumers in a realistic discussion environment, with transparent prompts, multiple distinct tool designs, and rich usage data. The study's strengths include the realistic platform, the comparison of four different AI-assistance paradigms, the inclusion of both behavioral and perceptual measures, and the unusually detailed supplementary documentation of prompts, questionnaires, and regression analyses. The main result—that AI can boost participation metrics while degrading perceived quality—is plausible and worth publishing, but the current manuscript states at least one central claim (negative spillover) that the design cannot identify, and the statistical inference ignores the group-level randomization structure. These issues need to be resolved before the paper can be accepted.","major_comments":[{"comment":"The abstract's claim that AI tools 'introduce a negative spill-over effect on conversations' and the stronger statement in 'How to move forward' that AI lowers 'the quality of subsequent conversations within threads, even among users not using the AI themselves' are not identifiable from the experimental design. As stated in Methods ('Platform Design'), each group was randomly assigned to one condition and participants interacted only within their assigned group, so AI availability is constant within a group and there is no within-group control of non-AI users. The only non-users in treatment groups are participants who voluntarily chose not to click the tool, a self-selected subset. No analysis in the paper compares conversation quality before and after AI-assisted comments or isolates non-AI-user outcomes within treatments. This claim should be removed from the abstract or explicitly reframed as a hypothesis, unless the authors add a dedicated within-treatment analysis that defines non-users and compares their conversation quality in a way that is not confounded by self-selection.","section":"Abstract; 'How to move forward' (p. 10)"},{"comment":"The statistical inference appears to ignore the group-level randomization and the dependence among participants within a group. The treatment is assigned at the group level, and participants within a group interact with one another, so individual-level observations are not independent. The paper reports t-tests and permutation tests on individual-level metrics without stating the resampling unit of the bootstrap and without cluster-robust standard errors or group-level aggregation for most outcomes. This likely inflates significance levels, especially for outcomes such as reaction distributions and perceived-AI-use ratings in Fig. 12. In addition, the paper makes many comparisons across conditions and outcomes without any multiple-comparison correction, so some of the reported 'significant' effects may be chance findings. The authors should either reanalyze at the group level, use cluster-robust inference, or explicitly label the results as exploratory and unadjusted; this is consequential for the central claim that 'no single AI tool enhances both producer and consumer experiences,' which depends on several non-significant and marginal contrasts.","section":"Methods, 'Evaluation'; Supplementary Material, 'Statistical Tests and Bootstrapping'"},{"comment":"The abstract claims that AI tools increase 'volume of generated content,' but the only volume-related measure reported is average comment length in words (Fig. 2c). No result on the total number or rate of comments per participant is presented. If comment counts did not differ between treatment and control, the word 'volume' overstates the finding. The authors should report the comment-count outcome explicitly, or temper the claim to 'longer comments' rather than 'volume of generated content.'","section":"Abstract; Results (p. 3)"}],"minor_comments":[{"comment":"The caption lists two entries labeled 'e' ('Suggestions tool' and 'Chat assistant'); the second should be 'f'.","section":"Fig. 1 caption"},{"comment":"The sentence 'the proportion of participants believed by the other users to use AI tools ranged from from 13.8%...' contains a duplicated 'from'.","section":"p. 11, 'How to move forward'"},{"comment":"The phrase 'already designed optimized for sustained engagement' appears to be missing a word ('already designed and optimized' or 'already designed to be optimized').","section":"p. 13, 'An ethical deployment'"},{"comment":"The reference to 'Figs. 3a-d and 4h-i' is slightly confusing because Figs. 3a-d and 4h-i do not cover all usage analyses; consider referring to the full set of subfigures or specifying the relevant panels more precisely.","section":"p. 6, 'How are the AI tools used?'"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of cs.HC and the experimental core is worth publishing after revision. The spillover claim in the abstract is likely to draw sharp scrutiny from reviewers and readers, and the absence of cluster-level inference is a substantive statistical weakness that should be addressed head-on. I have no concerns about novelty disclosure or citation practices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nYou should know two things about arXiv:2506.14295. First, it is a solid, transparent experimental study: 680 participants, five arms, four different AI assistance tools, three topics, and both producer and consumer measures. The main contrast — some tools increase engagement and comment volume while perceived quality and authenticity drop — is real and well supported. Second, the abstract and the “How to move forward” section claim a “negative spill-over effect on conversations,” and that claim is not identifiable from the design. Because groups are randomly assigned to one condition, there are no non-AI users inside treatment groups. Every treatment participant has the same tool available; the only “non-users” are self-selected treatment participants who chose not to click. Comparing them to users who did click is confounded with motivation and engagement. The spill-over sentence should be removed or explicitly framed as a hypothesis.\n\nWhat is new: the integration of four AI assistance paradigms in one platform with measurement of both content producers and consumers, plus the usage analysis (which prompts, which stances, how often tools are ignored). That is a genuinely new combination and a useful foundation. The paper also earns credit for open reporting: full prompts, model settings, regression tables, and questionnaire instruments are in the supplementary material.\n\nSoft spots, in proportion. The multiple-comparison issue is real but not damning: with about a dozen outcome comparisons, some of the p<0.05 markers would not survive a correction. The group-level clustering is only partially addressed: they bootstrap on group-level means for entropy, but individual-level tests for comment length and ratings may overstate significance; the paper should report cluster-robust or group-level analyses. The demographic regressions are underpowered and the authors say so. These are fixable.\n\nThe central trade-off claim — no single tool enhances both producer and consumer experience — holds up as an average treatment effect on whole groups. What does not hold up is the stronger claim about degrading subsequent conversation among non-users. That is the one load-bearing overreach.\n\nWho this is for: anyone working on AI-mediated communication, platform design, or content moderation. It deserves a serious referee. I would send it out, with a request to fix the spill-over language and the statistics before acceptance.","headline":"A well-run experiment whose central trade-off is supported, but the abstract's 'negative spill-over' claim is not identifiable from the group-level design.","tokens_in":30264,"tokens_out":2306,"would_cite":true,"duration_ms":29713,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Some generative-AI tools on social media increase participation and content volume while lowering perceived quality and authenticity of discussion, a controlled experiment with 680 U.S. participants finds.","keywords":["Generative AI","Social media","Controlled experiment","Large language models","User engagement","Content quality","Authenticity","Online discussions"],"falsifier":"A field experiment on a production social media platform that randomly enables a similar AI suggestion tool would falsify the central trade-off if engagement rises without a significant drop in users' ratings of comment quality and authenticity, or without a negative spillover onto non-AI threads.","tokens_in":29371,"feed_emoji":"🤖","tokens_out":5362,"duration_ms":53815,"temperature":0.7,"pith_summary":"This paper tries to establish how generative AI assistance changes online discussion when it is actually embedded in a social media interface. Across five randomly assigned conditions—no AI, open-ended chat, conversation starters, draft feedback, and reply suggestions—the authors find a consistent split: AI tools increase willingness to participate, comment length, and participation equality, but they do not improve how readers judge the conversation. Consumers rate AI-assisted comments as less informative and lower in quality, give more Dislike reactions, and rate replies to their own comments lower in most conditions. The paper concludes that no single AI tool enhances both the producer and the consumer experience, and it draws design principles for ethical deployment.","feed_headline":"AI assistance boosts posts but lowers discussion quality, test finds","feed_subtitle":"In a controlled study, every AI tool raised engagement but none improved how readers rated the conversation.","key_machinery":"The central object is a controlled five-condition experiment on a custom-built discussion platform: 680 U.S. participants in 136 groups of five, each group randomly assigned to control or one of four GPT-4o-backed tools (open-ended chat, conversation starters, draft feedback, stance-based reply suggestions), discussing three topics—trivial, scientific, and political—for ten minutes each. The mechanism that carries the argument is the joint measurement of producer-side behavioral metrics (comment length, participation entropy, reply likelihood, self-reported willingness to participate) and consumer-side perception metrics (ratings of comment informativeness and quality, ratings of replies received, reaction distributions, perceived AI use), with bootstrapped confidence intervals and permutation or t-tests comparing each treatment to control.","core_discovery":"The paper claims that AI assistance in social media discussion produces a split outcome: on the producer side, tools like chat assistance and reply suggestions increase willingness to participate, comment length, and participation equality; on the consumer side, none of the four tools improved perceived quality—comments were rated less informative and lower in quality in the Chat and Conversation Starter conditions, replies to one's own comments were rated lower in all but the Suggestions condition, and Dislike reactions rose across all treatments. The authors summarize this as a trade-off: no single AI tool enhances both producer and consumer experience.","pith_inferences":["The authors do not test longer-term exposure; if the quality decline persists or worsens over weeks, even tools that initially raise engagement could erode trust in platform discourse over time.","Because participants already guessed AI use at 38–44% across treatments, the paper's disclosure recommendation may need to cover not just copied text but any AI-assisted wording that readers can detect.","A direct extension would be to randomize the same four tools on an existing platform and track whether the negative spillover onto human-only threads observed here reproduces at scale."],"forward_implications":["Platforms adding AI assistance should expect more participation and longer comments but lower perceived quality, more Dislikes, and reduced authenticity ratings.","Because no tested tool improved both producer and consumer experiences, deployment choices involve a real trade-off rather than a straightforward win.","AI-assisted content can spill over negatively, degrading the quality of subsequent human conversation in the same thread even for users who did not use the AI.","Users selectively adopt AI suggestions and prefer agreement in higher-stakes topics, so tools may nudge discussion toward consensus rather than diverse viewpoints.","Transparent disclosure of directly copied AI content, alongside optional tools and personalization, is the paper's proposed route to preserving authenticity."],"supporting_citations":[{"why":"supplies the chat-intervention paradigm the Chat tool builds on and its claim that AI can improve political conversation.","marker":"[17]"},{"why":"motivates stance-based reply suggestions by showing co-writing with opinionated language models can shift users' views.","marker":"[7]"},{"why":"provides communication-strategy evidence behind the Conversation Starter intervention.","marker":"[20]"},{"why":"frames AI feedback for learning, which the Feedback tool applies to comment drafts.","marker":"[5]"},{"why":"supplies the LLM-based rewriting approach that the Feedback tool adapts.","marker":"[19]"},{"why":"motivates LLM-based expansion and suggestion of ideas for the Suggestions tool.","marker":"[18]"},{"why":"provides the virtual-lab platform used to run the controlled experiment.","marker":"[28]"},{"why":"supply the real Reddit seed threads used as starting content for the three discussion topics.","marker":"[29-31]"},{"why":"introduces the 'semantic garbage' notion used to interpret the lower perceived quality of AI-assisted content.","marker":"[21]"}],"fun_headline_variants":["AI boosts posting but lowers comment quality in social test","Engagement rises, quality falls: AI in social media experiment","AI assistance spurs comments, but users rate them worse","Social AI tools increase output, decrease authenticity","In social groups, AI aids participation but undermines quality"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that four specific GPT-4o-powered tools, used by strangers in forced 10-minute discussions on three topics, represent the range of generative-AI assistance users actually encounter on real social media platforms.","fun_headline_variants_meta":{"raw":{"variants":["AI boosts posting but lowers comment quality in social test","Engagement rises, quality falls: AI in social media experiment","AI assistance spurs comments, but users rate them worse","Social AI tools increase output, decrease authenticity","In social groups, AI aids participation but undermines quality"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1460,"prompt_tokens":889,"completion_tokens":571,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":493}},"tokens_in":505,"tokens_out":571,"duration_ms":6479,"temperature":1.0,"reasoning_tokens":493,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:17:36.134071+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field experiment on a production social media platform that randomly enables a similar AI suggestion tool would falsify the central trade-off if engagement rises without a significant drop in users' ratings of comment quality and authenticity, or without a negative spillover onto non-AI threads.","supporting_citations":[{"cited_title":"Proceedings of the National Academy of Sciences120(41), 2311627120 (2023) https://doi.org/10","cited_arxiv_id":null,"evidence_quote":"supplies the chat-intervention paradigm the Chat tool builds on and its claim that AI can improve political conversation."},{"cited_title":"In: Proceedings of the 14th Conference on Creativity and Cognition","cited_arxiv_id":null,"evidence_quote":"motivates LLM-based expansion and suggestion of ideas for the Suggestions tool."},{"cited_title":"Behavior Research Methods 53(5), 2158–2171 (2021) https://doi.org/10.3758/ s13428-020-01535-9","cited_arxiv_id":null,"evidence_quote":"provides the virtual-lab platform used to run the controlled experiment."}],"review_version":1}