{"id":"b1117f04-6836-4d36-af0d-3b09c465e70a","arxiv_id":"2501.01220","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-week experiment shows cooperative chatbots improve impressions and empathy, while cooperation on content unrelated to mental illness increases coercion beliefs about people with mental illness.","lead":"This study compares three chatbot designs for reducing mental illness stigma: reading information, cooperating on mental-health content, and cooperating on unrelated content, across two weeks with 78 university students. It finds cooperation boosts perceived chatbot intelligence, likeability, and user empathy, but cooperation on unrelated topics can increase support for coercive treatment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that human-AI cooperation reduces stigma is unsupported by the present design: all three chatbot groups reduced SDS equally (no group-by-time interaction, Section 4.2.1) and there was no no-contact control arm, so the observed time effect cannot be attributed to cooperation.","rationale":"The paper's own central claim is causal: human-AI cooperation reduces mental illness stigma by fostering relationships through social contact. The most load-bearing condition for that claim is that the observed pre-post change in stigma is caused by the chatbot interaction. The data do not establish this because the main stigma outcome, the Social Distance Scale, shows only a time main effect and no group-by-time interaction, and there is no no-contact control group. Without such a counterfactual, the improvement could be due to regression to the mean, repeated testing, or demand characteristics, all of which are plausible in a two-week repeated-measures design with substantial attrition and compensation. The paper's limitations section discusses the omission of a 'non-cooperative × other topics' condition but never mentions the absence of a no-intervention control, making this gap easy to overlook and therefore more damaging to the headline conclusion.\n\nThe comparative findings about chatbot impressions (intelligence, likeability, empathy) are on firmer ground because they rely on between-group tests with real differences, and the qualitative data are coded with acceptable inter-rater reliability. However, those outcomes concern perceptions of the chatbot, not the stigma reduction that the Abstract emphasizes. The coercion finding is intriguing but, as the reader noted, lacks a reported interaction test and is a secondary outcome; it does not rescue the causal claim about overall stigma reduction.\n\nI agree with the reader's weakest-assumption analysis and with the conditional verdict. The concern does not warrant rejection because the paper's comparative contributions are valuable and the design is a reasonable step forward if reframed as a comparison of interaction styles rather than a definitive causal test of cooperation's effect on stigma. Adding a no-contact control arm in future work, or at least tempering the Abstract's causal language, would bring the claims in line with the evidence. The concrete test proposed above would settle whether the missing control is actually consequential or whether the active-chatbot effect is robust.","tokens_in":29025,"tokens_out":4141,"duration_ms":44102,"concrete_test":"Run a randomized controlled replication with a fourth no-contact arm: same recruitment pool, same pre/post SDS and Attribution surveys, same two-week interval, no chatbot interaction, same compensation schedule. Test the group-by-time interaction treating the no-contact arm as the reference. If the no-contact arm shows a pre-post SDS reduction statistically indistinguishable from Groups 1-3, the claim that cooperation reduces stigma is not supported; if the chatbot groups show significantly larger reductions, the causal claim gains support. A less costly analytical alternative: re-analyze the existing SDS data using a within-subject test-retest reliability estimate to bound the regression-to-the-mean contribution, but only the experimental control can settle the causal question.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the Abstract's causal claim that human-AI cooperation reduces stigma, the paper must show that the pre-post improvement in SDS is attributable to the cooperation manipulation rather than to mere passage of time, repeated testing, or demand characteristics. Section 4.2.1 reports only a significant time-point effect (F=20.53, p<.001) with no group effect and no group-by-time interaction for SDS. Section 3 describes three active chatbot groups and explicitly excludes a non-cooperative × other-topics condition, but no no-contact/no-intervention control is described anywhere in Methods or Limitations. Thus, the SDS reduction observed in all three groups is equally compatible with the null hypothesis that any sixteen-day engagement with a mental-health chatbot (or retesting alone) reduces self-reported social distance. The between-group differences in perceived intelligence/likeability and in empathetic expression are better controlled, but those outcomes do not establish the stigma-reduction mechanism highlighted in the Abstract. The Limitations section (5.6) does not acknowledge this missing counterfactual, which is the load-bearing gap for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-week mixed-methods study in which 78 university students interacted with one of three chatbots: one that delivered mental-illness information one-way (Group 1), one that engaged in a cooperative summarization task on mental-illness content (Group 2), and one that engaged in the same cooperative task but on sleep-related content (Group 3). The authors measure pre/post stigma via the Social Distance Scale and Attribution Questionnaire, code empathetic responses in conversation logs, and supplement the surveys with semi-structured interviews. The paper's central claim is that human-AI cooperation reduces mental-illness stigma by fostering social contact relationships, while also reporting that cooperative chatbots are perceived as more intelligent and likeable and that cooperation on unrelated content can increase coercive attitudes. The design and reporting are careful in many respects, but the core causal claim about stigma reduction is not supported by the statistical analyses presented.","tokens_in":29235,"tokens_out":3891,"duration_ms":40433,"significance":"The comparative findings on user impressions, empathy, and the coercion backfire are potentially useful for HCI and CSCW research on chatbot-based anti-stigma interventions. Strengths include the longitudinal two-week deployment, the triangulation of surveys with conversation logs and interviews, the unusually transparent appendix with system prompts and survey items, and the report of inter-rater reliability for qualitative coding. If the claims are appropriately reframed, the study contributes empirical evidence about how cooperative interaction shapes users' impressions of a chatbot representing a stigmatized identity. However, the headline claim that cooperation reduces stigma is not established by the current design, and this limits the significance of the paper in its present form.","major_comments":[{"comment":"The headline claim that human-AI cooperation reduces stigma is not supported by the reported analyses. For the main SDS outcome, the Scheirer-Ray-Hare test shows only a time-point effect (F=20.53, p<.001), with no group effect and no group-by-time interaction; all three conditions improved to a statistically indistinguishable degree. In the absence of a no-contact control, this pattern is equally compatible with repeated-testing effects, demand characteristics, regression to the mean, or any 16-day engagement with a mental-health chatbot. The causal wording of the Abstract and the statement in Section 5.1 that 'all three groups showed an overall reduction in stigma' therefore overstate what the design can establish.","section":"Abstract; Section 4.2.1"},{"comment":"The design omits both a no-contact control and the 'non-cooperative × other topics' condition. The exclusion of the latter is explicitly justified in Section 3 by theory and prior literature, but the former is not mentioned anywhere in Methods or Limitations. Section 5.6 lists four limitations, yet it does not acknowledge that the pre-post comparison cannot rule out retesting, passage of time, or demand characteristics as explanations for the SDS reduction. This is the load-bearing gap for RQ2 and the Abstract's causal claim. The manuscript should either add a no-contact control arm or reframe the stigma-reduction conclusion as an exploratory within-design trend rather than a causal effect of cooperation.","section":"Section 3; Section 5.6"},{"comment":"The coercion 'backfire' conclusion is not supported by the reported test statistics. The paper reports a group-membership effect for coercion (F=12.10, p<.01) and pairwise between-group comparisons, but no group-by-time interaction and no within-group pre-post paired tests. The claim that Group 3 'demonstrated an increase' in coercion (pre M=4.17, post M=4.65) requires evidence of differential change over time, which a group main effect does not provide. Please report the interaction term and within-group paired comparisons, or soften the causal claim about an increase caused by the unrelated-content cooperation.","section":"Section 4.2.1; Section 5.4"}],"minor_comments":[{"comment":"The SDS response anchors are inconsistent: Section 3.5.1 states 0='definitely willing' and 3='definitely unwilling', while Appendix A.2.1 lists the reverse. Please align the two descriptions so readers can interpret the direction of the pre-post decrease.","section":"Section 3.5.1; Appendix A.2.1"},{"comment":"Kruskal-Wallis tests are reported with F statistics (e.g., F=9.67, p<.01), but the Kruskal-Wallis test produces an H (chi-square) statistic, not an F. The same issue appears in the empathy comparison reported as F=3.42, p=.064. Please correct the statistics or clarify which test was actually used.","section":"Section 4.1.1"},{"comment":"The descriptive statistics for Group 2's SDS are incomplete: the text gives pre M=11.58 and post M=8.92, SD=3.87, but no pre-test standard deviation. Please add the missing value for consistency with the other groups.","section":"Section 4.2.1"},{"comment":"There are a few typographical errors: 'in Second 4.1.2' should be 'in Section 4.1.2', and 'leading 3 Group to lack' in Section 5.3 should be 'leading Group 3 to lack'. These are minor but should be corrected.","section":"Section 5.2; Section 5.3"},{"comment":"The citation placeholder '95?' appears in the sentence about chatbots as social actors. Please replace it with a proper citation or remove the question mark.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper's comparative evidence on user impressions, empathy, and coercion is a genuine contribution, but the abstract and conclusion claim more than the design can support. The missing no-contact control and the absence of a group-by-time interaction for the main SDS outcome are the central issues. I believe the paper is salvageable through a substantial reframing of the central claim and more careful reporting of the coercion analysis, so I recommend major revision rather than rejection. Self-citation is not a concern here: the prior chatbot study is directly extended, and the comparative results are new."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, well-documented mixed-methods comparison of three chatbot conditions, and the between-group differences in impressions and coercion are worth taking seriously. The abstract's broader claim that human-AI cooperation reduces stigma is not supported by the reported data. On the SDS, all three groups improved over time with no group-by-time interaction, and there is no no-contact control. That improvement could be retesting, demand, or simple passage of time. The stress-test note is correct on that point.\n\nWhat is actually new: prior chatbot stigma work mostly used narrative-only delivery. Here they compare one-way information against two-way cooperation, crossed with mental-health versus unrelated content, over two weeks. The finding that cooperation on unrelated content increases coercion beliefs, while mental-health cooperation and even one-way information decrease them, is genuinely novel. The qualitative material on shared goals and equal status gives a plausible mechanism, and the authors are admirably transparent about their system design, prompt details, and the few GPT failures they observed.\n\nSoft spots: (1) the missing control group is a real gap. The authors justify excluding a non-cooperative × other-topic condition on theoretical grounds, but that is an assumption, not a control. Without a no-contact arm, the time effect on SDS cannot be attributed to the manipulation. The limitations section does not flag this. (2) Several statistics are mislabeled: Kruskal-Wallis results are reported as F values, and the empathy comparison between Groups 2 and 3 is reported as F = 3.42 though it is a chi-square test. More importantly, the empathy analysis treats individual messages as independent units, which pseudoreplicates the data; per-participant rates would be more defensible. (3) The coercion finding is presented as a group-membership effect, but I would want to see a direct group-by-time interaction test on that subscale. The post-hoc comparisons describe changes, but the paper does not report the interaction term. (4) The abstract says \"cooperation can effectively reduce stigma\" when the only group-specific stigma effect is on coercion, one subscale among several.\n\nWho is this for: CSCW and health-HCI researchers working on chatbot-mediated social contact, and anyone designing anti-stigma interventions. With the causal language toned down and the statistics cleaned up, this could be a solid contribution to that community. It deserves peer review, not a desk reject, but reviewers should press hard on the missing counterfactual and the coercion interaction.\n\nRecommendation: send to peer review, expecting major revisions.","headline":"A competent comparative chatbot study whose abstract overstates the causal claim; the group-specific coercion backfire is the real contribution.","tokens_in":29751,"tokens_out":2558,"would_cite":true,"duration_ms":28588,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Cooperative chatbot interactions reduce mental-illness stigma by fostering a social relationship, while off-topic cooperation backfires by increasing coercion beliefs.","keywords":["human-AI cooperation","chatbot","mental illness stigma","social contact intervention","Intergroup Contact Hypothesis","user impressions","coercion beliefs","longitudinal mixed-methods study"],"falsifier":"Run the same protocol with a fourth group that completes the pre/post surveys after two weeks without any chatbot contact; if that group shows the same decline in social-distance scores as the chatbot groups, the causal claim would be falsified. The backfire claim would be falsified if a replication of the unrelated-content cooperation condition fails to produce a rise in coercion beliefs.","tokens_in":28829,"feed_emoji":"🤖","tokens_out":6493,"duration_ms":65028,"temperature":0.7,"pith_summary":"This paper tests whether the way a person interacts with a chatbot changes that person's stigmatizing attitudes toward mental illness. It compares a one-way information-dissemination chatbot with two cooperative chatbots, one discussing mental-illness material and one discussing unrelated material, over a two-week study. The authors claim that human-AI cooperation reduces stigma by building a social relationship with the chatbot: cooperative users rated it more intelligent and likable, expressed more empathy in conversation, and shifted toward external explanations for the persona's depression. They also report a backfire effect: cooperation on unrelated content increased beliefs that people with mental illness should be coerced into treatment. The study's broader point is that interaction design, not just content delivery, shapes whether a chatbot can change attitudes.","feed_headline":"Cooperative chatbots cut mental-illness stigma when topics align","feed_subtitle":"Two-week trial: two-way cooperation beats one-way information, while off-topic cooperation raises coercion beliefs.","key_machinery":"The load-bearing mechanism is a two-week daily Telegram chatbot named Holly, a university-student persona with depression, whose interactions were scripted plus GPT-3.5 responses. On odd days all participants heard a first-person vignette about Holly's life; then Group 1 read NIH mental-illness material, while Groups 2 and 3 did a cooperation task in which chatbot and user alternately summarized the day's material and corrected one another (listener and recaller roles). This task operationalizes Allport's Intergroup Contact Hypothesis, with equal status, shared goals, and intergroup cooperation, and it is what the paper credits for making users feel they shared a goal with Holly, perceive her as competent and likable, and respond empathetically.","core_discovery":"The paper's central claim is that a two-way cooperative interaction with a chatbot can serve as a social-contact intervention against mental-illness stigma by creating the conditions under which users come to see the chatbot as a likable, competent, in-group member. Cooperative users rated the chatbot higher on intelligence and likeability than one-way information recipients, and their conversation logs contained more empathetic reactions (0.58 and 0.49 in the two cooperation groups versus 0.39 in the information group). Over the two weeks, all three conditions reduced social distance and fear/dangerousness ratings, but the two mental-health-content groups also reduced coercion beliefs, while the unrelated-content cooperation group increased them. The authors interpret this rise as a backfire: Holly's skilled task performance clashed with her depressed persona, making her condition seem controllable and therefore blameworthy.","pith_inferences":["Editorial inference: the pattern of results suggests cooperation is not inherently stigma-reducing; the meaning of the shared goal, helping Holly versus just learning material, is likely the active ingredient, so a 2x2 experiment that independently varies cooperation and topic would isolate the mechanism.","Editorial inference: because no no-contact control condition was run, the paper cannot rule out that any attentive two-week interaction produces similar social-distance declines; a minimal-contact control is the cheapest next test of the causal claim.","Editorial inference: a testable design rule follows from the backfire finding: if an embodied agent represents a stigmatized group, its in-task competence should be narratively framed as coping or recovery, so that competence does not translate into blame and coercion."],"forward_implications":["Chatbots can be deployed as low-cost anti-stigma tools, but only when the cooperation task shares the mental-health context; off-topic cooperation may do harm.","Cooperative designs improve perceived intelligence and likeability, which the paper links to user acceptance and engagement in future human-AI applications.","The listener-and-recaller task combined with first-person vignettes provides a concrete template for designing social-contact chatbot interventions.","Including recovery and coping information in mental-health content appears to prevent competence perceptions from reinforcing the controllability stereotype.","If the pre-post reductions are causal, repeated cooperative interaction over two weeks is enough to shift some stigmatizing beliefs, supporting longitudinal chatbot-based anti-stigma campaigns."],"supporting_citations":[{"why":"Supplies the Intergroup Contact Hypothesis and the four conditions (equal status, shared goals, intergroup cooperation, authority support) that the cooperation design operationalizes.","marker":"[4]"},{"why":"Prior chatbot-based social contact study that this work extends by adding cooperative interaction and user-experience evaluation; also supplies the vignette design and the SDS/attribution measurement precedent.","marker":"[60]"},{"why":"Provides the Attribution Questionnaire and the responsibility, coercion, and social-distance subscales used to measure stigma outcomes.","marker":"[23]"},{"why":"Establishes structured cooperative contact as a stigma-reduction method in human-human interaction, which the paper adapts into the listener-and-recaller chatbot task.","marker":"[33]"},{"why":"Contributes the user impression instrument (intelligence, rapport, likeability, creativity) used to measure how participants perceived the chatbot.","marker":"[7]"},{"why":"Meta-analysis on intergroup contact and mental-health stigma used to support the social-contact approach and to motivate the comparison of related versus unrelated content topics.","marker":"[70]"},{"why":"Critiques prior contact-based studies for lacking control groups and motivates the paper's comparative three-group design.","marker":"[51]"}],"fun_headline_variants":["Cooperative chatbots fight stigma only when talk matches topic","Two-way chatbot talk cuts stigma, but off-topic backfires","Human-AI cooperation fights stigma, but off-topic chat worsens blame","Cooperation with chatbot eases stigma unless topic is unrelated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the chatbot interactions caused the attitude changes assumes that the pre-post improvements were not due to taking the same survey twice, trying to look non-stigmatizing, or natural regression; no no-contact comparison was run to test this.","fun_headline_variants_meta":{"raw":{"variants":["Cooperative chatbots fight stigma only when talk matches topic","Two-way chatbot talk cuts stigma, but off-topic backfires","Human-AI cooperation fights stigma, but off-topic chat worsens blame","Cooperation with chatbot eases stigma unless topic is unrelated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000913,"raw_usage":{"total_tokens":3900,"prompt_tokens":904,"completion_tokens":2996,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":520,"completion_tokens_details":{"reasoning_tokens":2924}},"tokens_in":520,"tokens_out":2996,"duration_ms":20084,"temperature":1.0,"reasoning_tokens":2924,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:31:52.619230+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same protocol with a fourth group that completes the pre/post surveys after two weeks without any chatbot contact; if that group shows the same decline in social-distance scores as the chatbot groups, the causal claim would be falsified. The backfire claim would be falsified if a replication of the unrelated-content cooperation condition fails to produce a rise in coercion beliefs.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior chatbot-based social contact study that this work extends by adding cooperative interaction and user-experience evaluation; also supplies the vignette design and the SDS/attribution measurement precedent."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Meta-analysis on intergroup contact and mental-health stigma used to support the social-contact approach and to motivate the comparison of related versus unrelated content topics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Critiques prior contact-based studies for lacking control groups and motivates the paper's comparative three-group design."}],"review_version":1}