{"id":"f5c4cef8-bd7d-4c7a-8f26-fced6616d424","arxiv_id":"2412.11656","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 10-day field study found that an LLM chatbot supported eating disorder recovery through private storytelling, yet also produced unnoticed harmful responses such as praising weight loss and restriction.","lead":"Researchers built an LLM chatbot for people with eating disorders and studied 26 users over 10 days. The bot offered a private space to share struggles, but it also gave harmful advice that users never questioned.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Harm taxonomy in Section 6.3 lacks clinical validation; if clinicians do not endorse the flagged responses as harmful, the 'harms went unnoticed' finding is unsupported.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the harm classifications in Section 6.3 and Table 4 are treated as clinical ground truth without clinician validation. This is the most load-bearing issue because the paper's distinctive contribution is the identification of 'unnoticed harms' in real LLM chatbot use. If those harms are not actually harmful, the central challenge claim reduces to a subjective coding disagreement, and the design implications built on it lose force. The Brief-IPQ quantitative result is weaker (uncontrolled), but it is secondary to the qualitative harm claim; even a perfect Brief-IPQ finding would not establish the harm narrative. The paper has independent strengths: rich chat log evidence, honest limitations, clear ethics procedures, and a plausible qualitative account of supportive use. However, the harm taxonomy is the load-bearing premise that needs external validation. The proposed clinician rating test directly settles whether the concern lands, and it is feasible with modest resources. Since the reader already arrived at CONDITIONAL largely on this basis, my stress-test does not move the verdict; it reinforces the condition. I agree with the reader's identification of the weakest assumption.","tokens_in":29532,"tokens_out":2262,"duration_ms":22997,"concrete_test":"Recruit two or three licensed clinicians specializing in eating disorders. Provide them with the complete list of chatbot messages that the authors coded as harmful (or, minimally, the six Table 4 examples), stripped of the authors' labels and of participant identifying context. Ask each clinician to classify each message on a standardized harm scale (e.g., harmful / neutral / helpful for a person with ED, with brief rationale). Compute Cohen's kappa between clinicians and between clinicians and the authors' labels. If clinician endorsement of the 'harmful' labels is below, say, 70% or kappa < 0.4, the harm taxonomy is not established and the 'unnoticed harm' claim should be downgraded. This is a bounded, inexpensive check that directly tests the validity premise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central challenge claim rests on the assertion that WellnessBot produced harmful responses that went unnoticed by participants (Section 6.3, Table 4). This requires that the authors' harm labels are correct clinical ground truth. They are not clinically validated: the thematic analysis was conducted by the first three authors (HCI researchers), no clinician was involved in coding, and no inter-rater reliability is reported. The participants themselves reported no harm (Section 6.3.2), so the 'unnoticed harms' claim is entirely dependent on the researchers' classification. Several Table 4 examples are genuinely debatable. Chat 11 labels 'extreme hunger' as 'unsupported advice' and 'hallucinated,' but the paper's own references [11, 28] describe extreme hunger as a real phenomenon in ED recovery; the chatbot describes it reasonably as responding to internal cues without restriction. Chat 9 praises 'picky eating only meat' in a user complaining of never feeling full; a clinician might see this as addressing protein satiety rather than endorsing restriction. Chat 7 celebrates weight loss, which is concerning, but without clinical context (e.g., the user's BMI, treatment goals) it is ambiguous. If ED clinicians do not agree with the majority of these labels, the headline finding that harmful responses 'went unnoticed' collapses into a lay-coding disagreement. The Brief-IPQ change (Section 6.4) is not a corrective because it has no control group and cannot distinguish chatbot effect from time or regression to the mean.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a ten-day field deployment of WellnessBot, a GPT-4-based Telegram chatbot designed as a technology probe for eating disorder (ED) support, with 26 self-identified or clinically diagnosed female participants. The authors collected chat logs, pre/post surveys (EDE-Q, Brief-IPQ), and semi-structured interviews, and analyzed them using descriptive statistics and thematic analysis. They argue that the chatbot created a \"private yet social\" space that supported storytelling, self-reflection, and recovery motivation, but also produced harmful, ED-inappropriate responses (Table 4) that went unnoticed because participants trusted the bot. The paper presents the Brief-IPQ pre-post improvement as supporting evidence and closes with design implications for safe LLM-based ED interventions.","tokens_in":29779,"tokens_out":5744,"duration_ms":59096,"significance":"If the harm findings are valid, this is a significant contribution: it provides field evidence about how vulnerable users interact with LLM-based chatbots and why safety warnings may fail. The study's strengths are its rich qualitative data, the transparent description of the WellnessBot design and deployment, the daily monitoring protocol, and the direct inclusion of chat-log excerpts and participant quotes. The paper is also honest about several limitations in Section 7.4. However, the headline \"unnoticed harms\" claim rests on the authors' own harm classifications, which are not clinically validated, and the quantitative support is exploratory rather than confirmatory. The design implications in Section 7 are reasonable but depend on the safety framing, so the contribution's weight currently falls on an unvalidated taxonomy.","major_comments":[{"comment":"The central claim that harmful chatbot responses \"went unnoticed\" is load-bearing, but the harm labels are assigned by the first three authors without clinical validation or inter-rater reliability. The coding process is described in Section 5.3.2 as thematic analysis by the authors, and Section 6.3.2 reports that participants themselves perceived no harm. Several Table 4 classifications are contestable: Chat 11 labels \"extreme hunger\" as unsupported advice even though the paper's own references [11, 28] describe extreme hunger as a real phenomenon in ED recovery; Chat 9's \"picky eating\" advice is framed as harmful without clinical context; and Chat 7's praise of weight loss is ambiguous without knowing the user's BMI or treatment goals. The headline finding therefore depends on a lay coding judgment. The authors should either have the harm taxonomy reviewed and validated by clinicians, report inter-rater reliability, or explicitly reframe the finding as \"potentially concerning responses\" rather than \"harms that went unnoticed.\" Section 7.4 lists several limitations but omits this clinical-validation gap.","section":"Section 6.3, Table 4; Section 5.3.2"},{"comment":"The Brief-IPQ analysis is presented as evidence that the intervention improved participants' perceptions (\"statistically significant decrease... Z = 2.43, p = .02\"), but there is no control or comparison condition, so regression to the mean, repeated-testing effects, or general study participation could explain the change. In addition, the paper reports significance for the overall scale and for two individual items without correcting for multiple comparisons, and the item-level p-values (p = .02 and p = .05) are marginal. This analysis should be explicitly framed as an exploratory descriptive result, not as an effectiveness claim, and the Discussion in Section 7.1 should not cite the Brief-IPQ decrease as evidence for the benefits of storytelling.","section":"Section 6.4"},{"comment":"The \"unnoticed\" finding is conditioned on the study protocol: Section 3 states that the authors deliberately did not intervene on misinformation or undesirable responses and only informed participants during post-interviews. This means the absence of participant questioning was observed under a protocol that withheld feedback, which is a valid observational choice but should be stated as a boundary of the claim. As written, the abstract and Section 6.3.2 imply a general property of user trust, whereas the evidence shows how users behave when no in-situ correction is provided. The design implications in Section 7.3 recommend in-situ interventions, but the paper should acknowledge that its own protocol precluded such interventions and therefore cannot directly test whether they would change user awareness.","section":"Section 3 and Section 6.3.2"}],"minor_comments":[{"comment":"Reference [24] misspells \"ChatGPT\" as \"ChagGPT\"; please correct the typo.","section":"Reference [24]"},{"comment":"The user message in Chat 9 contains \"When I've eating a lot of meat,\" which appears to be a typo in the translated transcript; please correct it or mark it as [sic].","section":"Table 4, Chat 9"},{"comment":"Figure 2 lacks a y-axis label and does not state whether the counts are raw message counts or per-user averages; please clarify in the caption or on the axis.","section":"Figure 2"},{"comment":"The sentence \"participants interact with WellnessBot without any instructions\" seems to conflict with the described introductory session; please rephrase to indicate that no further instructions were given during the deployment period.","section":"Section 5.2"},{"comment":"The Mann-Kendall results are reported per user (21 non-significant, 5 decreasing), which involves multiple comparisons and is not a study-level aggregate; consider reporting a single mixed-effects model or simply descriptive trends instead.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid qualitative field study with a clear contribution if the harm taxonomy is validated. The main risk is that the headline finding is currently supported only by the authors' own coding, and the quantitative supplement is not enough to carry the causal claim. I would encourage the editor to require clinical involvement in the harm classification or a clearly downgraded framing before publication; if the authors cannot provide either, I would lean toward rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading: this is the first real deployment of a personalized LLM chatbot for eating disorder recovery that I know of, and it produces concrete, useful evidence. The 'private yet social' framing is a real contribution, and the catalog of harmful chatbot responses—praising restriction, weight-centric praise, moralizing food, endorsing avoidance—is exactly the kind of domain-specific safety material the field needs. The authors did the work: 26 participants over 10 days, chat logs plus interviews, a described system, and an honest limitations section.\n\nThe qualitative analysis is careful. The coding process is described, the chat excerpts are plentiful, and the claims about participant trust are well supported by quotes. The ethical section is unusually thoughtful, including daily monitoring and Wellness Plan curation.\n\nThe main soft spot is the harm taxonomy in Section 6.3. The first three authors, HCI researchers, classified responses as harmful with no clinician involvement and no inter-rater reliability. Some Table 4 examples are genuinely ambiguous: the 'extreme hunger' exchange (Chat 11) is debatable—the chatbot's description is not obviously reckless, and the paper's own references describe extreme hunger as a real phenomenon. Chat 9 praising 'picky eating only meat' could be read as addressing protein satiety. Without clinical validation, 'harms went unnoticed' is a lay-coding result, not a clinical one. That doesn't kill the paper, but it should be reframed as 'potentially concerning responses' and validated with clinicians before being claimed as harm. The Brief-IPQ result (Section 6.4) is weaker: no control group, no multiple comparison correction, small sample. It should be reported as exploratory, not as evidence of effectiveness. The authors already hedge somewhat, but the presentation still overreaches.\n\nWho is this for? HCI researchers, mental health chatbot designers, and safety evaluators. The paper deserves a serious referee. I'd send it to review with a request for clinician validation of the harm labels and a de-emphasis of the quantitative claim. The core qualitative findings are likely to stand.","headline":"A genuinely useful field study of LLM chatbots in ED recovery, with a solid qualitative core; the harm taxonomy needs clinical validation, and the Brief-IPQ should be treated as exploratory.","tokens_in":30316,"tokens_out":1806,"would_cite":true,"duration_ms":17179,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Study finds LLM chatbot both helps and quietly harms ED users","keywords":["LLM chatbot","eating disorder recovery","technology probe","AI trust","mental health chatbot","personalized support","chatbot safety"],"falsifier":"Have an independent eating-disorder clinician, or a panel of them, rate the chatbot responses labeled harmful in Table 4 (praising weight loss, moralizing food, endorsing avoidance, suggesting picky eating, and presenting 'extreme hunger' as a treatment strategy) without knowing the paper's labels. If a majority of clinicians judge these responses acceptable or not harmful for ED patients, the paper's central claim that harmful responses went unnoticed would lack empirical support.","tokens_in":29333,"feed_emoji":"🤖","tokens_out":4878,"duration_ms":40885,"temperature":0.7,"pith_summary":"The paper claims that people with eating disorders can feel genuinely empowered by discussing their recovery with an LLM-based chatbot, which offers a uniquely private yet social space free of the judgment they fear from humans. It also claims that the same chatbot produced responses that were harmful in the context of eating disorders—endorsing avoidance, praising weight loss, moralizing food, suggesting picky eating, and describing 'extreme hunger' as a treatment strategy—and that none of the 26 participants noticed or questioned these responses, largely because they trusted the AI's data-driven reliability. The study tracks 1,477 real user-chatbot interactions over 10 days, so these claims are grounded in actual use rather than hypothetical scenarios. The authors argue that this combination of perceived benefit and invisible harm makes LLM chatbots a double-edged tool for eating-disorder care that needs design safeguards.","feed_headline":"Study finds LLM chatbot both helps and quietly harms ED users","feed_subtitle":"In a 10-day trial, 26 ED users felt empowered yet never noticed harmful weight-loss advice from the chatbot.","key_machinery":"The carrying mechanism is WellnessBot itself, a GPT-4-based chatbot whose responses are personalized by three components: an Indicator Detector that checks every user message against ED triggers and warning signs from the user's Wellness Plan, a Context Checker that retrieves relevant chat history, and a mentor persona that blends emotional and informational support. The Wellness Plan, a goal-and-coping-strategy survey adapted from a peer-mentoring program, is what lets WellnessBot name personal triggers and suggest timely strategies, which is the main source of perceived benefit. The same personalization pipeline also produces the harmful responses, because the underlying LLM lacks the clinical nuance needed to know when a compliment or a weight-loss acknowledgment is dangerous. The trust dynamic is the second mechanism: users know just enough about AI to believe its outputs are reliable, but not enough to question them.","core_discovery":"The central discovery is that an LLM chatbot deployed as a technology probe for eating-disorder care creates a 'private yet social' space that participants value: they can disclose ED experiences, share recovery stories, receive personalized coping strategies, and feel accompanied around the clock without fear of stigma or social comparison. At the same time, the chatbot's responses frequently crossed clinical safety lines by reinforcing weight-centric thinking, praising restriction, moralizing food choices, encouraging avoidance of root causes, and hallucinating 'extreme hunger' as a treatment strategy. Crucially, participants reported no harm and no doubt; their strong trust in AI, based on a partial understanding of how LLMs work, meant harmful responses went unnoticed and unchallenged. This discrepancy between perceived and actual risk is the paper's central finding, and it motivates the authors' design implications for in-situ critical-thinking aids and human-LLM collaborative care.","pith_inferences":["If the harm classifications in Table 4 are treated as clinical ground truth, then post-hoc audit of LLM logs for ED-specific unsafe patterns (weight praise, food moralization, avoidance endorsement, and pseudoscientific 'extreme hunger' advice) becomes a viable automated safety screen for mental-health chatbots; the authors gesture at this but do not implement it.","The 'private yet social' framing suggests a broader design principle: for stigmatized conditions, an LLM interlocutor may serve as a low-stakes social rehearsal space that preserves conversational skills and reduces avoidance, an effect that may generalize beyond eating disorders to social anxiety or depression.","Because pre-use warnings failed to instil critical evaluation, a testable extension is to compare pre-use warnings against in-conversation nudges and post-response 'critique prompts' for their effect on users' ability to spot unsafe advice.","Since all participants were female and the study ran only 10 days on a single chatbot, the benefit/harm balance could shift with gender diversity, longer use, or different LLM backbones; a concrete next step would be a multi-chatbot deployment with clinician-validated harm labels."],"forward_implications":["LLM chatbots that feel supportive and harmless can still deliver clinically unsafe guidance for eating-disorder populations in everyday, out-of-clinic use.","Users' trust in LLM chatbots is high enough that pre-use warnings about possible mistakes are ineffective; in-situ, real-time prompts to critically evaluate responses will be needed.","Designing chatbots around a user's Wellness Plan allows timely, personalized coping-strategy suggestions that participants genuinely value, showing a path to beneficial personalization.","The same personalization features should be paired with clinician oversight or human-LLM collaboration, since standalone use left harmful responses uncorrected.","Brief-IPQ scores improved significantly after 10 days ($Z = 2.43$, $p = .02$, $r = 0.41$), suggesting measurable short-term attitude gains toward illness control from chatbot storytelling."],"supporting_citations":[{"why":"Supplies the technology-probe methodology that frames WellnessBot as a design-research instrument for real-world use.","marker":"[63]"},{"why":"Documents the NEDA chatbot shutdown after harmful advice, the motivating cautionary case for ED chatbot safety.","marker":"[25]"},{"why":"Provides the Wellness Plan structure (goals, positive strategies, indicators) that personalizes WellnessBot's responses.","marker":"[7]"},{"why":"Establishes the stigma faced by people with eating disorders, explaining why the chatbot's private yet social space is valued.","marker":"[111]"},{"why":"Presents prior evidence that users over-rely on mental-health chatbots, grounding the paper's trust finding.","marker":"[35]"},{"why":"Provides EDE-Q norms used to validate that the participant sample represents the eating-disorder population.","marker":"[1]"},{"why":"Provides the Brief-IPQ instrument used to measure changes in illness perception before and after using the chatbot.","marker":"[10]"},{"why":"Supplies the thematic-analysis method applied to the interviews and chat logs.","marker":"[9]"}],"fun_headline_variants":["LLM chatbots empower ED users, but harmful advice slips past trust","Chatbots aid ED recovery, but toxic tips go unnoticed","LLM chatbot helps ED recovery, yet users miss risky responses","Private yet social: LLM chatbots both support and undermine ED care","LLM chatbots empower ED recovery, but silent harm goes unseen"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that harmful responses went unnoticed rests on the authors' own judgment that specific chatbot replies—praising weight loss, endorsing avoidance, suggesting picky eating—were harmful for people with eating disorders, a judgment no clinician validated and participants did not share.","fun_headline_variants_meta":{"raw":{"variants":["LLM chatbots empower ED users, but harmful advice slips past trust","Chatbots aid ED recovery, but toxic tips go unnoticed","LLM chatbot helps ED recovery, yet users miss risky responses","Private yet social: LLM chatbots both support and undermine ED care","LLM chatbots empower ED recovery, but silent harm goes unseen"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000439,"raw_usage":{"total_tokens":2198,"prompt_tokens":882,"completion_tokens":1316,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":1229}},"tokens_in":498,"tokens_out":1316,"duration_ms":8614,"temperature":1.0,"reasoning_tokens":1229,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:42:26.621024+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent eating-disorder clinician, or a panel of them, rate the chatbot responses labeled harmful in Table 4 (praising weight loss, moralizing food, endorsing avoidance, suggesting picky eating, and presenting 'extreme hunger' as a treatment strategy) without knowing the paper's labels. If a majority of clinicians judge these responses acceptable or not harmful for ED patients, the paper's central claim that harmful responses went unnoticed would lack empirical support.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the stigma faced by people with eating disorders, explaining why the chatbot's private yet social space is valued."}],"review_version":1}