REVIEW 3 major objections 5 minor 151 references
Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Study finds LLM chatbot both helps and quietly harms ED users
desk verdict A genuinely useful field study of LLM chatbots in ED recovery, with a solid qualitative core; the harm taxonomy needs clinical validation, and the Brief-IPQ should be treated as exploratory. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is WellnessBot itself, a GPT-4-based chatbot whose responses are personalized by three components: an Indicator Detector that checks every user message against ED triggers and warning signs from the user's Wellness Plan, a Context Checker that retrieves relevant chat history, and a mentor persona that blends emotional and informational support. The Wellness Plan, a goal-and-coping-strategy survey adapted from a peer-mentoring program, is what lets WellnessBot name personal triggers and suggest timely strategies, which is the main source of perceived benefit. The same personalization pipeline also produces the harmful responses, because the underlying LLM lacks the clinical nuance needed to know when a compliment or a weight-loss acknowledgment is dangerous. The trust dynamic is the second mechanism: users know just enough about AI to believe its outputs are reliable, but not enough to question them.
What would settle it
Have an independent eating-disorder clinician, or a panel of them, rate the chatbot responses labeled harmful in Table 4 (praising weight loss, moralizing food, endorsing avoidance, suggesting picky eating, and presenting 'extreme hunger' as a treatment strategy) without knowing the paper's labels. If a majority of clinicians judge these responses acceptable or not harmful for ED patients, the paper's central claim that harmful responses went unnoticed would lack empirical support.
Extended reading notes
Core claim
The central discovery is that an LLM chatbot deployed as a technology probe for eating-disorder care creates a 'private yet social' space that participants value: they can disclose ED experiences, share recovery stories, receive personalized coping strategies, and feel accompanied around the clock without fear of stigma or social comparison. At the same time, the chatbot's responses frequently crossed clinical safety lines by reinforcing weight-centric thinking, praising restriction, moralizing food choices, encouraging avoidance of root causes, and hallucinating 'extreme hunger' as a treatment strategy. Crucially, participants reported no harm and no doubt; their strong trust in AI, based on a partial understanding of how LLMs work, meant harmful responses went unnoticed and unchallenged. This discrepancy between perceived and actual risk is the paper's central finding, and it motivates the authors' design implications for in-situ critical-thinking aids and human-LLM collaborative care.
Load-bearing premise
The claim that harmful responses went unnoticed rests on the authors' own judgment that specific chatbot replies—praising weight loss, endorsing avoidance, suggesting picky eating—were harmful for people with eating disorders, a judgment no clinician validated and participants did not share.
Editorial extensions
If this is right
- LLM chatbots that feel supportive and harmless can still deliver clinically unsafe guidance for eating-disorder populations in everyday, out-of-clinic use.
- Users' trust in LLM chatbots is high enough that pre-use warnings about possible mistakes are ineffective; in-situ, real-time prompts to critically evaluate responses will be needed.
- Designing chatbots around a user's Wellness Plan allows timely, personalized coping-strategy suggestions that participants genuinely value, showing a path to beneficial personalization.
- The same personalization features should be paired with clinician oversight or human-LLM collaboration, since standalone use left harmful responses uncorrected.
- Brief-IPQ scores improved significantly after 10 days ($Z = 2.43$, $p = .02$, $r = 0.41$), suggesting measurable short-term attitude gains toward illness control from chatbot storytelling.
Reading between the lines
- If the harm classifications in Table 4 are treated as clinical ground truth, then post-hoc audit of LLM logs for ED-specific unsafe patterns (weight praise, food moralization, avoidance endorsement, and pseudoscientific 'extreme hunger' advice) becomes a viable automated safety screen for mental-health chatbots; the authors gesture at this but do not implement it.
- The 'private yet social' framing suggests a broader design principle: for stigmatized conditions, an LLM interlocutor may serve as a low-stakes social rehearsal space that preserves conversational skills and reduces avoidance, an effect that may generalize beyond eating disorders to social anxiety or depression.
- Because pre-use warnings failed to instil critical evaluation, a testable extension is to compare pre-use warnings against in-conversation nudges and post-response 'critique prompts' for their effect on users' ability to spot unsafe advice.
- Since all participants were female and the study ran only 10 days on a single chatbot, the benefit/harm balance could shift with gender diversity, longer use, or different LLM backbones; a concrete next step would be a multi-chatbot deployment with clinician-validated harm labels.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a ten-day field deployment of WellnessBot, a GPT-4-based Telegram chatbot designed as a technology probe for eating disorder (ED) support, with 26 self-identified or clinically diagnosed female participants. The authors collected chat logs, pre/post surveys (EDE-Q, Brief-IPQ), and semi-structured interviews, and analyzed them using descriptive statistics and thematic analysis. They argue that the chatbot created a "private yet social" space that supported storytelling, self-reflection, and recovery motivation, but also produced harmful, ED-inappropriate responses (Table 4) that went unnoticed because participants trusted the bot. The paper presents the Brief-IPQ pre-post improvement as supporting evidence and closes with design implications for safe LLM-based ED interventions.
Significance. If the harm findings are valid, this is a significant contribution: it provides field evidence about how vulnerable users interact with LLM-based chatbots and why safety warnings may fail. The study's strengths are its rich qualitative data, the transparent description of the WellnessBot design and deployment, the daily monitoring protocol, and the direct inclusion of chat-log excerpts and participant quotes. The paper is also honest about several limitations in Section 7.4. However, the headline "unnoticed harms" claim rests on the authors' own harm classifications, which are not clinically validated, and the quantitative support is exploratory rather than confirmatory. The design implications in Section 7 are reasonable but depend on the safety framing, so the contribution's weight currently falls on an unvalidated taxonomy.
major comments (3)
- [Section 6.3, Table 4; Section 5.3.2] The central claim that harmful chatbot responses "went unnoticed" is load-bearing, but the harm labels are assigned by the first three authors without clinical validation or inter-rater reliability. The coding process is described in Section 5.3.2 as thematic analysis by the authors, and Section 6.3.2 reports that participants themselves perceived no harm. Several Table 4 classifications are contestable: Chat 11 labels "extreme hunger" as unsupported advice even though the paper's own references [11, 28] describe extreme hunger as a real phenomenon in ED recovery; Chat 9's "picky eating" advice is framed as harmful without clinical context; and Chat 7's praise of weight loss is ambiguous without knowing the user's BMI or treatment goals. The headline finding therefore depends on a lay coding judgment. The authors should either have the harm taxonomy reviewed and validated by clinicians, report inter-rater reliability, or explicitly reframe the finding as "potentially concerning responses" rather than "harms that went unnoticed." Section 7.4 lists several limitations but omits this clinical-validation gap.
- [Section 6.4] The Brief-IPQ analysis is presented as evidence that the intervention improved participants' perceptions ("statistically significant decrease... Z = 2.43, p = .02"), but there is no control or comparison condition, so regression to the mean, repeated-testing effects, or general study participation could explain the change. In addition, the paper reports significance for the overall scale and for two individual items without correcting for multiple comparisons, and the item-level p-values (p = .02 and p = .05) are marginal. This analysis should be explicitly framed as an exploratory descriptive result, not as an effectiveness claim, and the Discussion in Section 7.1 should not cite the Brief-IPQ decrease as evidence for the benefits of storytelling.
- [Section 3 and Section 6.3.2] The "unnoticed" finding is conditioned on the study protocol: Section 3 states that the authors deliberately did not intervene on misinformation or undesirable responses and only informed participants during post-interviews. This means the absence of participant questioning was observed under a protocol that withheld feedback, which is a valid observational choice but should be stated as a boundary of the claim. As written, the abstract and Section 6.3.2 imply a general property of user trust, whereas the evidence shows how users behave when no in-situ correction is provided. The design implications in Section 7.3 recommend in-situ interventions, but the paper should acknowledge that its own protocol precluded such interventions and therefore cannot directly test whether they would change user awareness.
minor comments (5)
- [Reference [24]] Reference [24] misspells "ChatGPT" as "ChagGPT"; please correct the typo.
- [Table 4, Chat 9] The user message in Chat 9 contains "When I've eating a lot of meat," which appears to be a typo in the translated transcript; please correct it or mark it as [sic].
- [Figure 2] Figure 2 lacks a y-axis label and does not state whether the counts are raw message counts or per-user averages; please clarify in the caption or on the axis.
- [Section 5.2] The sentence "participants interact with WellnessBot without any instructions" seems to conflict with the described introductory session; please rephrase to indicate that no further instructions were given during the deployment period.
- [Section 6.1] The Mann-Kendall results are reported per user (21 non-significant, 5 decreasing), which involves multiple comparisons and is not a study-level aggregate; consider reporting a single mixed-effects model or simply descriptive trends instead.
Circularity Check
No significant circularity; empirical field study whose findings derive from observed chat logs, interviews, and independent survey instruments.
full rationale
This paper is a technology-probe field study, not a derivation with equations or fitted parameters, so the circularity patterns based on 'prediction equals fit' or 'uniqueness theorem imported from authors' do not apply. WellnessBot is an external system (GPT-4 API) rather than a model constructed by the authors to reproduce a target result. The central findings are qualitative: participants reported empowerment, and researchers identified potentially harmful responses from the chat logs. The 'harms went unnoticed' claim depends on the researchers' own classification of Table 4 responses, which is a validity or construct-concern (clinical validation, inter-rater reliability) rather than a circularity, because the harm labels are not defined in terms of participants' reports nor fitted to any outcome; indeed the paper explicitly notes that participants reported no harm, so the research inference is not equivalent to its input. The Brief-IPQ analyses use a standard independent instrument, and the lack of a control group limits causal inference but does not make any result circular. The only self-citations are background or related-work citations (e.g., the authors' prior work on food content moderation and stress topics), and none is load-bearing for the paper's main claims. No quoted step reduces to its own input by construction, so the appropriate finding is no significant circularity with score 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Self-identification as having an eating disorder is a valid inclusion criterion.
- domain assumption The EDE-Q scores confirm the sample represents the ED population.
- domain assumption GPT-4's behavior in WellnessBot is representative of current LLM chatbots.
- domain assumption Thematic analysis coding by the authors is a reliable measure of benefits and harms.
- standard math Brief-IPQ is a valid measure of illness perception change.
Cite this review
Pith. "Pith review of Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery." pith.science (2026). https://pith.science/paper/OSOJ3UTU
@misc{pith2026241211656,
author = {Pith},
title = {Pith review of: Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery},
year = {2026},
howpublished = {\url{https://pith.science/paper/OSOJ3UTU}},
note = {Machine review of arXiv:2412.11656}
}
read the original abstract
Eating disorders (ED) are complex mental health conditions that require long-term management and support. Recent advancements in large language model (LLM)-based chatbots offer the potential to assist individuals in receiving immediate support. Yet, concerns remain about their reliability and safety in sensitive contexts such as ED. We explore the opportunities and potential harms of using LLM-based chatbots for ED recovery. We observe the interactions between 26 participants with ED and an LLM-based chatbot, WellnessBot, designed to support ED recovery, over 10 days. We discovered that our participants have felt empowered in recovery by discussing ED-related stories with the chatbot, which served as a personal yet social avenue. However, we also identified harmful chatbot responses, especially concerning individuals with ED, that went unnoticed partly due to participants' unquestioning trust in the chatbot's reliability. Based on these findings, we provide design implications for safe and effective LLM-based interventions in ED management.
Figures
Reference graph
Works this paper leans on
-
[1]
Jiska J Aardoom, Alexandra E Dingemans, Margarita CT Slof Op’t Landt, and Eric F Van Furth. 2012. Norms and discriminative validity of the Eating Disorder Examination Questionnaire (EDE-Q). Eating behaviors 13, 4 (2012), 305–309
2012
-
[2]
Jonathan M Adler. 2012. Living into the story: agency and coherence in a longitudinal study of narrative identity development and mental health over the course of psychotherapy. Journal of personality and social psychology 102, 2 (2012), 367
2012
-
[3]
Fahad Alanezi. 2024. Assessing the effectiveness of ChatGPT in delivering men- tal health support: a qualitative study. Journal of Multidisciplinary Healthcare (2024), 461–471
2024
-
[4]
Anna M Bardone-Cone, Katrina Sturm, Melissa A Lawson, D Paul Robinson, and Roma Smith. 2010. Perfectionism across stages of recovery from eating disorders. International journal of eating disorders 43, 2 (2010), 139–148
2010
-
[5]
Kristian González Barman, Nathan Wood, and Pawel Pawlowski. 2024. Beyond transparency and explainability: on the need for adequate and contextualized user guidelines for LLM use. Ethics and Information Technology 26, 3 (2024), 47
2024
-
[6]
Francesca Beilharz, Suku Sukunesan, Susan L Rossell, Jayashri Kulkarni, Gemma Sharp, et al. 2021. Development of a positive body image chatbot (KIT) with young people and parents/carers: qualitative focus group study. Journal of medical Internet research 23, 6 (2021), e27807
2021
-
[7]
Jennifer Beveridge, Andrea Phillipou, Zoe Jenkins, Richard Newton, Leah Bren- nan, Freya Hanly, Benjamin Torrens-Witherow, Narelle Warren, Kelly Edwards, and David Castle. 2019. Peer mentoring for eating disorders: results from the evaluation of a pilot program. Journal of Eating Disorders 7 (2019), 1–10
2019
-
[8]
Alexei A Birkun and Adhish Gautam. 2023. Large Language Model (LLM)- powered chatbots fail to generate guideline-consistent content on resuscitation and may provide potentially harmful advice. Prehospital and Disaster Medicine 38, 6 (2023), 757–763
2023
Show all 151 references
-
[9]
Virginia Braun and Victoria Clarke. 2012. Thematic analysis. American Psycho- logical Association
2012
-
[10]
Elizabeth Broadbent, Keith J Petrie, Jodie Main, and John Weinman. 2006. The brief illness perception questionnaire. Journal of psychosomatic research 60, 6 (2006), 631–637
2006
-
[11]
Christine Byrne. 2024. How to Deal With Extreme Hunger in Eating Disor- der Recovery. https://rubyoaknutrition.com/extreme-hunger-eating-disorder- recovery. Accessed: December 17, 2024
2024
-
[12]
Fary M Cachelin, Ramona Rebeck, Catherine Veisel, and Ruth H Striegel-Moore
-
[13]
Carl S Carlson. 2012. Effective FMEAs: Achieving safe, reliable, and economical products and processes using failure mode and effects analysis . Vol. 1. John Wiley & Sons
2012
-
[14]
William W Chan, Ellen E Fitzsimmons-Craft, Arielle C Smith, Marie-Laure Firebaugh, Lauren A Fowler, Bianca DePietro, Naira Topooco, Denise E Wilfley, C Barr Taylor, and Nicholas C Jacobson. 2022. The challenges in designing a prevention chatbot for eating disorders: observatio...
2022
-
[15]
Stevie Chancellor, Yannis Kalantidis, Jessica A Pater, Munmun De Choudhury, and David A Shamma. 2017. Multimodal classification of moderated online pro-eating disorder content. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems . 3213–3226
2017
-
[16]
Stevie Chancellor, Zhiyuan Lin, Erica L Goodman, Stephanie Zerwas, and Mun- mun De Choudhury. 2016. Quantifying and predicting mental illness severity in online pro-eating disorder communities. In Proceedings of the 19th ACM con- ference on computer-supported cooperative work ...
2016
-
[17]
Stevie Chancellor, Jessica Annette Pater, Trustin Clear, Eric Gilbert, and Munmun De Choudhury. 2016. # thyghgapp: Instagram content moderation and lexical variation in pro-eating disorder communities. In Proceedings of the 19th ACM conference on computer-supported cooperative...
2016
-
[18]
Siyuan Chen, Mengyue Wu, Kenny Q Zhu, Kunyao Lan, Zhiling Zhang, and Lyuchun Cui. 2023. LLM-empowered chatbots for psychiatrist and patient simulation: application and evaluation. arXiv preprint arXiv:2305.13614 (2023)
2023 arXiv
-
[19]
Dasom Choi, Sunok Lee, Sung-In Kim, Kyungah Lee, Hee Jeong Yoo, Sangsu Lee, and Hwajung Hong. 2024. Unlock Life with a Chat (GPT): Integrating Conversational AI with Large Language Models into Everyday Lives of Autistic Individuals. In Proceedings of the CHI Conference on Huma...
2024
-
[20]
Ryuhaerang Choi, Subin Park, Sujin Han, and Sung-Ju Lee. 2024. FoodCensor: Promoting Mindful Digital Food Content Consumption for People with Eating Disorders. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18
2024
-
[21]
Ryuhaerang Choi, Chanwoo Yun, Hyunsung Cho, Hwajung Hong, Uichin Lee, and Sung-Ju Lee. 2022. You Are Not Alone: How Trending Stress Topics Brought# Awareness and# Resonance on Campus. Proceedings of the ACM on Human- Computer Interaction 6, CSCW2 (2022), 1–30
2022
-
[22]
Chia-Fang Chung, Kristin Dew, Allison Cole, Jasmine Zia, James Fogarty, Julie A Kientz, and Sean A Munson. 2016. Boundary negotiating artifacts in personal informatics: patient-provider collaboration with patient-generated data. In Pro- ceedings of the 19th ACM conference on c...
2016
-
[23]
2022 Claremont. 2023. Tess AI Chatbot. https://www.claremonteap.com/ employees-and-families/tess-ai-chatbot/. Accessed: December 17, 2024
2022
-
[24]
CNET. 2024. Apple Partners With OpenAI for ChagGPT on iPhones, iPads and Macs. https://www.cnet.com/tech/mobile/apple-partners-with-openai-for- chatgpt-on-iphones-ipads-and-macs/. Accessed: December 17, 2024
2024
-
[25]
CNN. 2023. National Eating Disorders Association takes its AI chatbot offline after complaints of ‘harmful’ advice. https://edition.cnn.com/2023/06/01/tech/ eating-disorder-chatbot/index.html. Accessed: December 17, 2024
2023
-
[26]
Simon Coghlan, Kobi Leins, Susie Sheldrick, Marc Cheong, Piers Gooding, and Simon D’Alfonso. 2023. To chat or bot to chat: Ethical issues with using chatbots in mental health. Digital health 9 (2023), 20552076231183542
2023
-
[27]
Naver cop. 2018. South Korean Online Social Support Community for People with Eating Disorders. https://cafe.naver.com/jahayun. Accessed: December 17, 2024
2018
-
[28]
Sophie Corbett. 2023. Extreme Hunger in ED Recovery. https://www. mentalhealthdietitians.com/extreme-hunger-in-ed-recovery/. Accessed: De- cember 17, 2024
2023
-
[29]
Marguerite Corvini, Casey Golomski, and John Burns. 2024. The Impact of Sharing Recovery Stories in Public: Stigma, Trauma Response, and the Need for Multiple Pathways. Journal of Social Service Research 50, 3 (2024), 481–493
2024
-
[30]
Jennifer Couturier, Melissa Kimber, and Peter Szatmari. 2013. Efficacy of family- based treatment for adolescents with eating disorders: A systematic review and meta-analysis. International Journal of Eating Disorders 46, 1 (2013), 3–11
2013
-
[31]
Michele L Crossley. 2000. Narrative psychology, trauma and the study of self/identity. Theory & Psychology 10, 4 (2000), 527–546
2000
-
[32]
Munmun De Choudhury, Michael Gamon, Scott Counts, and Eric Horvitz. 2013. Predicting depression via social media. In Proceedings of the international AAAI conference on web and social media , Vol. 7. 128–137
2013
-
[33]
Julian De Freitas, Ahmet Kaan Uğuralp, Zeliha Oğuz-Uğuralp, and Stefano Puntoni. 2024. Chatbots and mental health: Insights into the safety of generative AI. Journal of Consumer Psychology 34, 3 (2024), 481–491
2024
-
[34]
SM De la Rie, G Noordenbos, and EF Van Furth. 2005. Quality of life and eating disorders. Quality of life research 14 (2005), 1511–1521
2005
-
[35]
Kerstin Denecke, Alaa Abd-Alrazaq, and Mowafa Househ. 2021. Artificial intelligence for chatbots in mental health: opportunities and challenges.Multiple perspectives on artificial intelligence in healthcare: Opportunities and challenges (2021), 115–128
2021
-
[36]
Anjali Devakumar, Jay Modh, Bahador Saket, Eric PS Baumer, and Munmun De Choudhury. 2021. A review on strategies for data collection, reflection, and communication in eating disorder apps. InProceedings of the 2021 CHI conference on human factors in computing systems . 1–19
2021
-
[37]
Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[38]
AE Dingemans, MJ Bruna, and EF Van Furth. 2002. Binge eating disorder: a review. International journal of obesity 26, 3 (2002), 299–307
2002
-
[39]
eternnoir. 2024. pyTelegramBotAPI 4.22.1. https://pypi.org/project/ pyTelegramBotAPI. Accessed: December 17, 2024
2024
-
[40]
Christopher G Fairburn and Zafra Cooper. 2011. Eating disorders, DSM–5 and clinical reality. The British journal of psychiatry 198, 1 (2011), 8–10
2011
-
[41]
Christopher G Fairburn, Zafra Cooper, Helen A Doll, Patricia Norman, and Marianne O’Connor. 2000. The natural course of bulimia nervosa and binge eating disorder in young women. Archives of general psychiatry 57, 7 (2000), 659–665
2000
-
[42]
Christopher G Fairburn and David M Garner. 1986. The diagnosis of bulimia nervosa. International Journal of Eating Disorders 5, 3 (1986), 403–419
1986
-
[43]
Jessica L Feuston, Michael Ann DeVito, Morgan Klaus Scheuerman, Katy Weath- ington, Marianna Benitez, Bianca Z Perez, Lucy Sondheim, and Jed R Brubaker
-
[44]
Ellen E Fitzsimmons-Craft, William W Chan, Arielle C Smith, Marie-Laure Firebaugh, Lauren A Fowler, Naira Topooco, Bianca DePietro, Denise E Wilfley, C Barr Taylor, and Nicholas C Jacobson. 2022. Effectiveness of a chatbot for eating disorders prevention: a randomized clinical...
2022
-
[45]
Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Ver- ena Rieser, Hasan Iqbal, Nenad Tomašev, Ira Ktena, Zachary Kenton, Mikel Rodriguez, et al . 2024. The ethics of advanced ai assistants. arXiv preprint arXiv:2404.16244 (2024)
2024 arXiv
-
[46]
Saadia Gabriel, Isha Puri, Xuhai Xu, Matteo Malgaroli, and Marzyeh Ghassemi
-
[47]
Liza Gak, Seyi Olojo, and Niloufar Salehi. 2022. The distressing ads that persist: Uncovering the harms of targeted weight-loss ads among users with histories of disordered eating. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–23
2022
-
[48]
Marie Galmiche, Pierre Déchelotte, Grégory Lambert, and Marie Pierre Tavolacci
-
[49]
David M Garner and Paul E Garfinkel. 1997. Handbook of treatment for eating disorders. Guilford Press
1997
-
[50]
Suyu Ge, Chunting Zhou, Rui Hou, Madian Khabsa, Yi-Chia Wang, Qifan Wang, Jiawei Han, and Yuning Mao. 2023. Mart: Improving llm safety with multi-round automatic red-teaming. arXiv preprint arXiv:2311.07689 (2023)
2023 arXiv
-
[51]
Scott Glassman, Petra Kottsieper, Allan Zuckoff, and Elizabeth A. Gosch. 2013. Motivational interviewing and recovery: experiences of hope, meaning, and empowerment. Advances in dual diagnosis 6, 3 (2013), 106–120
2013
-
[52]
Garth Graham. 2023. An updated approach to eating disorder-related con- tent. https://blog.youtube/news-and-events/an-updated-approach-to-eating- disorder-related-content/. Accessed: December 17, 2024
2023
-
[53]
Nick Grey, Suzanne Byrne, Tracey Taylor, Avi Shmueli, Cathy Troupp, Peter Stratton, Aaron Sefi, Roslyn Law, and Mick Cooper. 2018. Goal-oriented practice across therapies. (2018)
2018
-
[54]
Samantha L Hahn, Katherine W Bauer, Niko Kaciroti, Daniel Eisenberg, Sarah K Lipson, and Kendrin R Sonneville. 2021. Relationships between patterns of weight-related self-monitoring and eating disorder symptomology among un- dergraduate and graduate students. International Jou...
2021
-
[55]
Katrin Hartwig, Tom Biselli, Franziska Schneider, and Christian Reuter. 2024. From Adolescents’ Eyes: Assessing an Indicator-Based Intervention to Combat Misinformation on TikTok. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
2024
-
[56]
2023 WoeBot Health. 2023. WoeBot Health. https://woebothealth.com. Accessed: December 17, 2024
2023
-
[57]
Woebot Health. 2024. Woebot for Adults. Instructions for Use. https://woebothealth.com/img/2024/07/Woebot-for-Adults_Instructions- for-Use_Users_July-18th-2024.pdf. Accessed: 2024-11-27
2024
-
[58]
Sharon Hillege, Barbara Beale, and Rose McMaster. 2006. Impact of eating disorders on family life: Individual parents’ stories. Journal of Clinical Nursing 15, 8 (2006), 1016–1022
2006
-
[59]
Tiancheng Hu and Nigel Collier. 2024. Quantifying the persona effect in llm simulations. arXiv preprint arXiv:2402.10811 (2024)
2024 arXiv
-
[60]
Wenyue Hua, Xianjun Yang, Mingyu Jin, Zelong Li, Wei Cheng, Ruixiang Tang, and Yongfeng Zhang. 2024. TrustAgent: Towards Safe and Trustworthy LLM- based Agents. InFindings of the Association for Computational Linguistics: EMNLP
2024
-
[61]
See your doctor
Jina Huh. 2015. Clinical questions in online health communities: the case of" See your doctor" threads. In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work & Social Computing . 1488–1499
2015
-
[62]
Martin Huschens, Martin Briesch, Dominik Sobania, and Franz Rothlauf. 2023. Do You Trust ChatGPT?–Perceived Credibility of Human and AI-Generated Content. arXiv preprint arXiv:2309.02524 (2023)
2023 arXiv
-
[63]
Hilary Hutchinson, Wendy Mackay, Bo Westerlund, Benjamin B Bederson, Allison Druin, Catherine Plaisant, Michel Beaudouin-Lafon, Stéphane Conversy, Helen Evans, Heiko Hansen, et al. 2003. Technology probes: inspiring design for and with families. In Proceedings of the SIGCHI co...
2003
-
[64]
Farnaz Jahanbakhsh and David R Karger. 2024. A Browser Extension for in- place Signaling and Assessment of Misinformation. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–21
2024
-
[65]
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung
-
[66]
Eunkyung Jo, Yuin Jeong, SoHyun Park, Daniel A Epstein, and Young-Ho Kim
-
[67]
Douglas Johnson, Rachel Goodman, J Patrinely, Cosby Stone, Eli Zimmerman, Rebecca Donald, Sam Chang, Sean Berkowitz, Avni Finn, Eiman Jahangir, et al
-
[68]
KakaoTalk. 2022. KakaoTalk Online Social Support Chatroom for People with Eating Disorders. https://open.kakao.com/o/guHeu70b. Accessed: 2023-07-15
2022
-
[69]
Jess Kerr-Gaffney, Amy Harrison, and Kate Tchanturia. 2018. Social anxiety in the eating disorders: a systematic review and meta-analysis. Psychological medicine 48, 15 (2018), 2477–2491
2018
-
[70]
Ronald C Kessler, Patricia A Berglund, Wai Tat Chiu, Anne C Deitz, James I Hudson, Victoria Shahly, Sergio Aguilar-Gaxiola, Jordi Alonso, Matthias C Angermeyer, Corina Benjet, et al. 2013. The prevalence and correlates of binge eating disorder in the World Health Organization ...
2013
-
[71]
Zoha Khawaja and Jean-Christophe Bélisle-Pipon. 2023. Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots. Frontiers in Digital Health 5 (2023), 1278186
2023
-
[72]
In Proceedings of the CHI Conference on Human Factors in Computing Systems
Understanding the Impact of Long-Term Memory on Self-Disclosure with Large Language Model-Driven Chatbots for Public Health Intervention. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–21
-
[73]
Taewan Kim, Seolyeong Bae, Hyun Ah Kim, Su-woo Lee, Hwajung Hong, Chanmo Yang, and Young-Ho Kim. 2024. MindfulDiary: Harnessing Large Language Model to Support Psychiatric Patients’ Journaling. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–20
2024
-
[74]
Research square (2023)
Assessing the accuracy and reliability of AI-generated medical responses: an evaluation of the Chat-GPT model. Research square (2023)
2023
-
[75]
Zeljko Kraljevic, Anthony Shek, Daniel Bean, Rebecca Bendayan, James Teo, and Richard Dobson. 2021. MedGPT: Medical concept prediction from clinical narratives. arXiv preprint arXiv:2107.03134 (2021)
2021 arXiv
-
[76]
Harsh Kumar, Ruiwei Xiao, Benjamin Lawson, Ilya Musabirov, Jiakai Shi, Xinyuan Wang, Huayin Luo, Joseph Jay Williams, Anna N Rafferty, John Stam- per, et al. 2024. Supporting Self-Reflection at Scale with Large Language Models: Insights from Randomized Field Experiments in Cla...
2024
-
[78]
Tin Lai, Yukun Shi, Zicong Du, Jiajie Wu, Ken Fu, Yichao Dou, and Ziqi Wang
-
[79]
Jennifer G Kim, Hwajung Hong, and Karrie Karahalios. 2018. Understanding identity presentation in medical crowdfunding. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–12
2018
-
[80]
Minha Lee, Sander Ackermans, Nena Van As, Hanwen Chang, Enzo Lucas, and Wijnand IJsselsteijn. 2019. Caring for Vincent: a chatbot for self-compassion. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–13
2019
-
[81]
René F Kizilcec. 2016. How much information? Effects of transparency on trust in an algorithmic interface. In Proceedings of the 2016 CHI conference on human factors in computing systems . 2390–2395
2016
-
[82]
Nancy G Leveson. 2016. Engineering a safer world: Systems thinking applied to safety. The MIT Press
2016
-
[83]
Q Vera Liao and Jennifer Wortman Vaughan. 2023. Ai transparency in the age of llms: A human-centered research roadmap. arXiv preprint arXiv:2306.01941 (2023), 5368–5393
2023 arXiv
-
[84]
José Francisco López-Gil, Antonio Garcia-Hermoso, Lee Smith, Joseph Firth, Mike Trott, Arthur Eumann Mesas, Estela Jimenez-Lopez, Hector Gutierrez- Espinoza, Pedro J Tarraga-Lopez, and Desiree Victoria-Montesinos. 2023. Global proportion of disordered eating in children and ad...
2023
-
[85]
arXiv preprint arXiv:2307.11991 (2023)
Psy-llm: Scaling up global mental health psychological services with ai-based large language models. arXiv preprint arXiv:2307.11991 (2023)
2023 arXiv
-
[86]
Ryan Lowe, Nissan Pow, Iulian Vlad Serban, Laurent Charlin, Chia-Wei Liu, and Joelle Pineau. 2017. Training end-to-end dialogue systems with the ubuntu dialogue corpus. Dialogue & Discourse 8, 1 (2017), 31–65
2017
-
[87]
BioMedInformatics 4, 1 (2023), 8–33
Supporting the Demand on Mental Health Services with AI-Based Con- versational Large Language Models (LLMs). BioMedInformatics 4, 1 (2023), 8–33
2023
-
[88]
Andrea LaMarre and Carla Rice. 2021. Healthcare providers’ engagement with eating disorder recovery narratives: opening to complexity and diversity. Medi- cal Humanities 47, 1 (2021), 78–86
2021
-
[89]
Zilin Ma, Yiyang Mei, Yinru Long, Zhaoyuan Su, and Krzysztof Z Gajos. 2024. Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health Support. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15
2024
-
[90]
I Hear You, I Feel You
Yi-Chieh Lee, Naomi Yamashita, Yun Huang, and Wai Fu. 2020. " I Hear You, I Feel You": encouraging deep self-disclosure through a chatbot. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–12
2020
-
[91]
Lena Mamykina, Andrew D Miller, Elizabeth D Mynatt, and Daniel Greenblatt
-
[92]
Mark Matthews and Gavin Doherty. 2011. My mobile story: therapeutic story- telling for children. InCHI’11 Extended Abstracts on Human Factors in Computing Systems. 2059–2064. Private Yet Social: How LLM Chatbots Support and Challenge Eating Disorder Recovery Conference’17, Jul...
2011
-
[93]
Dan P McAdams. 1993. The stories we live by: Personal myths and the making of the self. William Morrow (1993)
1993
-
[94]
I felt her company
Kate Loveys, Catherine Hiko, Mark Sagar, Xueyuan Zhang, and Elizabeth Broad- bent. 2022. “I felt her company”: A qualitative study on factors affecting close- ness and emotional support seeking with an embodied conversational agent. International Journal of Human-Computer Stud...
2022
-
[95]
Scott Monteith, Tasha Glenn, John R Geddes, Peter C Whybrow, Eric Achtyes, and Michael Bauer. 2024. Artificial intelligence and increasing misinformation. The British Journal of Psychiatry 224, 2 (2024), 33–35
2024
-
[96]
Wysa Ltd. 2023. Wysa. https://www.wysa.com. Accessed: December 17, 2024
2023
-
[97]
Kai Lukoff, Taoxi Li, Yuan Zhuang, and Brian Y Lim. 2018. TableChat: mobile food journaling to facilitate family support for healthy eating. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–28
2018
-
[98]
Marcia Nißen, Dominik Rüegger, Mirjam Stieger, Christoph Flückiger, Mathias Allemand, Florian v Wangenheim, and Tobias Kowatsch. 2022. The effects of health care chatbot personas with different social roles on the client-chatbot bond and usage intentions: development of a desi...
2022
-
[99]
Zilin Ma, Yiyang Mei, and Zhaoyuan Su. 2023. Understanding the benefits and challenges of using large language model-based conversational agents for mental well-being support. In AMIA Annual Symposium Proceedings , Vol. 2023. American Medical Informatics Association, 1105
2023
-
[100]
Jooyoung Oh, Sooah Jang, Hyunji Kim, and Jae-Jin Kim. 2020. Efficacy of mobile app-based interactive cognitive behavioral therapy using a chatbot for panic disorder. International journal of medical informatics 140 (2020), 104171
2020
-
[101]
Yoo Jung Oh, Jingwen Zhang, Min-Lin Fang, and Yoshimi Fukuoka. 2021. A sys- tematic review of artificial intelligence chatbots for promoting physical activity, healthy diet, and weight loss. International Journal of Behavioral Nutrition and Physical Activity 18 (2021), 1–25
2021
-
[102]
OpenAI. 2024. OpenAI Privacy Policy. https://openai.com/policies/row-privacy- policy. Accessed: December 17, 2024
2024
-
[103]
OpenAI. 2024. OpenAI Usage Policies. https://openai.com/policies/usage- policies/. Accessed: 2024-11-27
2024
-
[104]
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishin- skaya, Maja Trebacz, and Jan Leike. 2024. Llm critics help catch llm bugs. arXiv preprint arXiv:2407.00215 (2024)
2024 arXiv
-
[105]
Krisna Patel, Kate Tchanturia, and Amy Harrison. 2016. An exploration of social functioning in young people with eating disorders: a qualitative study. PloS one 11, 7 (2016), e0159910
2016
-
[106]
Subigya Nepal, Arvind Pillai, William Campbell, Talie Massachi, Eunsol Soul Choi, Xuhai Xu, Joanna Kuc, Jeremy F Huckins, Jason Holden, Colin Depp, et al
-
[107]
In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems
Contextual AI Journaling: Integrating LLM and Time Series Behavioral Sensing Technology to Promote Self-Reflection and Well-being using the Mind- Scape App. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. 1–8
-
[108]
Jingping Nie, Hanya Shao, Yuang Fan, Qijia Shao, Haoxuan You, Matthias Preindl, and Xiaofan Jiang. 2024. LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices. arXiv preprint arXiv:2403.10779 (2024)
2024 arXiv
-
[109]
Janet Polivy and C Peter Herman. 2004. Sociocultural idealization of thin female body shapes: An introduction to the special issue on body image and eating disorders. , 6 pages
2004
-
[110]
Kate P Nurser, Imogen Rushworth, Tom Shakespeare, and Deirdre Williams
-
[111]
Rebecca Puhl and Young Suh. 2015. Stigma and eating and weight disorders. Current psychiatry reports 17 (2015), 1–10
2015
-
[112]
Taivo Pungas. 2023. GPT-3.5 and GPT-4 response times. https://www.taivo.ai/ __gpt-3-5-and-gpt-4-response-times/. Accessed: December 17, 2024
2023
-
[113]
Matthew Renze and Erhan Guven. 2024. Self-Reflection in LLM Agents: Effects on Problem-Solving Performance. arXiv preprint arXiv:2405.06682 (2024)
2024 arXiv
-
[114]
MIT Technology Review. 2024. A chatbot helped more people access mental- health services. https://www.technologyreview.com/2024/02/05/1087690/a- chatbot-helped-more-people-access-mental-health-services/. Accessed: De- cember 17, 2024
2024
-
[115]
Deborah Richards, Ayse Aysin Bilgin, and Hedieh Ranjbartabar. 2018. Users’ perceptions of empathic dialogue cues: A data-driven approach to provide tailored empathy. InProceedings of the 18th International Conference on Intelligent Virtual Agents. 35–42
2018
-
[116]
P Parmar, J Ryu, S Pandya, J Sedoc, and S Agarwal. 2022. Health-focused conversational agents in person-centered care: A review of apps. npj Digital Medicine 5, 21
2022
-
[117]
S Roller. 2020. Recipes for building an open-domain chatbot. arXiv preprint arXiv:2004.13637 (2020)
2020 arXiv
-
[118]
Jessica Pater, Fayika Farhat Nova, Amanda Coupe, Lauren E Reining, Connie Kerrigan, Tammy Toscos, and Elizabeth D Mynatt. 2021. Charting the unknown: Challenges in the clinical assessment of patients’ technology use related to eating disorders. In Proceedings of the 2021 CHI c...
2021
-
[119]
Manjiri Pawaskar, Edward A Witt, Dylan Supina, Barry K Herman, and Thomas A Wadden. 2017. Impact of binge eating disorder on functional impair- ment and work productivity in an adult community sample in the United States. International journal of clinical practice 71, 7 (2017), e12970
2017
-
[120]
Janet Polivy and C Peter Herman. 2002. Causes of eating disorders. Annual review of psychology 53, 1 (2002), 187–213
2002
-
[121]
Rose Stackpole, Danyelle Greene, Elizabeth Bills, and Sarah J Egan. 2023. The association between eating disorders and perfectionism in adults: A systematic review and meta-analysis. Eating behaviors 50 (2023), 101769
2023
-
[122]
Katie Prizeman, Netta Weinstein, and Ciara McCabe. 2023. Effects of mental health stigma on loneliness, social isolation, and relationships in young people with depression symptoms. BMC psychiatry 23, 1 (2023), 527
2023
-
[123]
Richard Sutcliffe. 2023. A Survey of Personality, Persona, and Profile in Conver- sational Agents and Chatbots. arXiv preprint arXiv:2401.00609 (2023)
2023 arXiv
-
[124]
I Sutskever. 2014. Sequence to Sequence Learning with Neural Networks. arXiv preprint arXiv:1409.3215 (2014)
2014 arXiv
-
[125]
Telegram FZ LLC and Telegram Messenger Inc. [n. d.]. Telegram. https:// telegram.org
-
[126]
The New York Times. 2023. A Wellness Chatbot Is Offline After Its ’Harmful’ Focus on Weight Loss. https://www.nytimes.com/2023/06/08/us/ai-chatbot- tessa-eating-disorders-association.html. Accessed: December 17, 2024
2023
-
[127]
Tracy L Tylka and Ashley M Kroon Van Diest. 2013. The Intuitive Eating Scale–2: Item refinement and psychometric evaluation with college women and men. Journal of counseling psychology 60, 1 (2013), 137
2013
-
[128]
Shalaleh Rismani, Renee Shelby, Andrew Smart, Edgar Jatho, Joshua Kroll, AJung Moon, and Negar Rostamzadeh. 2023. From plane crashes to algorithmic harm: applicability of safety engineering frameworks for responsible ML. In Proceedings of the 2023 CHI Conference on Human Facto...
2023
-
[129]
Tara Wadhwa. 2021. Supporting #NEDAwareness and body inclusivity on TikTok. https://newsroom.tiktok.com/en-us/supporting-nedawareness-and- body-inclusivity-on-tiktok. Accessed: December 17, 2024
2021
-
[130]
Sivan Schwartz, Avi Yaeli, and Segev Shlomov. 2023. Enhancing trust in LLM- based AI automation agents: New considerations and future challenges. arXiv preprint arXiv:2308.05391 (2023)
2023 arXiv
-
[131]
Jillian Shah, Bianca DePietro, Laura D’Adamo, Marie-Laure Firebaugh, Olivia Laing, Lauren A Fowler, Lauren Smolar, Shiri Sadeh-Sharvit, C Barr Taylor, Denise E Wilfley, et al. 2022. Development and usability testing of a chatbot to promote mental health services use among indi...
2022
-
[132]
Inhwa Song, Sachin R Pendse, Neha Kumar, and Munmun De Choudhury. 2024. The typing cure: Experiences with large language model chatbots for mental health support. arXiv preprint arXiv:2401.14362 (2024)
2024 arXiv
-
[133]
Thomas Wolf, Victor Sanh, Julien Chaumond, and Clement Delangue. 2019. Transfertransfo: A transfer learning approach for neural network based conver- sational agents. arXiv preprint arXiv:1901.08149 (2019)
2019 arXiv
-
[134]
Sharon Stovezky. 2021. You are not alone. https://blog.youtube/news-and- events/you-are-not-alone/. Accessed: December 17, 2024
2021
-
[135]
Xuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel, Hong Yu, James Hendler, Marzyeh Ghassemi, Anind K Dey, and Dakuo Wang. 2024. Mental-llm: Leveraging large language models for mental health prediction via online text data. Proceedings of the ACM on Interactive, Mobile, We...
2024
-
[136]
Joel Yager and Pauline S Powers. 2008. Clinical manual of eating disorders . American Psychiatric Pub
2008
-
[137]
It’s a Fair Game
Zhiping Zhang, Michelle Jia, Hao-Ping Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024. “It’s a Fair Game”, or Is It? Examining How Users Navigate Disclosure Risks and Benefits When Using LLM-Based Conversational Agents. In Proceedings of the CHI Co...
2024
-
[138]
Yuxin Zhao, Jiawei Wu, Ping Qu, Beibei Zhang, and Hao Yan. 2024. Assessing User Trust in LLM-based Mental Health Applications: Perceptions of Reliability and Effectiveness. Journal of Computer Technology and Applied Mathematics 1, 2 (2024), 19–26
2024
-
[139]
Zhonghua Zheng, Lizi Liao, Yang Deng, and Liqiang Nie. 2023. Building emo- tional support chatbots in the era of llms.arXiv preprint arXiv:2308.11584 (2023)
2023 arXiv
-
[140]
A Vaswani. 2017. Attention is all you need. Advances in Neural Information Processing Systems (2017)
2017
-
[142]
B Timothy Walsh and Michael J Devlin. 1998. Eating disorders: progress and problems. Science 280, 5368 (1998), 1387–1390
1998
-
[143]
As an AI language model, I cannot
Joel Wester, Tim Schrills, Henning Pohl, and Niels van Berkel. 2024. “As an AI language model, I cannot”: Investigating LLM Denials of User Requests. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–14
2024
-
[144]
WinnieNwanne. 2024. Comparing GPT-3.5 and GPT-4: A Thought Framework on When To Use Each Model. https://techcommunity.microsoft.com/t5/ai-azure- ai-services-blog/comparing-gpt-3-5-amp-gpt-4-a-thought-framework-on- when-to-use/ba-p/4088645. Accessed: December 17, 2024
2024
-
[146]
Siyi Wu, Feixue Han, Bingsheng Yao, Tianyi Xie, Xuan Zhao, and Dakuo Wang
-
[147]
arXiv preprint arXiv:2405.13803 (2024)
Sunnie: An Anthropomorphic LLM-Based Conversational Agent for Mental Well-Being Activity Recommendation. arXiv preprint arXiv:2405.13803 (2024)
2024 arXiv
-
[2001]
International Journal of Eating Disorders 30, 3 (2001), 269–278
Barriers to treatment for eating disorders among ethnically diverse women. International Journal of Eating Disorders 30, 3 (2001), 269–278
2001
-
[2010]
In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems
Constructing identities through storytelling in diabetes management. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 1203–1212
-
[2018]
Mental Health Review Journal 23, 1 (2018), 25–36
Personal storytelling in mental health recovery. Mental Health Review Journal 23, 1 (2018), 25–36
2018
-
[2019]
The American journal of clinical nutrition 109, 5 (2019), 1402– 1413
Prevalence of eating disorders over the 2000–2018 period: a systematic literature review. The American journal of clinical nutrition 109, 5 (2019), 1402– 1413
2019
-
[2022]
Do You Ladies Relate?
" Do You Ladies Relate?": Experiences of Gender Diverse People in Online Eating Disorder Communities. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2 (2022), 1–32
2022
-
[2023]
In Findings of the Association for Computational Linguistics: EMNLP 2023
Towards mitigating LLM hallucination via self reflection. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 1827–1843
2023
-
[2024]
arXiv preprint arXiv:2405.12021 (2024)
Can AI Relate: Testing Large Language Model Response for Mental Health Support. arXiv preprint arXiv:2405.12021 (2024). Conference’17, July 2017, Washington, DC, USA Ryuhaerang Choi et al
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.