REVIEW 4 major objections 5 minor 45 references
An empathic GPT-based chatbot to talk about mental disorders with Spanish teenagers
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper reports a Telegram chatbot with a vulnerable-teen persona that uses self-disclosure to engage Spanish adolescents in discussions about mental disorders, and finds that most users who conversed opened up emotionally.
desk verdict A useful Spanish-language feasibility pilot for a teen mental-health chatbot; its self-disclosure causal claim is the one real overreach. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a two-layer dialogue engine: a controlled dialogue built from psychologist-designed triage questions and topic prompts, and an open dialogue using GPT-3, specifically the Davinci-002 model, with DeepL translating between Spanish and English. The bot's persona is itself the key instrument: it appears as a teenager named Ada, Hugo, or Big, lets the user choose the bot's gender, reveals personal worries, asks the user for advice, and reciprocates, following the disclosure layers of Social Penetration Theory. Every five user turns a specialist prompt steers the conversation back on track, and language suggesting self-harm or suicidal ideation triggers an alert to a human. This hybrid of controlled and open dialogue is what the paper credits for both conversational fluency and safety.
What would settle it
Assign teenagers at random to the self-disclosing bot versus a warm but emotionally neutral bot that asks identical questions, and count how many share personal concerns; if the neutral bot produces a similar disclosure rate, the claim that self-disclosure drives openness is refuted.
Extended reading notes
Core claim
The study's discovery is that a chatbot designed as a vulnerable, disclosing peer, rather than as a therapist or a neutral assistant, can draw Spanish teenagers into sustained and emotionally open conversations about mental health. Of the 44 users who moved past onboarding, 31 helped and advised the bot and 22 shared their own worries, and in all 22 cases empathetic engagement with the bot preceded the disclosure. The authors interpret this as evidence that self-disclosure is consistently effective, and they combine usage statistics, linguistic feature analysis, manual conversation review, and a user survey to support the view that the system was well received and potentially useful for raising awareness.
Load-bearing premise
The argument rests on the assumption that the bot's self-revelation is what causes teenagers to open up, yet the evidence only shows that in every chat the user's disclosure came after empathetic engagement with the bot, with no neutral comparison chatbot to rule out other causes.
Editorial extensions
If this is right
- If the finding holds, a freely available chatbot on a familiar messaging platform can serve as a low-stigma first step for teenagers to talk about depression, anxiety, eating disorders, and related topics in Spanish.
- A bot that reveals its own worries before asking about the user's can create an atmosphere in which many teenagers reciprocate with personal concerns, supporting self-disclosure as an engagement strategy rather than a purely scripted interview.
- Combining a controlled, psychologist-designed dialogue with an open GPT-3 conversation keeps the chat on topic while letting users drift toward what actually worries them, such as friendship, break-ups, and school, suggesting that rigid disorder-focused scripts miss the real content.
- The Spanish-to-English translation loop is workable, but colloquial speech and grammatical gender are recurring failure points, so native-Spanish models or better handling of gendered language would improve fluency and comfort.
- The system's risk-alert mechanism, together with the safe framing of topics under soft names, offers a template for how generative chatbots can address highly sensitive mental-health content with teenagers without impersonating a clinician.
Reading between the lines
- Editorial inference: the paper's 100 percent disclosure result shows temporal precedence, not causation; a randomized trial with a neutral comparison bot would be needed to prove that self-disclosure itself drives openness.
- Editorial inference: because the manual coding of engagement and openness was done by the authors without a second, independent rater, replicating the 70 percent emotional-engagement figure with inter-rater reliability checks would harden the main outcome.
- Editorial inference: the Spanish-English translation loop, despite its errors, suggests a reusable recipe for deploying English-centric large language models in lower-resource languages for sensitive dialogue, provided colloquialisms and gender agreement are handled.
- Editorial inference: the collected corpus of 1,860 messages and 94 linguistic features could feed early-detection models, but the deliberate anonymity that separates survey responses from chat data means user-satisfaction findings cannot yet be linked to clinical or triage status.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a field deployment of a GPT-3-based chatbot ('Ada', 'Hugo', 'Big') on Telegram for Spanish-speaking teenagers aged 12–18, mixing psychologist-authored controlled dialogue, triage questions, and open dialogue with Spanish↔English translation. Usage statistics from 44 active users, a post-hoc anonymous survey, NLP feature correlations, and a manual reading of all 44 chats are presented. The authors conclude that the self-disclosure persona led most users to engage emotionally, that 100% of users who disclosed concerns first helped or advised the bot, and hence that such systems interest young people and can raise awareness of mental disorders.
Significance. If the descriptive claims hold, the study contributes a concrete, privacy-preserving deployment with psychologist involvement, a Spanish-language corpus from adolescent interactions, and a design pattern (a vulnerable peer bot with selectable gender) that could inform future mental-health chatbots for youth. The paper's strengths include a real deployment over roughly three and a half months, ethics approval with parental consent, explicit acknowledgment of some limitations in Section 4.4, and enough architectural detail to permit approximate reproduction. There are no fitted models or formal derivations whose circularity could threaten the central result; the main risk is causal overreading of observational data rather than internal inconsistency in the system design.
major comments (4)
- [Section 4.3] The paper's distinctive contribution is the self-disclosure design, and the load-bearing sentence — 'the self-disclosure technique is consistently effective, as 100% of users expressing concerns have previously engaged in empathetic interactions with the chatbot' — is not supported by the evidence as reported. The observation is a temporal ordering in 44 chats that were read by the authors, with no published coding rubric, no inter-rater reliability check, and no comparison condition such as a warm bot that does not self-disclose. Longer conversations mechanically allow more opportunities both to advise the bot and to disclose concerns, so the ordering is compatible with a conversation-length confound. Please reframe this as an observational association, report a reliability analysis or blinded re-coding, and soften the corresponding claim in the abstract.
- [Sections 4.3 and 5] The numerical summary is internally inconsistent. The text reports 31/44 users 'care about the bot and get involved in advising and helping it' and only 22/44 'open up and talk about their concerns'; the statement that 'more than 70% of the users engaged emotionally with the bot, sharing their concerns and worries' conflates the 70% helping figure with the 50% disclosure figure, and Section 5 repeats this as '70% emotional openness.' Because the self-disclosure claim concerns the 22 users who disclosed, please report the two rates separately and correct the Discussion text.
- [Section 4.2] The NLP analysis computes Pearson correlations across 94 linguistic features using n=44 users and then interprets the top correlations (e.g., coordinating-conjunction frequency above 0.4) as evidence that 'there are differences in language between people classified as healthy and people classified as having some form of mental disorder.' With 94 features and no multiple-testing correction, top correlations of this size are expected under noise, and no confidence intervals or effect sizes are provided. Please label this analysis explicitly exploratory and use an adjustment such as false-discovery-rate control, or remove the inferential wording.
- [Sections 4.4 and 6] The survey-based evaluation lacks the information needed to support the positive-attitude conclusions. The percentages 'over 50%', '66.7%', '60%', and so on are not accompanied by a response count or response rate, and the survey was limited to interviewed users whose contact details were available, creating a self-selection risk that the paper acknowledges only indirectly. The paper itself states in Section 4.4 that the survey cannot be linked to user data because of anonymity, so it cannot validate any risk-related or clinical benefit. In addition, the Conclusions' statement that the approach 'can help to leverage the impact of Cognitive-Behavioral Therapies through chatbots' is unsupported because no CBT outcome was measured in this study; it should be explicitly marked as speculation or removed.
minor comments (5)
- [Section 6] The phrase 'To the best of your knowledge' should read 'To the best of our knowledge.'
- [Table 1] The row for '/noTengoAlias,/noAliases' contains a stray '95' before 'Starts.'
- [Figure 6] The Pearson correlation matrix is likely illegible at print resolution; please provide a high-resolution version with feature names and a clear legend.
- [Section 4.3] '1-gramas' should be '1-grams,' and the word-cloud analysis would benefit from stating the stop-word list used.
- [Section 5] 'With this experimental studio' should be 'With this experimental study.'
Circularity Check
No significant circularity: the central self-disclosure claim is an empirical coded observation, not a derived or fitted prediction, and the only self-citation is a non-load-bearing NLP toolkit.
full rationale
No circularity found. The paper makes no formal derivation claim that reduces an equation or fitted parameter to its own input. Its main assertion is the observational result in Section 4.3: “the self-disclosure technique is consistently effective, as 100% of users expressing concerns have previously engaged in empathetic interactions with the chatbot.” This is a coded temporal pattern from conversation logs, not a quantity fitted from data and then renamed as a prediction. The coding categories are distinct enough that the paper reports 9 chats where users helped the bot but did not open up, so the “100%” claim is not true by definition. The only self-citation is the TextFlow feature-extraction toolkit (ref [37]) used in Section 4.2; the engagement claim does not depend on it, so it is not load-bearing. No uniqueness theorem or ansatz is imported from the authors’ prior work, and the paper does not rename a known result as a new derivation. Methodological limitations—no control chatbot, no inter-rater reliability for chat coding, and no pre/post clinical re-evaluation—are evidentiary threats to the causal reading, but they are not circularity.
Assumptions & free parameters
free parameters (3)
- GPT-3 sampling temperature =
0.9
- Maximum response tokens =
170
- Controlled prompt refresh interval =
every 5 user messages
assumptions (4)
- domain assumption Self-disclosure by the bot leads to reciprocal self-disclosure by the user, following Social Penetration Theory.
- domain assumption Manual qualitative coding of conversations accurately distinguishes emotional engagement.
- domain assumption DeepL translation preserves meaning and affect sufficiently for GPT-3 to conduct empathic dialogue in Spanish.
- ad hoc to paper Pearson correlations computed on 44 users across 94 features are interpretable without adjustment for multiple testing.
Cite this review
Pith. "Pith review of An empathic GPT-based chatbot to talk about mental disorders with Spanish teenagers." pith.science (2026). https://pith.science/paper/IOO6SDO6
@misc{pith2026250505828,
author = {Pith},
title = {Pith review of: An empathic GPT-based chatbot to talk about mental disorders with Spanish teenagers},
year = {2026},
howpublished = {\url{https://pith.science/paper/IOO6SDO6}},
note = {Machine review of arXiv:2505.05828}
}
read the original abstract
This paper presents a chatbot-based system to engage young Spanish people in the awareness of certain mental disorders through a self-disclosure technique. The study was carried out in a population of teenagers aged between 12 and 18 years. The dialogue engine mixes closed and open conversations, so certain controlled messages are sent to focus the chat on a specific disorder, which will change over time. Once a set of trial questions is answered, the system can initiate the conversation on the disorder under the focus according to the user's sensibility to that disorder, in an attempt to establish a more empathetic communication. Then, an open conversation based on the GPT-3 language model is initiated, allowing the user to express themselves with more freedom. The results show that these systems are of interest to young people and could help them become aware of certain mental disorders.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
W. H. Organization, et al., World mental health report: transforming mental health for all, World Health Organization (2022)
work page 2022
-
[2]
J. ´Sniadach, S. Szymkowiak, P. Osip, N. Waszkiewicz, Increased De- pression and Anxiety Disorders during the COVID-19 Pandemic in Children and Adolescents: A Literature Review, Life 11 (11) (2021). doi:10.3390/life11111188. URL https://www.mdpi.com/2075-1729/11/11/1188
-
[3]
M. T. Hawes, A. K. Szenczy, D. N. Klein, G. Hajcak, B. D. Nelson, Increases in depression and anxiety symptoms in adolescents and young 35 adults during the COVID-19 pandemic, Psychological Medicine (2021) 1–9
work page 2021
-
[4]
CDC (Centers for Disease Control and Prevention), The youth risk behavior survey data summary and trends report: 2011–2021, https://www.cdc.gov/healthyyouth/data/yrbs/pdf/ YRBS Data-Summary-Trends Report2023 508.pdf (2023)
work page 2023
-
[5]
H. Chen, X. Liu, D. Yin, J. Tang, A survey on dialogue systems: Re- cent advances and new frontiers, SIGKDD Explor. Newsl. 19 (2) (2017) 25–35. doi:10.1145/3166054.3166058. URL https://doi.org/10.1145/3166054.3166058
arXiv 2017
-
[6]
J. Weizenbaum, Eliza—a computer program for the study of natural language communication between man and machine, Communications of the ACM 9 (1) (1966) 36–45
work page 1966
-
[7]
K. M. Colby, Modeling a paranoid mind, Behavioral and Brain Sciences 4 (4) (1981) 515–534
work page 1981
-
[8]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, in: Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Curran Associates Inc., 2017, p. 6000–6010
work page 2017
Show all 45 references
-
[9]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: Pre-training of deep bidirectional transformers for language understanding, in: Pro- ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologie...
2019 doi
-
[10]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Ch...
-
[11]
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, H. Peng, J. Li, J. Wu, Z. Liu, P. Xie, C. Xiong, J. Pei, P. S. Yu, L. Sun, A comprehensive survey on pretrained foundation models: A history from bert to chatgpt, arXiv preprint arXiv:2302.09419 (2023)
2023 arXiv
-
[12]
Thoppilan, D
R. Thoppilan, D. D. Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, Y. Li, H. Lee, H. S. Zheng, A. Ghafouri, M. Menegali, Y. Huang, M. Krikun, D. Lepikhin, J. Qin, D. Chen, Y. Xu, Z. Chen, A. Roberts, M. Bosma, V. Zhao, Y. Zhou, C.-...
2022 arXiv
-
[13]
Touvron, T
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozi` ere, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, G. Lample, Llama: Open and efficient foundation language models (2023). arXiv:2302.13971
2023 arXiv
-
[14]
URL https://cdn.openai.com/papers/gpt-4.pdf
Open AI, GPT-4 Technical Report (2023). URL https://cdn.openai.com/papers/gpt-4.pdf
2023
-
[15]
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, I. Mordatch, Decision transformer: Reinforcement learning via sequence modeling, in: M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, J. W. Vaughan (Eds.), Advances in Neural Information Pro-...
2021
-
[16]
Schick, J
T. Schick, J. Dwivedi-Yu, R. Dess` ı, R. Raileanu, M. Lomeli, L. Zettle- moyer, N. Cancedda, T. Scialom, Toolformer: Language models can teach themselves to use tools (2023). arXiv:2302.04761. 37
2023 arXiv
-
[17]
Huang, L
S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, Q. Liu, K. Aggarwal, Z. Chi, J. Bjorck, V. Chaudhary, S. Som, X. Song, F. Wei, Language is not all you need: Aligning perception with language models, arXiv preprint arXiv:2302.14045 (2023)
2023 arXiv
-
[18]
Siddique, J
S. Siddique, J. C. Chow, Machine learning in healthcare communication, Encyclopedia 1 (1) (2021) 220–239
2021
-
[19]
J. C. Chow, V. Wong, L. Sanders, K. Li, Developing an ai-assisted educational chatbot for radiotherapy using the ibm watson assistant platform, in: Healthcare, Vol. 11, MDPI, 2023, p. 2417
2023
-
[20]
A. N. Vaidyam, H. Wisniewski, J. D. Halamka, M. S. Kashavan, J. B. Torous, Chatbots and conversational agents in mental health: a review of the psychiatric landscape, The Canadian Journal of Psychiatry 64 (7) (2019) 456–464
2019
-
[21]
Balaskas, S
A. Balaskas, S. M. Schueller, A. L. Cox, G. Doherty, Ecological momen- tary interventions for mental health: A scoping review, PloS one 16 (3) (2021) e0248152
2021
-
[22]
Ahmed, A
A. Ahmed, A. Hassan, S. Aziz, A. A. Abd-Alrazaq, N. Ali, M. Alzubaidi, D. Al-Thani, B. Elhusein, M. A. Siddig, M. Ahmed, et al., Chatbot fea- tures for anxiety and depression: A scoping review, Health Informatics Journal 29 (1) (2023)
2023
-
[23]
Skjuve, A
M. Skjuve, A. Følstad, K. I. Fostervold, P. B. Brandtzaeg, My chat- bot companion-a study of human-chatbot relationships, International Journal of Human-Computer Studies 149 (2021) 102601
2021
-
[24]
M. Luo, J. T. Hancock, Self-disclosure and social media: motivations, mechanisms and psychological well-being, Current opinion in psychology 31 (2020) 110–115
2020
-
[25]
Y.-C. Lee, N. Yamashita, Y. Huang, Designing a chatbot as a mediator for promoting deep self-disclosure to a real mental health professional, Proceedings of the ACM on Human-Computer Interaction 4 (CSCW1) (2020) 1–27. 38
2020
-
[26]
Skjuve, A
M. Skjuve, A. Følstad, K. I. Fostervold, P. B. Brandtzaeg, A longitudinal study of human–chatbot relationships, International Journal of Human- Computer Studies 168 (2022) 102903
2022
-
[27]
V. J. Derlega, S. Metts, S. Petronio, S. T. Margulis, Self-disclosure., Sage Publications, Inc, 1993
1993
-
[28]
Altman, D
I. Altman, D. A. Taylor, Social penetration: The development of inter- personal relationships., Holt, Rinehart & Winston, 1973
1973
-
[29]
Carpenter, K
A. Carpenter, K. Greene, Social penetration theory, John Wiley and Sons, Ltd, 2015, pp. 1–4. doi:10.1002/9781118540190.wbeic160
2015 doi
-
[30]
C. T. Hill, D. E. Stull, Gender and self-disclosure: Strategies for ex- ploring the issues, in: V. J. Derlega, J. H. Berg (Eds.), Self-disclosure: Theory, research, and therapy, Springer US, 1987, pp. 81–100. doi: 10.1007/978-1-4899-3523-6 \ 5
1987 doi
-
[31]
Q. Yu, T. Nguyen, S. Prakkamakul, N. Salehi, ”i almost fell in love with a machine” speaking with computers affects self-disclosure, in: Extended Abstracts of the 2019 CHI conference on human factors in computing systems, Association for Computing Machinery, 2019, pp. 1–6
2019
-
[32]
S. D. Hollon, A. T. Beck, Cognitive and cognitive-behavioral therapies, Bergin and Garfield’s handbook of psychotherapy and behavior change 6 (2013) 393–442
2013
-
[33]
K. K. Fitzpatrick, A. Darcy, M. Vierhile, Delivering cognitive behav- ior therapy to young adults with symptoms of depression and anx- iety using a fully automated conversational agent (woebot): A ran- domized controlled trial, JMIR Ment Health 4 (2) (2017) e19. doi: 10.2196/m...
2017 doi
-
[34]
Inkster, S
B. Inkster, S. Sarda, V. Subramanian, et al., An empathy-driven, conver- sational artificial intelligence agent (wysa) for digital mental well-being: real-world data evaluation mixed-methods study, JMIR mHealth and uHealth 6 (11) (2018)
2018
-
[35]
A. A. Abd-Alrazaq, M. Alajlani, A. A. Alalwan, B. M. Bewick, P. Gard- ner, M. Househ, An overview of the features of chatbots in mental 39 health: A scoping review, International Journal of Medical Informatics 132 (2019) 103978. doi:https://doi.org/10.1016/j.ijmedinf.2019. 103978
2019 doi
-
[36]
Coghlan, K
S. Coghlan, K. Leins, S. Sheldrick, M. Cheong, P. Gooding, S. D’Alfonso, To chat or bot to chat: Ethical issues with using chatbots in mental health, DIGITAL HEALTH 9 (2023) 20552076231183542.doi:10.1177/ 20552076231183542
2023
-
[37]
M. E. Vallecillo-Rodr´ ıguez, J. Collado-Monta˜ nez, A. Montejo-R´ aez, sinai-uja/textflow (2023). URL https://github.com/sinai-uja/textflow
2023
-
[38]
Shuster, S
K. Shuster, S. Poff, M. Chen, D. Kiela, J. Weston, Retrieval augmentation reduces hallucination in conversation, arXiv preprint arXiv:2104.07567 (2021)
2021 arXiv
-
[39]
J. A. Hall, J. A. Harrigan, R. Rosenthal, Nonverbal behavior in clin- ician—patient interaction, Applied and preventive psychology 4 (1) (1995) 21–37
1995
-
[40]
Mehrabian, J
A. Mehrabian, J. A. Russell, An approach to environmental psychology., the MIT Press, 1974
1974
-
[41]
Y. Jain, H. Gandhi, A. Burte, A. Vora, Mental and physical health management system using ml, computer vision and iot sensor network, in: 2020 4th International Conference on Electronics, Communication and Aerospace Technology (ICECA), IEEE, 2020, pp. 786–791
2020
-
[42]
Weiste, A
E. Weiste, A. Per¨ akyl¨ a, Prosody and empathic communication in psy- chotherapy interaction, Psychotherapy Research 24 (6) (2014) 687–701
2014
-
[43]
Garrido-Mu˜ noz, A
I. Garrido-Mu˜ noz, A. Montejo-R´ aez, F. Mart´ ınez-Santiago, L. A. Ure˜ na- L´ opez, A survey on bias in deep nlp, Applied Sciences 11 (7) (2021) 3184. doi:https://doi.org/10.3390/app11073184
2021 doi
-
[44]
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, P. Fung, Survey of hallucination in natural language gen- eration, ACM Computing Surveys 55 (12) (2023) 1–38. 40
2023
-
[45]
J. C. Chow, L. Sanders, K. Li, Impact of chatgpt on medical chatbots as a disruptive technology, Frontiers in artificial Intelligence 6 (2023) 1166014. doi:10.3389/frai.2023.1166014. 41
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.