{"id":"d612dda9-cfe8-4969-afa4-00e0ff16c69a","arxiv_id":"2505.05828","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A GPT-3-based chatbot with a vulnerable-teenager persona engaged 44 Spanish teenagers in mental-health conversations, with most users opening up emotionally, though the pilot lacked a control group and clinical measures.","lead":"Researchers built a Spanish-language Telegram chatbot, powered by GPT-3, that acts as a vulnerable teenage peer to get Spanish teenagers talking about mental health. In a small pilot, most of the 44 teenagers who chatted opened up emotionally and reported a positive experience, but the study had no comparison group and did not measure clinical outcomes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 100% self-disclosure claim in Section 4.3 rests on single-coder temporal ordering in 44 chats, with no control condition or inter-rater reliability; a blinded re-coding and a neutral-bot comparison would settle whether the causal mechanism is real.","rationale":"I read the paper as a small feasibility study of a GPT-3-based Telegram chatbot for Spanish teenagers, and the strongest defensible finding is that the system was usable, generally liked, and capable of eliciting substantial conversation. The architecture is described in enough detail to reimplement, the authors transparently note the small sample and the lack of a psychologist reevaluation (Section 4.4), and the ethics statement is appropriate. My concern targets the same weakest assumption the reader flagged: the causal role of self-disclosure. I sharpened it to the specific 100% statement in Section 4.3, which is an inference from a single-coder temporal pattern in 44 conversations with no control and no reliability check. The survey and usage data support interest but not the causal mechanism. This does not justify rejecting the paper, because the stated aim is feasibility, not a controlled clinical trial. However, the abstract and conclusions should be read as reporting user engagement and perceived value, not as evidence that the self-disclosure technique causes disclosure. The reader's CONDITIONAL verdict remains appropriate, with the condition being that claims about the self-disclosure mechanism be explicitly downgraded or supported by a control comparison.","tokens_in":19957,"tokens_out":3152,"duration_ms":37045,"concrete_test":"Re-code the 44 conversations with two independent annotators who are blind to the paper's hypotheses and who apply a pre-registered rubric for two behaviors: (a) the user helps, advises, or shows empathy toward the bot, and (b) the user discloses a personal concern. Compute Cohen's kappa for each category, and then test the temporal claim per user: for each disclosing user, record the turn index of the first disclosure and the first helping behavior. If kappa is below approximately 0.6, or if any disclosure occurs before a helping behavior, the Section 4.3 claim is falsified. If the ordering survives, run a randomized comparison with a warm but non-self-disclosing control bot on a new volunteer sample; if disclosure rates are statistically indistinguishable, then self-disclosure is not the active ingredient.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's distinctive contribution is the self-disclosure design, so the central claim is not merely that a chatbot can engage teenagers, but that the bot's self-disclosure causes teenagers to open up. Section 4.3 states: 'the self-disclosure technique is consistently effective, as 100% of users expressing concerns have previously engaged in empathetic interactions with the chatbot.' The evidence for this is a temporal pattern observed by the authors while manually reading 44 conversations: users who disclosed concerns had earlier helped or advised the bot. This inference is load-bearing and insecure for three reasons. First, there is no control condition: a warm, attentive chatbot that does not self-disclose could plausibly elicit the same disclosures, so the observed behavior cannot be attributed to self-disclosure specifically. Second, the coding was done by the authors with no reported inter-rater reliability, and the categories 'care about the bot,' 'get involved,' and 'open up' are not defined by a published rubric. Third, the 100% precedence claim is compatible with a simple engagement confound: longer conversations give users more opportunities both to help the bot and to disclose, so the ordering may reflect conversation length rather than the mechanism. The usage statistics and survey support interest and perceived friendliness, but they do not validate the causal reading. The conclusion's additional leap to CBT impact (Section 6) is also unsupported by any outcome measure. Thus the abstract's claim that such systems 'could help' teenagers become aware of disorders is plausible as a feasibility finding, but the stronger causal claim about self-disclosure is not established by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a field deployment of a GPT-3-based chatbot ('Ada', 'Hugo', 'Big') on Telegram for Spanish-speaking teenagers aged 12–18, mixing psychologist-authored controlled dialogue, triage questions, and open dialogue with Spanish↔English translation. Usage statistics from 44 active users, a post-hoc anonymous survey, NLP feature correlations, and a manual reading of all 44 chats are presented. The authors conclude that the self-disclosure persona led most users to engage emotionally, that 100% of users who disclosed concerns first helped or advised the bot, and hence that such systems interest young people and can raise awareness of mental disorders.","tokens_in":20161,"tokens_out":5747,"duration_ms":57717,"significance":"If the descriptive claims hold, the study contributes a concrete, privacy-preserving deployment with psychologist involvement, a Spanish-language corpus from adolescent interactions, and a design pattern (a vulnerable peer bot with selectable gender) that could inform future mental-health chatbots for youth. The paper's strengths include a real deployment over roughly three and a half months, ethics approval with parental consent, explicit acknowledgment of some limitations in Section 4.4, and enough architectural detail to permit approximate reproduction. There are no fitted models or formal derivations whose circularity could threaten the central result; the main risk is causal overreading of observational data rather than internal inconsistency in the system design.","major_comments":[{"comment":"The paper's distinctive contribution is the self-disclosure design, and the load-bearing sentence — 'the self-disclosure technique is consistently effective, as 100% of users expressing concerns have previously engaged in empathetic interactions with the chatbot' — is not supported by the evidence as reported. The observation is a temporal ordering in 44 chats that were read by the authors, with no published coding rubric, no inter-rater reliability check, and no comparison condition such as a warm bot that does not self-disclose. Longer conversations mechanically allow more opportunities both to advise the bot and to disclose concerns, so the ordering is compatible with a conversation-length confound. Please reframe this as an observational association, report a reliability analysis or blinded re-coding, and soften the corresponding claim in the abstract.","section":"Section 4.3"},{"comment":"The numerical summary is internally inconsistent. The text reports 31/44 users 'care about the bot and get involved in advising and helping it' and only 22/44 'open up and talk about their concerns'; the statement that 'more than 70% of the users engaged emotionally with the bot, sharing their concerns and worries' conflates the 70% helping figure with the 50% disclosure figure, and Section 5 repeats this as '70% emotional openness.' Because the self-disclosure claim concerns the 22 users who disclosed, please report the two rates separately and correct the Discussion text.","section":"Sections 4.3 and 5"},{"comment":"The NLP analysis computes Pearson correlations across 94 linguistic features using n=44 users and then interprets the top correlations (e.g., coordinating-conjunction frequency above 0.4) as evidence that 'there are differences in language between people classified as healthy and people classified as having some form of mental disorder.' With 94 features and no multiple-testing correction, top correlations of this size are expected under noise, and no confidence intervals or effect sizes are provided. Please label this analysis explicitly exploratory and use an adjustment such as false-discovery-rate control, or remove the inferential wording.","section":"Section 4.2"},{"comment":"The survey-based evaluation lacks the information needed to support the positive-attitude conclusions. The percentages 'over 50%', '66.7%', '60%', and so on are not accompanied by a response count or response rate, and the survey was limited to interviewed users whose contact details were available, creating a self-selection risk that the paper acknowledges only indirectly. The paper itself states in Section 4.4 that the survey cannot be linked to user data because of anonymity, so it cannot validate any risk-related or clinical benefit. In addition, the Conclusions' statement that the approach 'can help to leverage the impact of Cognitive-Behavioral Therapies through chatbots' is unsupported because no CBT outcome was measured in this study; it should be explicitly marked as speculation or removed.","section":"Sections 4.4 and 6"}],"minor_comments":[{"comment":"The phrase 'To the best of your knowledge' should read 'To the best of our knowledge.'","section":"Section 6"},{"comment":"The row for '/noTengoAlias,/noAliases' contains a stray '95' before 'Starts.'","section":"Table 1"},{"comment":"The Pearson correlation matrix is likely illegible at print resolution; please provide a high-resolution version with feature names and a clear legend.","section":"Figure 6"},{"comment":"'1-gramas' should be '1-grams,' and the word-cloud analysis would benefit from stating the stop-word list used.","section":"Section 4.3"},{"comment":"'With this experimental studio' should be 'With this experimental study.'","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper's real contribution is a feasibility result. A GPT-3-powered Telegram chatbot that role-plays a vulnerable teenage peer can get Spanish adolescents to chat about mental health topics, and most survey respondents liked it. That is worth having. The stronger statement in Section 4.3 — that self-disclosure is 'consistently effective' because 100% of users who opened up had first helped the bot — is not supported by the evidence as presented.\n\nWhat is new: first deployment I know of using a self-disclosure peer persona with GPT-3 for Spanish-speaking teenagers. The architecture is described in enough detail to reimplement: controlled psychologist-authored questions mixed with open dialogue, topic selection softened, a triage of sensitivity, DeepL translation to keep GPT-3 in English, and a human alert path for risk. The recruitment and ethics process is proper, with parental consent. The manual reading of conversations gives useful qualitative insight: teens wanted to talk about daily problems, relationships, and loneliness more than the controlled disorder questions, and they tended to thank the bot. Those observations have design value.\n\nSoft spots, in order of importance. First, the causal reading of self-disclosure. There is no control bot, no randomization, no inter-rater reliability, and the coding categories are not published. The '100%' is just a temporal precedence in 44 chats; it is compatible with the idea that longer, warmer conversations create more opportunities for both helping and disclosing. The paper would be stronger if it presented this as an observation that self-disclosure preceded openness in many cases, not as evidence the technique works. Second, the NLP correlations in Section 4.2 are exploratory and not corrected for multiple comparisons; the authors present them as suggestive, so this is a minor issue. Third, the conclusion's nod to CBT impact is speculative and should be cut or labeled as future work. The limitations section does acknowledge the pilot nature, so these are overstatements around a reasonable pilot, not fabrication.\n\nVerdict: yes, this deserves a serious referee. A good reviewer would ask for a neutral-bot comparison or at least a re-coding by a second annotator and softer causal language. The paper is for HCI and mental-health chatbot researchers, especially those working in non-English populations. I would bring it to a reading group and would cite it as a design reference for Spanish-language chatbot work.","headline":"A useful Spanish-language feasibility pilot for a teen mental-health chatbot; its self-disclosure causal claim is the one real overreach.","tokens_in":20788,"tokens_out":2329,"would_cite":true,"duration_ms":24387,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T50","97C99"],"pacs":["07.05.Wr"],"model":"deepseek-v4-flash","headline":"This paper reports a Telegram chatbot with a vulnerable-teen persona that uses self-disclosure to engage Spanish adolescents in discussions about mental disorders, and finds that most users who conversed opened up emotionally.","keywords":["Dialogue systems","Large Language Model","Mental Disorders","Natural Language Generation","GPT-3","Self-disclosure","chatbot","adolescents"],"falsifier":"Assign teenagers at random to the self-disclosing bot versus a warm but emotionally neutral bot that asks identical questions, and count how many share personal concerns; if the neutral bot produces a similar disclosure rate, the claim that self-disclosure drives openness is refuted.","tokens_in":19704,"feed_emoji":"🤖","tokens_out":6243,"duration_ms":64033,"temperature":0.7,"pith_summary":"The paper reports on a Telegram chatbot, built for Spanish teenagers aged 12 to 18, that combines psychologist-written questions with free conversation powered by GPT-3 and automatic translation. Its central claim is that the self-disclosure technique, in which the bot presents itself as a worried teenage peer and shares its own fears before asking about the user's, creates enough trust for adolescents to talk about sensitive topics such as depression, anxiety, eating disorders, and cyberbullying. In a pilot with 44 teenagers who actually conversed, over 70 percent engaged emotionally, and every user who expressed a personal concern had first interacted empathetically with the bot. The authors conclude that such systems interest young people and could help them become aware of mental disorders, while stressing that the bot is not a substitute for professional care.","feed_headline":"Self-disclosing chatbot gets Spanish teens to open up","feed_subtitle":"In a 44-user pilot, teens who revealed concerns had first engaged with the bot's own worries, supporting a trust-first design.","key_machinery":"The central mechanism is a two-layer dialogue engine: a controlled dialogue built from psychologist-designed triage questions and topic prompts, and an open dialogue using GPT-3, specifically the Davinci-002 model, with DeepL translating between Spanish and English. The bot's persona is itself the key instrument: it appears as a teenager named Ada, Hugo, or Big, lets the user choose the bot's gender, reveals personal worries, asks the user for advice, and reciprocates, following the disclosure layers of Social Penetration Theory. Every five user turns a specialist prompt steers the conversation back on track, and language suggesting self-harm or suicidal ideation triggers an alert to a human. This hybrid of controlled and open dialogue is what the paper credits for both conversational fluency and safety.","core_discovery":"The study's discovery is that a chatbot designed as a vulnerable, disclosing peer, rather than as a therapist or a neutral assistant, can draw Spanish teenagers into sustained and emotionally open conversations about mental health. Of the 44 users who moved past onboarding, 31 helped and advised the bot and 22 shared their own worries, and in all 22 cases empathetic engagement with the bot preceded the disclosure. The authors interpret this as evidence that self-disclosure is consistently effective, and they combine usage statistics, linguistic feature analysis, manual conversation review, and a user survey to support the view that the system was well received and potentially useful for raising awareness.","pith_inferences":["Editorial inference: the paper's 100 percent disclosure result shows temporal precedence, not causation; a randomized trial with a neutral comparison bot would be needed to prove that self-disclosure itself drives openness.","Editorial inference: because the manual coding of engagement and openness was done by the authors without a second, independent rater, replicating the 70 percent emotional-engagement figure with inter-rater reliability checks would harden the main outcome.","Editorial inference: the Spanish-English translation loop, despite its errors, suggests a reusable recipe for deploying English-centric large language models in lower-resource languages for sensitive dialogue, provided colloquialisms and gender agreement are handled.","Editorial inference: the collected corpus of 1,860 messages and 94 linguistic features could feed early-detection models, but the deliberate anonymity that separates survey responses from chat data means user-satisfaction findings cannot yet be linked to clinical or triage status."],"forward_implications":["If the finding holds, a freely available chatbot on a familiar messaging platform can serve as a low-stigma first step for teenagers to talk about depression, anxiety, eating disorders, and related topics in Spanish.","A bot that reveals its own worries before asking about the user's can create an atmosphere in which many teenagers reciprocate with personal concerns, supporting self-disclosure as an engagement strategy rather than a purely scripted interview.","Combining a controlled, psychologist-designed dialogue with an open GPT-3 conversation keeps the chat on topic while letting users drift toward what actually worries them, such as friendship, break-ups, and school, suggesting that rigid disorder-focused scripts miss the real content.","The Spanish-to-English translation loop is workable, but colloquial speech and grammatical gender are recurring failure points, so native-Spanish models or better handling of gendered language would improve fluency and comfort.","The system's risk-alert mechanism, together with the safe framing of topics under soft names, offers a template for how generative chatbots can address highly sensitive mental-health content with teenagers without impersonating a clinician."],"supporting_citations":[{"why":"Supplies evidence that self-disclosure in social media is a mechanism tied to psychological well-being, motivating the bot's confessional persona.","marker":"[24]"},{"why":"Shows a chatbot can be designed as a mediator to promote deep self-disclosure to a mental-health professional, a direct precedent for this design.","marker":"[25]"},{"why":"Provides Social Penetration Theory, the theoretical basis that progressive self-disclosure deepens trust and intimacy.","marker":"[29]"},{"why":"Documents that users perceive chatbots as companions, supporting the decision to give the bot a relatable teenage identity.","marker":"[23]"},{"why":"Validates chatbot-delivered cognitive-behavioral therapy with a randomized controlled trial for young adults, the key reference point that chatbots can affect mental-health outcomes.","marker":"[33]"},{"why":"Supplies GPT-3, the language model that generates the open-dialogue responses at the core of the system.","marker":"[10]"}],"fun_headline_variants":["Bot's own worries coax Spanish teens into sharing","Self-disclosing chatbot sparks teen mental health talks","For Spanish teens, a bot that confesses first gets truth","Teens open up after chatbot bares its own struggles","Vulnerable chatbot wins trust of Spanish teenagers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the assumption that the bot's self-revelation is what causes teenagers to open up, yet the evidence only shows that in every chat the user's disclosure came after empathetic engagement with the bot, with no neutral comparison chatbot to rule out other causes.","fun_headline_variants_meta":{"raw":{"variants":["Bot's own worries coax Spanish teens into sharing","Self-disclosing chatbot sparks teen mental health talks","For Spanish teens, a bot that confesses first gets truth","Teens open up after chatbot bares its own struggles","Vulnerable chatbot wins trust of Spanish teenagers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1232,"prompt_tokens":811,"completion_tokens":421,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":344}},"tokens_in":427,"tokens_out":421,"duration_ms":4657,"temperature":1.0,"reasoning_tokens":344,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:54:56.257513+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Assign teenagers at random to the self-disclosing bot versus a warm but emotionally neutral bot that asks identical questions, and count how many share personal concerns; if the neutral bot produces a similar disclosure rate, the claim that self-disclosure drives openness is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies evidence that self-disclosure in social media is a mechanism tied to psychological well-being, motivating the bot's confessional persona."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a chatbot can be designed as a mediator to promote deep self-disclosure to a mental-health professional, a direct precedent for this design."},{"cited_title":"Carpenter, K","cited_arxiv_id":null,"evidence_quote":"Provides Social Penetration Theory, the theoretical basis that progressive self-disclosure deepens trust and intimacy."},{"cited_title":"Skjuve, A","cited_arxiv_id":null,"evidence_quote":"Documents that users perceive chatbots as companions, supporting the decision to give the bot a relatable teenage identity."}],"review_version":1}