Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Evaluating an LLM-Powered Chatbot for Cognitive Restructuring: Insights from Mental Health Professionals

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A prompt-engineered chatbot can run cognitive restructuring sessions, but expert review shows relational missteps that could harm trust.

desk verdict A useful qualitative evaluation of a real-user LLM CR chatbot, with safety claims that need trimming to match the evidence. read the letter →

arxiv 2501.15599 v1 pith:6N6GWNWU submitted 2025-01-26 cs.HC

classification cs.HC
keywords cognitiverestructuringLLMchatbotmentalhealthprofessionalstherapeuticalliancepowerdynamicstoxicpositivityexpertevaluationGPT-4
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper evaluates whether a large language model can deliver cognitive restructuring (CR), a core technique of cognitive behavioral therapy, through real user conversations. Nineteen college students used a GPT-4-based chatbot, and four mental health professionals reviewed the resulting logs. Experts found the chatbot followed the three CR phases and asked Socratic questions, but they also identified repeated problems with excessive positivity, leading questions, evaluative praise, unsolicited advice, and misinterpretation of users' emotional cues. The authors argue these issues risk eroding therapeutic rapport and raise ethical concerns for autonomous AI therapy.

What carries the argument

The central object is CRBot, a web chatbot built on GPT-4 via Azure OpenAI Service and controlled by system prompts and ten few-shot examples co-designed with five mental health professionals. The few-shot scenarios cover successful completion, absence of negative thoughts, identification challenges, challenging barriers, and creation of alternative thoughts. The evaluation machinery is an expert review process in which four additional mental health professionals used think-aloud sessions to assess 95 conversation logs, with thematic analysis and iterative consensus-building to reconcile divergent assessments. This two-layer design—structured prompts on one side, expert qualitative judgment on the other—carries the argument that protocol adherence and relational quality can be separated.

What would settle it

If a replication study collected standardized user-reported measures of therapeutic alliance, perceived autonomy, and emotional safety after each session, and these measures showed no correlation with the expert-flagged problems, the claim that these conversational patterns risk eroding rapport would be directly contradicted.

Watch

Extended reading notes

Core claim

The paper claims that an LLM-based CR approach can adhere to the core CR protocol of exploration, evaluation, and substitution while keeping conversations natural and fluid. However, expert review of real user interactions reveals that the same system misuses positive regard, reinforces power imbalances through leading questions and evaluative language, gives advice without adequate context, and misreads subtle linguistic markers such as 'maybe' or short, disengaged replies. These behaviors, the authors conclude, can undermine user autonomy and trust, even when the chatbot structurally follows the protocol, so safe deployment requires expert oversight and careful alignment of language style.

Load-bearing premise

The study assumes that mental health professionals' qualitative judgments of conversation logs accurately reflect how users actually experienced the chatbot, without measuring user-reported outcomes or working alliance.

Editorial extensions

If this is right

  • If the protocol adherence holds in larger samples, prompt-engineered LLMs could offer scalable, low-cost scaffolding for structured self-help interventions.
  • If the relational risks are general, LLM therapy tools will need language-style guardrails, such as banning evaluative praise and 'right?' questions, before autonomous use.
  • If LLMs cannot reliably detect session-wide disengagement, standalone deployment is unsafe and human monitoring or context-aware models become necessary.
  • If expert judgments mirror user experience, redesigning the chatbot to avoid toxic positivity and unsolicited advice would reduce harm to the therapeutic alliance.
  • If the findings extend to other structured therapies, the design implications could guide safer LLM deployment beyond cognitive restructuring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension would measure user-reported alliance and autonomy alongside expert ratings to test whether the relational risks experts flagged actually matter to users.
  • The paper's observation about 'maybe' suggests future systems could add explicit uncertainty detection and a confirm-before-proceed step, an idea the authors raise but do not implement.
  • Because participants had only none-to-moderate symptoms, the findings may not transfer to clinical populations, so supervised testing with symptomatic users would sharpen safety conclusions.
  • The power-dynamics findings imply that AI therapists should treat language as a relational act, a principle that could generalize to other high-stakes AI assistants.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents CRBot, an LLM-powered chatbot for cognitive restructuring (CR), co-designed with five mental health professionals through prompt engineering and few-shot examples. Nineteen college students with none-to-moderate symptom scores used CRBot, and four additional mental health professionals reviewed 95 conversation logs using think-aloud sessions and thematic analysis. The authors report that CRBot can adhere to core CR steps, maintain a natural conversational flow, and pose Socratic questions, while also identifying limitations such as toxic positivity, evaluative language, advice-giving, and misunderstanding of user context. The paper concludes with design implications and calls for human oversight. The central contribution is an exploratory qualitative evaluation of real user-chatbot interactions with expert review, rather than an outcome study.

Significance. If the findings are appropriately qualified, the paper makes a timely and useful contribution to the underexplored area of LLM-enabled psychotherapy. Its strengths include the use of real user interaction logs rather than synthetic or imagined scenarios, a transparent account of the co-design process, and detailed expert quotes that illustrate the identified themes. The design implications in Section 5.2 are concrete and actionable. However, the paper's headline claims about eroded rapport and ethical concerns go beyond the evidence, which consists of expert interpretations of transcripts with no user-reported outcomes. As an exploratory qualitative study, the work is valuable; the safety-related conclusions need to be reframed as expert-identified risks rather than established effects.

major comments (3)
  1. [Abstract vs. Section 5.2] The abstract and Introduction state that the findings reveal 'ethical concerns' as a central result, but Section 5.2 explicitly says: 'Although mental health professionals (MHPs) did not identify ethical concerns in this study, they raised future concerns regarding some of the observed LLM behaviors.' This contradiction is load-bearing because the paper's safety implications are part of its main claim. The authors should either revise the abstract and conclusions to say that the study identified potential or future ethical risks, or provide specific evidence of ethical concerns actually identified by the MHPs in the reviewed logs.
  2. [Sections 3.2 and 4.4] The claims that CRBot's behaviors 'risk eroding rapport' and 'raise ethical concerns' rest entirely on mental health professionals' post-hoc interpretations of conversation logs. The manuscript reports no user-reported outcomes, no working alliance measure, no post-session questionnaire, and no follow-up assessment. Section 3.1.2 also restricts participants to those with none-to-moderate symptoms. Expert transcript review is a useful but unvalidated proxy for real user experience; statements such as 'this person doesn't want to engage' (MHP3, Section 4.4) are expert inferences, not observed user responses. The authors should explicitly label these as expert-identified risks and add a limitation stating that user-side validation is needed before drawing conclusions about actual rapport erosion or harm.
  3. [Sections 3.1.1, 3.2, and 4.1] There is a partial circularity in the positive CR-adherence finding. Section 3.1.1 states that the system prompts and few-shot examples were collaboratively designed with five MHPs specifically to ensure adherence to CR, and Section 3.2 then has four different MHPs evaluate the logs and confirm CR adherence (Section 4.1). This makes the adherence result partly a check on whether the prompt engineering succeeded, rather than an independent test of whether LLMs can inherently deliver CR. The paper should frame this finding as verification of prompt fidelity and, ideally, compare against a non-CR-specific baseline or a generic chatbot to support broader claims about LLM capability.
minor comments (5)
  1. [Section 4.2] The subsection title 'Misuse Positive Regards' contains a grammatical error; it should be 'Misuse of Positive Regard.'
  2. [Figure 3] All dialog snippets in Figure 3 begin with the same generic user line ('I feel anxious about tomorrow's conference talk') even though the captions describe different scenarios such as acne, a roommate, or a professor. Using actual anonymized excerpts from the corresponding logs would avoid confusion and better support the thematic claims.
  3. [Section 3.2.2] The thematic analysis section reports the General Inductive Approach and team discussions but does not describe any inter-rater reliability checks, independent coding, or an explicit audit trail. For a qualitative study this is not fatal, but a brief statement about coding consistency would strengthen confidence in the reported themes.
  4. [Table 1] Table 1 reports n=95 dialogs without explaining how this number relates to the 19 participants; adding one sentence in the text or table note clarifying that participants engaged in multiple sessions would improve interpretability.
  5. [References] Reference [40] contains a typo in the page numbers ('1𝜓7'), which should be corrected to the proper volume and page range.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the co-designed prompt and the independent MHP evaluation are distinct, and no fitted parameter is renamed as a prediction.

full rationale

The paper's claimed derivation chain is empirical rather than definitional. CRBot's prompts and few-shot examples were co-designed with five MHPs (Section 3.1.1) to implement cognitive restructuring, but the adherence findings were produced by four additional, non-overlapping MHPs who reviewed conversation logs (Section 3.2.1). The evaluators were not the designers, so the positive adherence finding is a genuine check on whether the prompt engineering worked; it could have failed, and in fact the evaluators reported unscripted problems such as toxic positivity, leading questions, advice-giving, and misinterpretation of user cues. No quantitative or qualitative parameter is fitted to the evaluators' judgments and then presented as an independent prediction, and no load-bearing claim depends on a self-citation or on an imported uniqueness theorem. The abstract's framing of 'ethical concerns' is weakened by Section 5.2's explicit statement that 'mental health professionals (MHPs) did not identify ethical concerns in this study,' and the rapport-erosion and safety implications rest on expert transcript review rather than user-outcome measures. These are validity and framing limitations, not circularity under the defined rubric, so no circular step is flagged.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No numeric free parameters are involved. The central claims rest on qualitative assumptions about the validity of expert judgment, the adequacy of the sample, and the representativeness of the chosen model. No new entities are introduced.

assumptions (4)
  • domain assumption Cognitive restructuring, as described by the three-step model (exploration, evaluation, substitution), is a valid and effective therapeutic technique.
    Underpins the entire evaluation; taken from the clinical literature without challenge (Section 2.1).
  • domain assumption Mental health professionals' qualitative review of conversation transcripts is a valid measure of therapeutic quality and risk.
    Section 3.2 uses MHP review as the primary evidence; no user outcome or alliance measures were collected.
  • domain assumption The sample of 19 college students with mild symptoms is sufficient to surface typical interaction patterns with the chatbot.
    Section 3.1.2 describes the small, homogeneous sample; generalizability to broader populations is uncertain.
  • domain assumption GPT-4 accessed via Azure OpenAI with the developed prompts is representative of LLM-based cognitive restructuring systems.
    Only one model and one prompt setup were tested (Section 3.1.1), yet the discussion generalizes to LLM-based psychotherapy.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating an LLM-Powered Chatbot for Cognitive Restructuring: Insights from Mental Health Professionals." pith.science (2026). https://pith.science/paper/6N6GWNWU

@misc{pith2026250115599,
  author       = {Pith},
  title        = {Pith review of: Evaluating an LLM-Powered Chatbot for Cognitive Restructuring: Insights from Mental Health Professionals},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6N6GWNWU}},
  note         = {Machine review of arXiv:2501.15599}
}
read the original abstract

Recent advancements in large language models (LLMs) promise to expand mental health interventions by emulating therapeutic techniques, potentially easing barriers to care. Yet there is a lack of real-world empirical evidence evaluating the strengths and limitations of LLM-enabled psychotherapy interventions. In this work, we evaluate an LLM-powered chatbot, designed via prompt engineering to deliver cognitive restructuring (CR), with 19 users. Mental health professionals then examined the resulting conversation logs to uncover potential benefits and pitfalls. Our findings indicate that an LLM-based CR approach has the capability to adhere to core CR protocols, prompt Socratic questioning, and provide empathetic validation. However, issues of power imbalances, advice-giving, misunderstood cues, and excessive positivity reveal deeper challenges, including the potential to erode therapeutic rapport and ethical concerns. We also discuss design implications for leveraging LLMs in psychotherapy and underscore the importance of expert oversight to mitigate these concerns, which are critical steps toward safer, more effective AI-assisted interventions.

Figures

Figures reproduced from arXiv: 2501.15599 by the authors.

Figure 1
Figure 1. Overall study flow Manuscript submitted to ACM [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. An example conversation and each column represents a step in CR. In the exploration step, [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. a) "That’s great" potentially overshadowed user’s negative experience b) Excessive positive regard in some sessions c) [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Embodied Empathy: A Multimodal AR and LLM-Powered System for Self-Attachment Psychotherapy with Self-Initiated Humour

    cs.HC 2026-08 conditional novelty 6.0 of 10

    A multimodal AR/LLM app for self-attachment therapy is feasible and engaging, but its emotion-mirroring and mood-improvement claims are only weakly evidenced.

Reference graph

Works this paper leans on

59 extracted references · 42 canonical work pages · cited by 1 Pith paper

  1. [1]

    Aseel Ajlouni, Abdallah Almahaireh, and Fatima Whaba. 2023. Students’ perception of using ChatGPT in counseling and mental health education: the benefits and challenges. International Journal of Emerging Technologies in Learning (iJET) 18, 20 (2023), 199–218

  2. [2]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862 (2022)

  3. [3]

    Pierre Baillargeon, Robert Coté, Lyne Douville, et al. 2012. Resolution process of therapeutic alliance ruptures: A review of the literature. Psychology 3, 12 (2012), 1049

  4. [4]

    Aaron T Beck. 1979. Cognitive therapy and the emotional disorders . Penguin

  5. [5]

    Judith S Beck. 2020. Cognitive behavior therapy: Basics and beyond . Guilford Publications

  6. [6]

    Olga V Berkout, Diana Tinsley, and Maureen K Flynn. 2019. A review of anger, hostility, and aggression from an ACT perspective. Journal of contextual behavioral science 11 (2019), 34–43

  7. [7]

    It is never okay to talk about suicide

    Matt Blanchard and Barry A Farber. 2020. “It is never okay to talk about suicide”: patients’ reasons for concealing suicidal ideation in psychotherapy. Psychotherapy Research 30, 1 (2020), 124–136

  8. [8]

    Pierre-William Breau. 2023. Low-resource suicide ideation and depression detection with multitask learning and large language models. (2023). https://hdl.handle.net/1866/32349

Show all 59 references
  1. [9]

    Laura E Captari, Steven J Sandage, Richard A Vandiver, Peter J Jankowski, and Joshua N Hook. 2023. Integrating positive psychology, reli- gion/spirituality, and a virtue focus within culturally responsive mental healthcare. Handbook of positive psychology, religion, and spirit...

  2. [10]

    Yujin Cho, Mingeon Kim, Seojin Kim, Oyun Kwon, Ryan Donghan Kwon, Yoonha Lee, and Dohyun Lim. 2023. Evaluating the efficacy of interactive language therapy based on LLM for high-functioning autistic adolescent psychological counseling. arXiv preprint arXiv:2311.09243 (2023)

  3. [11]

    David A Clark. 2013. Cognitive restructuring.The Wiley handbook of cognitive behavioral therapy (2013), 1–22. https://doi.org/10.1002/9781118528563. wbcbt02

  4. [12]

    Gavin I Clark and Sarah J Egan. 2015. The Socratic method in cognitive behavioural therapy: a narrative review. Cognitive Therapy and Research 39 (2015), 863–879

  5. [13]

    Torrey A Creed, Sarah A Frankel, Ramaris E German, Kelly L Green, Shari Jager-Hyman, Kristin P Taylor, Abby D Adler, Courtney B Wolk, Shannon W Stirman, Scott H Waltman, et al. 2016. Implementation of transdiagnostic cognitive therapy in community behavioral health: The Beck C...

  6. [14]

    Munmun De Choudhury, Sachin R Pendse, and Neha Kumar. 2023. Benefits and harms of large language models in digital mental health. arXiv preprint arXiv:2311.14693 (2023)

  7. [15]

    Jerry L Deffenbacher. 2011. Cognitive-behavioral conceptualization and treatment of anger. Cognitive and Behavioral Practice 18, 2 (2011), 212–221

  8. [16]

    Ismail Dergaa, Feten Fekih-Romdhane, Souheil Hallit, Alexandre Andrade Loch, Jordan M Glenn, Mohamed Saifeddin Fessi, Mohamed Ben Aissa, Nizar Souissi, Noomen Guelmami, Sarya Swed, et al. 2024. ChatGPT is not ready yet for use in providing mental health assessment and interven...

  9. [17]

    Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. 2024. Longrope: Extending llm context window beyond 2 million tokens. arXiv preprint arXiv:2402.13753 (2024)

  10. [18]

    Changming Duan, Sarah Knox, and Clara E Hill. 2018. Advice giving in psychotherapy. The Oxford handbook of advice (2018), 175–195

  11. [19]

    Daniil Filienko, Yinzhou Wang, Caroline El Jazmi, Serena Xie, Trevor Cohen, Martine De Cock, and Weichao Yuwen. 2024. Toward large language models as a therapeutic tool: Comparing prompting techniques to improve gpt-delivered problem-solving therapy. arXiv preprint arXiv:2409....

  12. [20]

    Gordon CN Hall, Janie J Hong, Nolan WS Zane, and Oanh L Meyer. 2011. Culturally competent treatments for Asian Americans: The relevance of mindfulness and acceptance-based psychotherapies. Clinical Psychology: Science and Practice 18, 3 (2011), 215

  13. [21]

    Steven C Hayes and Heather Pierson. 2005. Acceptance and commitment therapy . Springer

  14. [22]

    Steven C Hayes, Kirk D Strosahl, and Kelly G Wilson. 2011. Acceptance and commitment therapy: The process and practice of mindful change . Guilford press

  15. [23]

    Zainab Iftikhar, Sean Ransom, Amy Xiao, and Jeff Huang. 2024. Therapy as an NLP Task: Psychologists’ Comparison of LLMs and Human Peers in CBT. arXiv preprint arXiv:2409.02244 (2024)

  16. [24]

    Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, and Xia Hu. 2024. Llm maybe longlm: Self-extend llm context window without tuning. arXiv preprint arXiv:2401.01325 (2024)

  17. [25]

    Lana Kambeitz-Ilankovic, Uma Rzayeva, Laura Völkel, Julian Wenzel, Johanna Weiske, Frank Jessen, Ulrich Reininghaus, Peter J Uhlhaas, Mario Alvarez-Jimenez, and Joseph Kambeitz. 2022. A systematic review of digital and face-to-face cognitive behavioral therapy for depression. ...

  18. [26]

    Mina J Kian, Mingyu Zong, Katrin Fischer, Abhyuday Singh, Anna-Maria Velentza, Pau Sang, Shriya Upadhyay, Anika Gupta, Misha A Faruki, Wallace Browning, et al. 2024. Can an LLM-powered socially assistive robot effectively and safely deliver cognitive behavioral therapy? A stud...

  19. [27]

    Anton O Kris. 1984. The conflicts of ambivalence. The Psychoanalytic Study of the Child 39, 1 (1984), 213–234

  20. [28]

    Kurt Kroenke, Robert L Spitzer, and Janet BW Williams. 2001. The PHQ-9: validity of a brief depression severity measure. Journal of general internal medicine 16, 9 (2001), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x

  21. [29]

    Willem Kuyken, Christine A Padesky, and Robert Dudley. 2011. Collaborative case conceptualization: Working effectively with clients in cognitive- behavioral therapy. Guilford Press

  22. [30]

    Hannah R Lawrence, Renee A Schneider, Susan B Rubin, Maja J Matarić, Daniel J McDuff, and Megan Jones Bell. 2024. The opportunities and risks of large language models in mental health. JMIR Mental Health 11, 1 (2024), e59479

  23. [31]

    Joanna Zun Li, Alina Herderich, and Amit Goldenberg. 2024. Skill but not effort drive gpt overperformance over humans in cognitive reframing of negative scenarios. (2024)

  24. [32]

    Marsha Linehan. 1993. Cognitive-behavioral treatment of borderline personality disorder . Guilford press

  25. [33]

    Rakesh K Maurya, Steven Montesinos, Mikhail Bogomaz, and Amanda C DeDiego. 2025. Assessing the use of ChatGPT as a psychoeducational tool for mental health practice. Counselling and Psychotherapy Research 25, 1 (2025), e12759

  26. [34]

    Jingping Nie, Hanya Shao, Yuang Fan, Qijia Shao, Haoxuan You, Matthias Preindl, and Xiaofan Jiang. 2024. LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices. arXiv preprint arXiv:2403.10779 (2024)....

  27. [35]

    John C Norcross and Michael J Lambert. 2018. Psychotherapy relationships that work III. Psychotherapy 55, 4 (2018), 303

  28. [36]

    Nick Obradovich, Sahib S Khalsa, Waqas U Khan, Jina Suh, Roy H Perlis, Olusola Ajilore, and Martin P Paulus. 2024. Opportunities and risks of large language models in psychiatry. NPP—Digital Psychiatry and Neuroscience 2, 1 (2024), 8. https://doi.org/10.1038/s44277-024-00010-z

  29. [37]

    Rafael Zambelli Pinto, Manuela L Ferreira, Vinicius C Oliveira, Marcia R Franco, Roger Adams, Christopher G Maher, and Paulo H Ferreira. 2012. Patient-centred communication is associated with positive therapeutic alliance: a systematic review. Journal of physiotherapy 58, 2 (2...

  30. [38]

    Kenneth S Pope and Melba JT Vasquez. 2016. Ethics in psychotherapy and counseling: A practical guide . John Wiley & Sons

  31. [39]

    Paolo Raile. 2024. The usefulness of ChatGPT for psychotherapists and patients. Humanities and Social Sciences Communications 11, 1 (2024), 1–8

  32. [40]

    Richard M Ryan, Martin F Lynch, Maarten Vansteenkiste, and Edward L Deci. 2011. Motivation and autonomy in counseling, psychotherapy, and behavior change: A look at theory and practice 1𝜓7. The Counseling Psychologist 39, 2 (2011), 193–260

  33. [41]

    Ashish Sharma, Kevin Rushton, Inna Wanyin Lin, Theresa Nguyen, and Tim Althoff. 2024. Facilitating Self-Guided Mental Health Interven- tions Through Human-Language Model Interaction: A Case Study of Cognitive Restructuring. https://doi.org/10.48550/arXiv.2310.15461 arXiv:2310....

  34. [42]

    Lucas, Adam S

    Ashish Sharma, Kevin Rushton, Inna Wanyin Lin, David Wadden, Khendra G. Lucas, Adam S. Miner, Theresa Nguyen, and Tim Althoff. 2023. Cognitive Reframing of Negative Thoughts through Human-Language Model Interaction. https://doi.org/10.48550/arXiv.2305.02466 arXiv:2305.02466 [cs.CL]

  35. [43]

    Inhwa Song, Sachin R Pendse, Neha Kumar, and Munmun De Choudhury. 2024. The typing cure: Experiences with large language model chatbots for mental health support. arXiv preprint arXiv:2401.14362 (2024)

  36. [44]

    Robert L Spitzer, Kurt Kroenke, Janet BW Williams, and Bernd Löwe. 2006. A brief measure for assessing generalized anxiety disorder: the GAD-7. Archives of internal medicine 166, 10 (2006), 1092–1097. https://doi.org/10.1001/archinte.166.10.1092

  37. [45]

    Elizabeth C Stade, Shannon Wiltsey Stirman, Lyle H Ungar, Cody L Boland, H Andrew Schwartz, David B Yaden, Joao Sedoc, Robert J DeRubeis, Robb Willer, and Johannes C Eichstaedt. 2024. Large language models could change the future of behavioral healthcare: A proposal for respon...

  38. [46]

    Streamlit. 2024. A faster way to build and share data apps. Retrieved 2024-9-4 from https://streamlit.io/

  39. [47]

    Xin Sun, Jan de Wit, Zhuying Li, Jiahuan Pei, Abdallah El Ali, and Jos A Bosch. 2024. Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for Psychotherapy. arXiv preprint arXiv:2411.06723 (2024)

  40. [48]

    Jessica Y Suzuki. 2018. A Qualitative Investigation of Psychotherapy Clients’ Perceptions of Positive Regard . Columbia University

  41. [49]

    David R Thomas. 2006. A general inductive approach for analyzing qualitative evaluation data. American journal of evaluation 27, 2 (2006), 237–246. https://doi.org/10.1177/109821400528374

  42. [50]

    Charles B Truax. 1970. Therapist’s evaluative statements and patient outcome in psychotherapy. Journal of Clinical Psychology 26, 4 (1970)

  43. [51]

    Conal Twomey, Gary O’Reilly, and Michael Byrne. 2015. Effectiveness of cognitive behavioural therapy for anxiety and depression in primary care: a meta-analysis. Family practice 32, 1 (2015), 3–15

  44. [52]

    Marleen Van de Kerkhof. 2006. Making a difference: on the constraints of consensus building and the relevance of deliberation in stakeholder dialogues. Policy Sciences 39, 3 (2006), 279–299

  45. [53]

    Kimberly A Van Orden, Tracy K Witte, Kelly C Cukrowicz, Scott R Braithwaite, Edward A Selby, and Thomas E Joiner Jr. 2010. The interpersonal theory of suicide. Psychological review 117, 2 (2010), 575

  46. [54]

    Maarten Van Someren, Yvonne F Barnard, and J Sandberg. 1994. The think aloud method: a practical approach to modelling cognitive. London: AcademicPress 11, 6 (1994). Manuscript submitted to ACM 19

  47. [55]

    Milton L Wainberg, Pamela Scorza, James M Shultz, Liat Helpman, Jennifer J Mootz, Karen A Johnson, Yuval Neria, Jean-Marie E Bradford, Maria A Oquendo, and Melissa R Arbuckle. 2017. Challenges and opportunities in global mental health: a research-to-practice perspective. Curre...

  48. [56]

    Xiaomeng Wang, Dharmendra Sharma, and Dinesh Kumar. 2024. Cognitive Reframing via Large Language Models for Enhanced Linguistic Attributes. In The Second Tiny Papers Track at ICLR 2024 . https://openreview.net/forum?id=IIus8CSwsU

  49. [57]

    Mengxi Xiao, Qianqian Xie, Ziyan Kuang, Zhicheng Liu, Kailai Yang, Min Peng, Weiguang Han, and Jimin Huang. 2024. HealMe: Harnessing Cognitive Reframing in Large Language Models for Psychotherapy. arXiv preprint arXiv:2403.05574 (2024). https://doi.org/10.48550/arXiv.2403.05574

  50. [58]

    Xue-li Yao and Wen Ma. 2017. Question resistance and its management in Chinese psychotherapy. Discourse Studies 19, 2 (2017), 216–233

  51. [59]

    That’s great

    Peng Zhang, Yanhe Deng, Xue Yu, Xin Zhao, and Xiangping Liu. 2016. Social anxiety, stress type, and conformity among adolescents. Frontiers in psychology 7 (2016), 760. Manuscript submitted to ACM 20 Wang et al. A Appendix A.1 Dialog Snippets by Themes I feel anxious about tom...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.