REVIEW 6 cited by
Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) are being used as ad-hoc therapists. Research suggests that LLMs outperform human counselors when generating a single, isolated empathetic response; however, their session-level behavior remains understudied. In this study, we compare the session-level behaviors of human counselors with those of an LLM prompted by a team of peer counselors to deliver single-session Cognitive Behavioral Therapy (CBT). Our three-stage, mixed-methods study involved: a) a year-long ethnography of a text-based support platform where seven counselors iteratively refined CBT prompts through self-counseling and weekly focus groups; b) the manual simulation of human counselor sessions with a CBT-prompted LLM, given the full patient dialogue and contextual notes; and c) session evaluations of both human and LLM sessions by three licensed clinical psychologists using CBT competence measures. Our results show a clear trade-off. Human counselors excel at relational strategies -- small talk, self-disclosure, and culturally situated language -- that lead to higher empathy, collaboration, and deeper user reflection. LLM counselors demonstrate higher procedural adherence to CBT techniques but struggle to sustain collaboration, misread cultural cues, and sometimes produce "deceptive empathy," i.e., formulaic warmth that can inflate users' expectations of genuine human care. Taken together, our findings imply that while LLMs might outperform counselors in generating single empathetic responses, their ability to lead sessions is more limited, highlighting that therapy cannot be reduced to a standalone natural language processing (NLP) task. We call for carefully designed human-AI workflows in scalable support: LLMs can scaffold evidence-based techniques, while peers provide relational support. We conclude by mapping concrete design opportunities and ethical guardrails for such hybrid systems.
Forward citations
Cited by 6 Pith papers
-
Understanding Attitudes and Trust of Generative AI Chatbots for Social Anxiety Support
People with severe social anxiety symptoms report greater trust in and willingness to use GenAI chatbots, valuing emotional connection, while milder-symptom users emphasize technical reliability.
-
Evaluating an LLM-Powered Chatbot for Cognitive Restructuring: Insights from Mental Health Professionals
A GPT-4 chatbot followed cognitive restructuring steps for 19 users, but mental health experts flagged toxic positivity, advice-giving, and context misunderstandings.
-
"It was Mentally Painful to Try and Stop": Design Opportunities for Just-in-Time Interventions for People with Obsessive-Compulsive Disorder in the Real World
An interview study identifies trigger properties, compulsion patterns, and coping-strategy conflicts that should shape future just-in-time OCD self-management technologies.
-
Detecting Conversational Mental Manipulation with Intent-Aware Prompting
Adding per-speaker intent summaries to an LLM prompt reduces false negatives in mental manipulation detection by 30.5% versus zero-shot prompting on the MentalManip dataset.
-
"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore
Volunteer peer supporters in Singapore experience emotional labour, organisational gaps, and ambivalence toward AI, yielding design implications for human-centred support technologies.
-
Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning
An open mental-well-being chatbot demo with retrieval-augmented responses, synthetic dialogue generation, and agentic self-care planning reports a small human study favoring its retrieval-augmented version.
Discussion (0). Continue with ORCID to comment.