Pith. sign in

REVIEW 7 cited by

Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.02244 v2 pith:KMQLHYPL submitted 2024-09-03 cs.HC cs.CL

classification cs.HCcs.CL
keywords counselorshumanllmslanguagesessionssupporttherapycollaboration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) are being used as ad-hoc therapists. Research suggests that LLMs outperform human counselors when generating a single, isolated empathetic response; however, their session-level behavior remains understudied. In this study, we compare the session-level behaviors of human counselors with those of an LLM prompted by a team of peer counselors to deliver single-session Cognitive Behavioral Therapy (CBT). Our three-stage, mixed-methods study involved: a) a year-long ethnography of a text-based support platform where seven counselors iteratively refined CBT prompts through self-counseling and weekly focus groups; b) the manual simulation of human counselor sessions with a CBT-prompted LLM, given the full patient dialogue and contextual notes; and c) session evaluations of both human and LLM sessions by three licensed clinical psychologists using CBT competence measures. Our results show a clear trade-off. Human counselors excel at relational strategies -- small talk, self-disclosure, and culturally situated language -- that lead to higher empathy, collaboration, and deeper user reflection. LLM counselors demonstrate higher procedural adherence to CBT techniques but struggle to sustain collaboration, misread cultural cues, and sometimes produce "deceptive empathy," i.e., formulaic warmth that can inflate users' expectations of genuine human care. Taken together, our findings imply that while LLMs might outperform counselors in generating single empathetic responses, their ability to lead sessions is more limited, highlighting that therapy cannot be reduced to a standalone natural language processing (NLP) task. We call for carefully designed human-AI workflows in scalable support: LLMs can scaffold evidence-based techniques, while peers provide relational support. We conclude by mapping concrete design opportunities and ethical guardrails for such hybrid systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles

    cs.HC 2026-04 unverdicted novelty 6.0 of 10

    LLMs produce interpretive closure in 87.5% of ambiguous social scenarios through narrative alignment, reversal, or normative advice, with first-person perspectives increasing alignment tendencies.

  2. "I'm Not Able to Be There for You": Emotional Labour, Responsibility, and AI in Peer Support

    cs.HC 2026-04 conditional novelty 6.0 of 10

    For sparse G(n,d/n), Δ(G^r)∼log n/log^{(r+1)} n and χ(G^2)=Δ(G)+1; for denser p, χ(G^r)=Θ(d^r/log d) w.h.p.

  3. "I'm Not Able to Be There for You": Emotional Labour, Responsibility, and AI in Peer Support

    cs.HC 2026-04 unverdicted novelty 5.0 of 10

    Peer supporters bear concentrated emotional labor from institutional ambiguity and judge AI by its effects on redistributing responsibility and risk within fragile support roles.

  4. "I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    An interview study with 20 Singapore peer supporters maps their emotional, organisational, and sociocultural practices and derives design directions for culturally responsive digital tools and responsible AI augmentat...

  5. "Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    Mixed-methods studies of an LLM-supported peer support system uncover systematic misalignments where mental health experts flag critical safety and fidelity issues in peer responses that the supporters themselves do n...

  6. Tell Me: An LLM-powered Mental Well-being Assistant with RAG, Synthetic Dialogue Generation, and Agentic Planning

    cs.CL 2025-11 conditional novelty 4.0 of 10

    An open mental-well-being chatbot demo with retrieval-augmented responses, synthetic dialogue generation, and agentic self-care planning reports a small human study favoring its retrieval-augmented version.

  7. Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects

    cs.HC 2025-10 unverdicted novelty 2.0 of 10

    A holistic survey of affective computing for intelligent agents covering emotion understanding via multimodal data, affective cognition, emotional expression synthesis, key challenges, and future directions emphasizin...

Pith tools