Pith. sign in

REVIEW 6 cited by

A Computational Framework for Behavioral Assessment of LLM Therapists

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00820 v2 pith:W2O7KMVL submitted 2024-01-01 cs.CL cs.HC

classification cs.CLcs.HC
keywords therapyllmstherapistsbehaviorlow-qualityassesscareclients
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of large language models (LLMs) like ChatGPT has increased interest in their use as therapists to address mental health challenges and the widespread lack of access to care. However, experts have emphasized the critical need for systematic evaluation of LLM-based mental health interventions to accurately assess their capabilities and limitations. Here, we propose BOLT, a proof-of-concept computational framework to systematically assess the conversational behavior of LLM therapists. We quantitatively measure LLM behavior across 13 psychotherapeutic approaches with in-context learning methods. Then, we compare the behavior of LLMs against high- and low-quality human therapy. Our analysis based on Motivational Interviewing therapy reveals that LLMs often resemble behaviors more commonly exhibited in low-quality therapy rather than high-quality therapy, such as offering a higher degree of problem-solving advice when clients share emotions. However, unlike low-quality therapy, LLMs reflect significantly more upon clients' needs and strengths. Our findings caution that LLM therapists still require further research for consistent, high-quality care.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. DiaCBT: A Long-Periodic Dialogue Corpus Guided by Cognitive Conceptualization Diagram for CBT-based Psychological Counseling

    cs.CL 2025-09 conditional novelty 6.0 of 10

    DiaCBT introduces 108 multi-session CBT counseling cases with CCD-guided generation; a Qwen2.5-7B model fine-tuned on it outperforms prior chatbots on simulated and human evaluation.

  2. Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses

    cs.AI 2025-08 conditional novelty 6.0 of 10

    LLMs only become competitive at spatial data integration when given pre-computed geometric features; a review-and-refine prompt then exceeds hand-tuned heuristics.

  3. Visually grounded emotion regulation via diffusion models and user-driven reappraisal

    cs.LG 2025-07 conditional novelty 6.0 of 10

    AI-generated images made from a person's own spoken reappraisal reduced self-reported negative affect more than reappraisal alone in a 20-person lab experiment.

  4. Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial

    cs.CY 2025-07 conditional novelty 6.0 of 10

    In 915 unscreened adults, a single tailored GPT-4.1 message matched an MI chatbot on colorectal screening intent and beat expert materials for stool-test intent but not colonoscopy intent.

  5. "Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions

    cs.HC 2025-06 unverdicted novelty 6.0 of 10

    Mixed-methods studies of an LLM-supported peer support system uncover systematic misalignments where mental health experts flag critical safety and fidelity issues in peer responses that the supporters themselves do n...

  6. Behavioral Fingerprinting of Large Language Models

    cs.CL 2025-09 reject novelty 5.0 of 10

    Reports a behavioral fingerprint framework for 18 LLMs, claiming reasoning converges while alignment behaviors diverge, but the measurements rest on a single unvalidated judge that is itself one of the graded models.

Pith tools