REVIEW 6 cited by
A Computational Framework for Behavioral Assessment of LLM Therapists
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The emergence of large language models (LLMs) like ChatGPT has increased interest in their use as therapists to address mental health challenges and the widespread lack of access to care. However, experts have emphasized the critical need for systematic evaluation of LLM-based mental health interventions to accurately assess their capabilities and limitations. Here, we propose BOLT, a proof-of-concept computational framework to systematically assess the conversational behavior of LLM therapists. We quantitatively measure LLM behavior across 13 psychotherapeutic approaches with in-context learning methods. Then, we compare the behavior of LLMs against high- and low-quality human therapy. Our analysis based on Motivational Interviewing therapy reveals that LLMs often resemble behaviors more commonly exhibited in low-quality therapy rather than high-quality therapy, such as offering a higher degree of problem-solving advice when clients share emotions. However, unlike low-quality therapy, LLMs reflect significantly more upon clients' needs and strengths. Our findings caution that LLM therapists still require further research for consistent, high-quality care.
Forward citations
Cited by 6 Pith papers
-
DiaCBT: A Long-Periodic Dialogue Corpus Guided by Cognitive Conceptualization Diagram for CBT-based Psychological Counseling
DiaCBT introduces 108 multi-session CBT counseling cases with CCD-guided generation; a Qwen2.5-7B model fine-tuned on it outperforms prior chatbots on simulated and human evaluation.
-
Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
LLMs only become competitive at spatial data integration when given pre-computed geometric features; a review-and-refine prompt then exceeds hand-tuned heuristics.
-
Visually grounded emotion regulation via diffusion models and user-driven reappraisal
AI-generated images made from a person's own spoken reappraisal reduced self-reported negative affect more than reappraisal alone in a 20-person lab experiment.
-
Effect of Static vs. Conversational AI-Generated Messages on Colorectal Cancer Screening Intent: a Randomized Controlled Trial
In 915 unscreened adults, a single tailored GPT-4.1 message matched an MI chatbot on colorectal screening intent and beat expert materials for stool-test intent but not colonoscopy intent.
-
"Is This Really a Human Peer Supporter?": Misalignments Between Peer Supporters and Experts in LLM-Supported Interactions
Mixed-methods studies of an LLM-supported peer support system uncover systematic misalignments where mental health experts flag critical safety and fidelity issues in peer responses that the supporters themselves do n...
-
Behavioral Fingerprinting of Large Language Models
Reports a behavioral fingerprint framework for 18 LLMs, claiming reasoning converges while alignment behaviors diverge, but the measurements rest on a single unvalidated judge that is itself one of the graded models.
Discussion (0). Sign in to comment.