Pith. sign in

REVIEW 2 cited by

Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.20201 v2 pith:NFUGAIQH submitted 2025-05-26 cs.CL

Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations

classification cs.CL
keywords mentalhealthllmsconversationsmulti-turnframeworkcapabilitiesconversation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Limited access to mental healthcare, extended wait times, and increasing capabilities of Large Language Models (LLMs) has led individuals to turn to LLMs for fulfilling their mental health needs. However, examining the multi-turn mental health conversation capabilities of LLMs remains under-explored. Existing evaluation frameworks typically focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations. To address this, we introduce MedAgent, a novel framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and use it to create the Mental Health Sensemaking Dialogue (MHSD) dataset, comprising over 2,200 patient-LLM conversations. Additionally, we present MultiSenseEval, a holistic framework to evaluate the multi-turn conversation abilities of LLMs in healthcare settings using human-centric criteria. Our findings reveal that frontier reasoning models yield below-par performance for patient-centric communication and struggle at advanced diagnostic capabilities with average score of 31%. Additionally, we observed variation in model performance based on patient's persona and performance drop with increasing turns in the conversation. Our work provides a comprehensive synthetic data generation framework, a dataset and evaluation framework for assessing LLMs in multi-turn mental health conversations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CALM-IT: Generating Realistic Long-Form Motivational Interviewing Dialogues with Dual-Actor Conversational Dynamics Tracking

    cs.CL 2026-01 conditional novelty 6.0

    Tracking both client and therapist states as they evolve produces 8,232 generated MI dialogues that outscore three baselines on MI-quality rubrics and stay stable at 100 turns.

  2. User Perceptions of an LLM-Based Chatbot for Cognitive Reappraisal of Stress: Feasibility Study

    cs.HC 2026-01 conditional novelty 4.0

    A GPT-4o chatbot guiding employees through an 11-step reappraisal script was associated with small short-term reductions in self-reported stress and improved stress mindset in an uncontrolled feasibility study.