REVIEW 15 cited by
Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Creating effective dialogue systems for mental health support requires high-quality multi-turn counseling dialogue data, yet collecting real counselor-client conversations presents significant challenges, including privacy concerns, high costs, and limited scalability. We present \textbf{Interactive Agents}, a novel framework that simulates naturalistic counseling dialogues through controlled LLM-to-LLM interactions. The framework introduces two key innovations: (1) a personalized client agent that maintains consistent psychological characteristics throughout a session, and (2) a counselor agent that implements a theoretically grounded three-stage therapeutic model comprising the exploration, insight, and action phases. Through rigorous evaluation using both automatic metrics and professional-counselor assessments based on the Working Alliance Inventory, we demonstrate that our framework generates therapeutically valid dialogues that are comparable in quality to human-generated sessions. Models fine-tuned on our proposed synthetic dataset (SimPsyDial) achieve state-of-the-art performance in a standard pairwise chatbot-arena evaluation of LLM-based counselors. Our framework provides a scalable, privacy-preserving method for generating high-quality counseling dialogue data while maintaining professional therapeutic standards.
Forward citations
Cited by 15 Pith papers
-
A Comprehensive Study of Implementation Bugs in Multi-modal Agents
First systematic taxonomy of 158 multi-modal agent bugs plus a runtime analyzer that recovers most open issues and surfaces 31 new ones.
-
When Seekers Are Hard to Help: Evaluating Emotional Support Dialogue Systems in Worst-Case Interactions
Worst-case seeker simulations show that emotional support dialogue systems suffer substantial performance drops, with large general LLMs more robust than specialized models but still limited in sustaining engagement.
-
ProEvent: An Event-centric Benchmark for Proactive Agents
ProEvent is a benchmark showing LLM agents keep a user's event timetable from chats poorly, with the best fully-correct score at 27.2%.
-
GenPT: Beyond Self-Report for Reliable LLM Psychometrics via Generative Projective Testing
GenPT applies generative projective testing to LLM agents and reports lower directional bias plus greater longitudinal sensitivity than self-report questionnaires.
-
Reframe Your Life Story: Interactive Narrative Therapist and Innovative Moment Assessment with Large Language Models
A stage-planning LLM therapist, INT, and an Innovative Moment metric, IMA, improve narrative therapy dialogue quality over direct LLM role-playing in simulated and human evaluations.
-
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
Dialogue agents aligned via DPO on preference pairs mined from simulated conversations improve engagement scores against the same simulator, with smaller and partially inconsistent human evaluation evidence.
-
Psychological Counseling Cannot Be Achieved Overnight: Automated Psychological Counseling Through Multi-Session Conversations
A new multi-session CBT counseling dataset and model show improved simulated client outcomes over single-session baselines, but the fully LLM-based evaluation leaves real-world validity unproven.
-
Examining Spanish Counseling with MIDAS: a Motivational Interviewing Dataset in Spanish
MIDAS, a new expert-annotated Spanish motivational interviewing dataset, reveals language-specific counselor behaviors and supports Spanish-language behavior classification.
-
Exploring the Inquiry-Diagnosis Relationship with Advanced Patient Simulators
A dialogue-strategy-trained patient simulator improves realism in AI medical consultations and shows that inquiry quality and diagnostic skill jointly limit diagnostic accuracy.
-
Resonant Minds: Closed-Loop Social Avatars with Theory of Mind
A dual-agent closed-loop system integrates Theory of Mind reasoning with multimodal video generation to create social avatars that outperform full-information baselines on dialogue quality under information asymmetry.
-
LLM4Sweat: A Trustworthy Large Language Model for Hyperhidrosis Support
LLM4Sweat reports high accuracy on hyperhidrosis MCQ tasks after fine-tuning on synthetic data generated from the test set, conflating memorization with generalization.
-
H2HTalk: Evaluating Large Language Models as Emotional Companion
H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.
-
Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
Current large language models show stigma and give clinically inappropriate responses to common mental health symptoms, so they should not be deployed as replacement therapists.
-
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
A survey and position paper mapping privacy threats in mental health AI and recommending a pipeline of anonymization, synthetic data, and differential privacy.
-
Scientific Hypothesis Generation and Validation: Methods, Datasets, and Future Directions
A survey of LLM-based hypothesis generation and validation whose taxonomy is useful in outline but whose citations and tool descriptions are unreliable.
Discussion (0). Continue with ORCID to comment.