Pith. sign in

REVIEW 3 major objections 5 minor 50 references

Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read In a fixed set of Japanese AI-to-AI counseling dialogues, expert raters preferred the structured-prompt condition over a minimal prompt on four of five quality scales, while newer LLM judges remained systematically lenient.

desk verdict A transparent, well-scoped Japanese AI-counseling benchmark with a robust LLM-calibration caution, but the expert reference needs a nonauthor-only sensitivity check. read the letter →

arxiv 2507.02950 v3 pith:PRFSMRN3 submitted 2025-06-28 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords motivationalinterviewinglargelanguagemodelsstructuredpromptingJapanese-languagecounselingautomatedevaluationMITIglobalratingsAI-to-AIdialoguesimulation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper examines whether a structured dialogue prompt changes how counselors in an AI-to-AI simulation are rated by human experts, and whether newer LLMs can rate counseling transcripts in place of experts. In 18 fixed Japanese-language transcripts, expert raters judged dialogues produced under the Structured Multi-step Dialogue Prompt (SMDP) higher on cultivating change talk, partnership, empathy, and overall quality than dialogues produced with a minimal instruction, while the two SMDP counselor models did not differ. The same transcripts were then rated repeatedly by three newer LLMs, which produced reproducible but systematically more lenient scores than the expert panel, most visibly for softening sustain talk and overall quality. The paper's message is that prompt structure can shape observable counseling behavior in a fixed stimulus set, but automated evaluations still need expert-anchored calibration before they can serve as counseling-quality evidence.

What carries the argument

The central mechanism is the Structured Multi-step Dialogue Prompt (SMDP), which imposes both a macro-level question sequence (present concern, ideal outcome, prior attempts, current efforts, resources, immediate steps, and first-step commitment) and a micro-level turn rule requiring two reflective paraphrases before any question. This procedural scaffolding is what the paper varies against a minimal role instruction within the same counselor model, and it is the manipulation that carried the expert-rated advantages in cultivating change talk, partnership, empathy, and overall quality. On the measurement side, the paper's machinery is the adapted MITI global rating scales plus an overall-quality item, used by experts and by three LLM evaluators, which let the authors separate three properties that are often conflated: repeatability across scoring iterations, calibration to expert-reference levels, and discrimination among conditions.

What would settle it

A decisive test would be a preregistered replication with a fully crossed panel of formally trained, calibrated raters, randomized evaluator assignment, and fresh independently generated transcripts per condition; if the SMDP-condition dialogues are not rated higher on change talk, partnership, empathy, and overall quality, the paper's central claim about prompt structure would be overturned, and if LLM evaluators provided with expert-scored anchor exemplars no longer overrate sustain talk and overall quality, the leniency finding would be shown to be a calibration artifact rather than a fixed property of the models.

Watch

Extended reading notes

Core claim

Within this controlled set of 18 synthetic Japanese-language counseling dialogues, the paper's central discovery is that the SMDP condition—a macro-level script moving from the present concern to a desired future, prior attempts, resources, and next steps, plus a micro-rule requiring two reflective listen-backs before each question—produced dialogues that fifteen counseling experts rated higher for change-talk cultivation, partnership, empathy, and overall quality than the minimally prompted counselor. The contrast between the two GPT-4-turbo conditions isolates prompt structure, since model family is held constant, and the paper reports that this direct prompt comparison showed the advantage on four of five outcomes. A parallel finding is that three newer LLM evaluators, though highly repeatable across scoring iterations, were lenient relative to the expert reference, especially on softening sustain talk and overall quality, with one evaluator system closer to the expert reference on change talk and partnership than the others. The paper presents the result as an expert-referenced benchmark for Japanese-language AI counseling simulations rather than as evidence of clinical effectiveness or of general prompt effects across all possible dialogues.

Load-bearing premise

The central claim rests on the assumption that the aggregated ratings of the fifteen experts are a valid and comparable yardstick for counseling quality and for calibrating the automated judges; if those ratings were noisy or biased, both the prompt advantage and the LLM leniency findings could be artifacts.

Editorial extensions

If this is right

  • Structured prompting appears to be a usable design element for Japanese-language counseling simulations: the SMDP-condition dialogues were preferred by expert raters on four of five quality dimensions in this fixed set.
  • Automated LLM ratings should not be treated as calibrated counseling-quality evidence on their own, because reproducible scores in this study were systematically more lenient than expert ratings, particularly for softening sustain talk and overall quality.
  • The two SMDP counselor models, one GPT-based and one Claude-based, did not differ on any expert-rated outcome, so within this set the prompt framework, rather than model family, was the visible driver of the gain over the minimal condition.
  • The softness of the softening-sustain-talk result means that construct cannot be fairly benchmarked until simulated clients provide more resistance and ambivalence; expert agreement on that scale was very low.
  • Simulated clients were rated below the naturalness midpoint, so the stimulus set is best understood as a controlled benchmark for training materials, not as realistic human-client interaction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A directly testable extension is to hold client utterances constant across counselor conditions; the present design evaluates whole interactions, so part of the SMDP advantage may come from the client responses it elicits rather than from counselor behavior alone.
  • The SMDP rule of two reflections before each question could be tested separately from the macro-sequence: a component study that keeps the turn rule while scrambling the question order would tell whether the ratings come from pacing or from the evocation of change talk.
  • The 18-transcript benchmark could be reused as a fixed calibration set for LLM evaluators: anchoring prompts with expert-rated exemplars and asking for score justifications might shrink the systematic leniency observed for sustain talk and overall quality.
  • Because expert agreement on softening sustain talk was near zero and clients were compliant, a richer client simulator with hesitation, disagreement, and negative affect is the natural precondition for any future SST benchmark.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports an evaluation of 18 fixed Japanese-language AI-to-AI counseling transcripts generated under three counselor conditions: GPT-minimal, GPT-SMDP, and Claude-SMDP. Fifteen counseling experts rated each transcript on four adapted MITI-style global scales and an overall-quality item, and three newer LLM evaluators independently rated the same transcripts three times each. The central claims are that, within this fixed stimulus set, the two SMDP-condition dialogue sets received higher expert ratings than GPT-minimal on cultivating change talk, partnership, empathy, and overall quality, that the two SMDP counselor models did not differ reliably, and that LLM ratings were reproducible but systematically more lenient than expert-reference ratings, particularly for softening sustain talk and overall quality. The paper carefully limits its inference to the fixed transcript set and does not claim clinical effectiveness or generalization to the broader population of possible dialogues.

Significance. If the central expert-rating contrasts hold, the paper makes a useful contribution by providing a Japanese-language, expert-referenced benchmark for structured-prompt counseling simulations and by separating three evaluator properties that are often conflated: repeatability, calibration to expert reference levels, and discrimination among conditions. The preregistration, public data and code archive, fixed-stimulus paired-difference sensitivity analyses, ordinal-model checks, and explicit scoping caveats are genuine strengths. The main significance risk is that every substantive contrast and every LLM calibration comparison is anchored to an expert reference whose validity is not independently established; this is a fixable but load-bearing weakness rather than a circularity in the derivations.

major comments (3)
  1. [Section 3.2; Section 5.4] The expert reference is the sole basis for the central counselor-condition contrasts and for LLM calibration, yet the panel included coauthor evaluators who knew the design and the fixed Moon/Star/Sun-to-condition mapping. Display-order randomization does not blind an evaluator who knows that Moon always denotes Claude-SMDP, Star always denotes GPT-minimal, and Sun always denotes GPT-SMDP. The manuscript acknowledges this possibility in Section 5.4, but it does not provide the empirical check that would settle it. I request a nonauthor-only subgroup analysis: re-estimate the expert-only mixed models and the within-evaluation-unit paired-difference contrasts for CCT, PAR, EMP, and OVR using only nonauthor evaluators, and report whether the adjusted effect sizes and multiplicity-adjusted p-values remain comparable to those in Table 3 and S6.7. If the nonauthor-only contrasts are attenuated or nonsignificant, the main claim would lose its independent support.
  2. [Section 4.4; Supporting Information S5.5 and S6.11] The LLM calibration claim is quantified by transcript-level LLM-minus-expert differences with bootstrap intervals in S6.11, but those intervals resample transcripts, not expert raters. Given the low single-rater ICCs reported in S5.5 (0.20-0.26 for CCT, PAR, EMP, OVR; 0.03 for SST) and the incomplete rating matrix with no formal rater calibration, the expert-reference means used in S6.11 are noisy, and the reported intervals likely understate uncertainty about the true expert reference. Please add a rater-resampling bootstrap or an evaluator-random-effect calibration model so that the precision of the LLM-minus-expert differences reflects both transcript and rater variation, and interpret the calibration results in light of the resulting wider intervals.
  3. [Section 3.1; Section 5.4] Because client utterances were generated dynamically in response to each counselor, the condition contrast is a contrast of complete counselor-client interactions rather than of counselor responses to identical client turns. The manuscript states this limitation clearly and uses a consistent estimand, so this is not an error. However, the Abstract's phrasing 'structured prompting improves observable counseling quality' can be read as a causal prompt effect. I recommend that the Abstract and Conclusion consistently use the interaction-level wording ('SMDP-condition dialogues were rated higher'), which the paper mostly does, to prevent an unintended causal reading that the design cannot support.
minor comments (5)
  1. [Section 4.2; Table 3] The text states that the omnibus counselor-condition term was significant for all five outcomes, but Table 3 does not report the omnibus likelihood-ratio statistics. Adding chi-square values, degrees of freedom, and p-values for the five omnibus tests would make the table self-contained.
  2. [Figure 1] The caption says error bars show 95% confidence intervals but does not state whether these are model-based intervals from the mixed model or raw-script means. Please specify the method and, if the bars are raw means, clarify that they do not account for evaluator and transcript random effects.
  3. [Section 4.2] The SST primary contrast is reported as 'narrowly missed' the adjusted threshold (p = .051). Since the paired-difference sensitivity analysis also leaves SST uncertain and the expert SST ICC is 0.03, I suggest stating more directly that SST is not interpretable as evidence for or against a condition effect, rather than describing it only as directionally consistent.
  4. [S5.5] The average-measure ICC(2,k) values for the expert panel are labeled as requiring caution because no common panel rated every transcript. This caveat is good; I suggest moving that sentence into the main text near Section 4.4, since readers may otherwise mistake the average-measure values for the reliability of the observed 8-10-rater means.
  5. [Section 3.3] The provider-native message-role differences are acknowledged in the text and S3.3. The term 'identical task prompts and rubrics' in S3.3 is accurate, but the subsequent sentence about message-role structure could be placed immediately after the first mention to avoid any impression of full procedural identity across providers.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the expert ratings and LLM calibration are external references, and the prompt contrast is measured, not fitted.

full rationale

The paper's derivation chain is empirical rather than analytic: SMDP and GPT-minimal prompts were fixed inputs, the 18 transcripts were generated by AI-to-AI interaction, and the outcomes were expert ratings and LLM ratings of those transcripts. No equation or model is fitted against its own output, and no prediction is statistically forced by construction. The LLM calibration analysis compares LLM scores to expert-reference means, which are separate measurements, not fitted parameters from the LLM runs. The central claim that SMDP-condition dialogues received higher expert ratings than GPT-minimal on CCT, PAR, EMP, and OVR is a direct comparison of observed ratings, not a quantity derived from the definition of SMDP or from the rating scales. The only near-self-citation is Ohtsubo et al. (2016), which includes a coauthor and supports rater-training considerations, but it is not load-bearing for any main contrast. Acknowledged limitations, including coauthor participation in the expert panel, nonrandomized evaluator assignment, low single-rater ICCs, and the absence of a formal rater-calibration exercise, are threats to validity and independence, not circularity of derivation. The paper also explicitly frames its results as evaluations of a fixed stimulus set and disclaims generalization, which further separates the observed findings from any fitted or definitionally forced claim. Accordingly, no circular step meeting the quoted-evidence standard can be identified.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The study does not fit numerical parameters to data in the sense of the ledger; the statistical models estimate fixed effects and variance components from the observed ratings. The main unstated costs are domain assumptions about the validity of adapted MITI scales, the adequacy of an uncalibrated partially coauthor expert panel as a reference, and the interpretability of the fixed 18-transcript design. No invented entities are introduced.

assumptions (4)
  • domain assumption Adapted MITI global rating scales are valid for assessing AI-generated counseling dialogues in Japanese.
    The study uses four adapted MITI 4.2.1 global scales, acknowledged as not full MITI fidelity coding because behavior counts and human-session calibration are absent (Section 3.2, S2.1).
  • domain assumption The panel of 15 experts, despite no formal calibration and partially coauthor composition, provides a usable reference for counseling quality.
    Expert ratings are the reference against which LLM scores are calibrated; the paper reports low ICC(2,1) values and no documented rater training (Section 4.4, Section 5.4).
  • domain assumption The fixed 18-transcript design, with one transcript per counselor-by-client cell, supports inference about expert evaluations of this set.
    The paper explicitly limits claims to the fixed stimulus set and does not generalize to the population of possible dialogues (Section 3.4, Section 5.4).
  • domain assumption Client utterances generated in response to each counselor condition do not invalidate counselor-condition contrasts.
    The estimand is the complete interaction; client speech is not held constant across conditions, a confound the paper acknowledges and interprets around (Section 3.1, Section 5.4).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues." pith.science (2026). https://pith.science/paper/PRFSMRN3

@misc{pith2026250702950,
  author       = {Pith},
  title        = {Pith review of: Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PRFSMRN3}},
  note         = {Machine review of arXiv:2507.02950}
}
read the original abstract

Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings remains limited. We examined 18 fixed Japanese-language counseling transcripts generated through artificial intelligence (AI)-to-AI interactions under three counselor conditions: GPT-minimal (GPT-4-turbo with a minimal role instruction), GPT-SMDP (GPT-4-turbo with the Structured Multi-step Dialogue Prompt [SMDP]), and Claude-SMDP (Claude-3-Opus with SMDP). Fifteen counseling experts rated transcripts on four adapted global scales from the Motivational Interviewing Treatment Integrity coding manual and an overall-quality item; three newer LLMs independently rated the same transcripts in three iterations. In this fixed stimulus set, SMDP-condition dialogues received higher expert ratings for cultivating change talk, partnership, empathy, and overall quality than GPT-minimal dialogues; the two SMDP counselor models did not differ. LLM ratings were reproducible but generally more lenient than expert-reference ratings, particularly for softening sustain talk and overall quality. Simulated-client naturalness was below the scale midpoint. These findings provide an expert-referenced benchmark for Japanese-language AI counseling simulations and show that reproducible LLM ratings should not be treated as calibrated counseling-quality evidence without expert validation. This study does not test clinical effectiveness or human-client outcomes.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

50 extracted references · 30 canonical work pages

  1. [1]

    Perils and opportunities in using large language models in psychological research

    Suhaib Abdurahman, Mohammad Atari, Farzan Karimi-Malekabadi , Mona J Xue, Jackson Trager, Peter S Park, Preni Golazizian, Ali Omrani, and Morteza Dehghani. Perils and opportunities in using large language models in psychological research. PNAS Nexus, 3 0 (7): 0 pgae245, July 2024. ISSN 2752-6542. doi:10.1093/pnasnexus/pgae245

  2. [2]

    Guilherme F. C. F. Almeida, Jos \'e Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo \ de\ Ara \'u jo. Exploring the psychology of LLMs ' moral and legal reasoning. Artificial Intelligence, 333: 0 104145, August 2024. ISSN 0004-3702. doi:10.1016/j.artint.2024.104145

  3. [3]

    Habit Coach : Customising RAG-based chatbots to support behavior change, December 2024

    Arian Fooroogh Mand Arabi, Cansu Koyuturk, Michael O'Mahony, Raffaella Calati, and Dimitri Ognibene. Habit Coach : Customising RAG-based chatbots to support behavior change, December 2024

  4. [4]

    SFMSS : Service Flow aware Medical Scenario Simulation for Conversational Data Generation

    Zhijie Bao, Qingyun Liu, Xuanjing Huang, and Zhongyu Wei. SFMSS : Service Flow aware Medical Scenario Simulation for Conversational Data Generation . In Luis Chiruzzo, Alan Ritter, and Lu Wang, editors, Findings of the Association for Computational Linguistics : NAACL 2025 , pages 4586--4604, Albuquerque, New Mexico, April 2025. Association for Computatio...

  5. [5]

    Exploring the Efficacy of Robotic Assistants with ChatGPT and Claude in Enhancing ADHD Therapy : Innovating Treatment Paradigms

    Santiago Berrezueta-Guzman , Mohanad Kandil, Mar \'i a-Luisa Mart \'i n-Ruiz , Iv \'a n Pau \ de la\ Cruz , and Stephan Krusche. Exploring the Efficacy of Robotic Assistants with ChatGPT and Claude in Enhancing ADHD Therapy : Innovating Treatment Paradigms . In 2024 International Conference on Intelligent Environments ( IE ) , pages 25--32, June 2024. doi...

  6. [6]

    Prompt engineering, 2024

    Lee Boonstra. Prompt engineering, 2024

  7. [7]

    Empowering Psychotherapy with Large Language Models : Cognitive Distortion Detection through Diagnosis of Thought Prompting

    Zhiyu Chen, Yujie Lu, and William Wang. Empowering Psychotherapy with Large Language Models : Cognitive Distortion Detection through Diagnosis of Thought Prompting . In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Findings of the Association for Computational Linguistics : EMNLP 2023 , pages 4295--4304, Singapore, February 2023. Association for Com...

  8. [8]

    Motivational Interviewing Transcripts Annotated with Global Scores

    Ben Cohen, Moreah Zisquit, Stav Yosef, Doron Friedman, and Kfir Bar. Motivational Interviewing Transcripts Annotated with Global Scores . In Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings of the 2024 Joint International Conference on Computational Linguistics , Language Resources ...

Show all 50 references
  1. [9]

    Cook, Joshua Overgaard, V

    David A. Cook, Joshua Overgaard, V. Shane Pankratz, Guilherme Del Fiol, and Chris A. Aakre. Virtual Patients Using Large Language Models : Scalable , Contextualized Simulation of Clinician-Patient Dialogue With Feedback . Journal of Medical Internet Research, 27 0 (1): 0 e6848...

  2. [10]

    Yeager, Christopher J

    Dorottya Demszky, Diyi Yang, David S. Yeager, Christopher J. Bryan, Margarett Clapper, Susannah Chandhok, Johannes C. Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann Johnson, Michaela Jones, Danielle Krettek-Cobb , Leslie Lai, Nirel JonesMitchell, Desmond C. Ong, Carol S. ...

  3. [11]

    LLMs Can Simulate Standardized Patients via Agent Coevolution , December 2024

    Zhuoyun Du, Lujie Zheng, Renjun Hu, Yuyang Xu, Xiawei Li, Ying Sun, Wei Chen, Jian Wu, Haolei Cai, and Haohao Ying. LLMs Can Simulate Standardized Patients via Agent Coevolution , December 2024

  4. [12]

    Determinants of LLM-assisted Decision-Making , February 2024

    Eva Eigner and Thorsten H \"a ndler. Determinants of LLM-assisted Decision-Making , February 2024

  5. [13]

    From LLM to NMT : Advancing Low-Resource Machine Translation with Claude , April 2024

    Maxim Enis and Mark Hopkins. From LLM to NMT : Advancing Low-Resource Machine Translation with Claude , April 2024

  6. [14]

    The architecture of language: Understanding the mechanics behind LLMs

    Andrea Filippo Ferraris, Davide Audrito, Luigi Di Caro, and Cristina Poncib \`o . The architecture of language: Understanding the mechanics behind LLMs . Cambridge Forum on AI: Law and Governance, 1: 0 e11, January 2025. ISSN 3033-3733. doi:10.1017/cfl.2024.16

  7. [15]

    The Pile : An 800GB Dataset of Diverse Text for Language Modeling , December 2020

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The Pile : An 800GB Dataset of Diverse Text for Language Modeling , December 2020

  8. [16]

    Increasing Realism and Variety of Virtual Patient Dialogues for Prenatal Counseling Education Through a Novel Application of ChatGPT : Exploratory Observational Study

    Megan Gray, Austin Baird, Taylor Sawyer, Jasmine James, Thea DeBroux, Michelle Bartlett, Jeanne Krick, and Rachel Umoren. Increasing Realism and Variety of Virtual Patient Dialogues for Prenatal Counseling Education Through a Novel Application of ChatGPT : Exploratory Observat...

  9. [17]

    Machine psychology

    Thilo Hagendorff, Ishita Dasgupta, Marcel Binz, Stephanie CY Chan, Andrew Lampinen, Jane X Wang, Zeynep Akata, and Eric Schulz. Machine psychology. arXiv preprint arXiv:2303.13988, 2023

  10. [18]

    Helping the Helper : Supporting Peer Counselors via AI-Empowered Practice and Feedback , March 2025

    Shang-Ling Hsu, Raj Sanjay Shah, Prathik Senthil, Zahra Ashktorab, Casey Dugan, Werner Geyer, and Diyi Yang. Helping the Helper : Supporting Peer Counselors via AI-Empowered Practice and Feedback , March 2025

  11. [19]

    PsycoLLM : Enhancing LLM for Psychological Understanding and Evaluation

    Jinpeng Hu, Tengteng Dong, Gang Luo, Hui Ma, Peng Zou, Xiao Sun, Dan Guo, Xun Yang, and Meng Wang. PsycoLLM : Enhancing LLM for Psychological Understanding and Evaluation . IEEE Transactions on Computational Social Systems, 12 0 (2): 0 539--551, April 2025. ISSN 2329-924X. doi...

  12. [20]

    Empowerment of Large Language Models in Psychological Counseling through Prompt Engineering

    Shanshan Huang, Fuxiang Fu, Ke Yang, Ke Zhang, and Fan Yang. Empowerment of Large Language Models in Psychological Counseling through Prompt Engineering . In 2024 IEEE 4th International Conference on Software Engineering and Artificial Intelligence ( SEAI ) , pages 220--225, J...

  13. [21]

    Can Large Language Models be Used to Provide Psychological Counselling ? An Analysis of GPT-4-Generated Responses Using Role-play Dialogues , February 2024

    Michimasa Inaba, Mariko Ukiyo, and Keiko Takamizo. Can Large Language Models be Used to Provide Psychological Counselling ? An Analysis of GPT-4-Generated Responses Using Role-play Dialogues , February 2024

  14. [22]

    Response Generation for Cognitive Behavioral Therapy with Large Language Models : Comparative Study with Socratic Questioning , January 2024

    Kenta Izumi, Hiroki Tanaka, Kazuhiro Shidara, Hiroyoshi Adachi, Daisuke Kanayama, Takashi Kudo, and Satoshi Nakamura. Response Generation for Cognitive Behavioral Therapy with Large Language Models : Comparative Study with Socratic Questioning , January 2024

  15. [23]

    Onno P. Kampman, Ye Sheng Phang, Stanley Han, Michael Xing, Xinyi Hong, Hazirah Hoosainsah, Caleb Tan, Genta Indra Winata, Skyler Wang, Creighton Heaukulani, Janice Huiqin Weng, and Robert JT Morris. A Multi-Agent Dual Dialogue System to Support Mental Health Care Providers , ...

  16. [24]

    Exploring the Frontiers of LLMs in Psychological Applications : A Comprehensive Review , March 2024

    Luoma Ke, Song Tong, Peng Cheng, and Kaiping Peng. Exploring the Frontiers of LLMs in Psychological Applications : A Comprehensive Review , March 2024

  17. [25]

    Koo and Mae Y

    Terry K. Koo and Mae Y. Li. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research . Journal of Chiropractic Medicine, 15 0 (2): 0 155--163, June 2016. ISSN 1556-3707. doi:10.1016/j.jcm.2016.02.012

  18. [26]

    Comparative Study on the Performance of LLM-based Psychological Counseling Chatbots via Prompt Engineering Techniques

    Aram Lee, Sehwan Moon, Min Jhon, Ju-Wan Kim, Dae-Kwang Kim, Jeong Eun Kim, Kiwon Park, and Eunkyoung Jeon. Comparative Study on the Performance of LLM-based Psychological Counseling Chatbots via Prompt Engineering Techniques . In 2024 IEEE International Conference on Bioinform...

  19. [27]

    Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs

    Anqi Li, Yu Lu, Nirui Song, Shuai Zhang, Lizhi Ma, and Zhenzhong Lan. Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs . In Yaser Al-Onaizan , Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Associat...

  20. [28]

    Rakesh K. Maurya. A qualitative content analysis of ChatGPT 's client simulation role-play for practising counselling skills. Counselling & Psychotherapy Research, 24 0 (2): 0 614--630, 2024 a . ISSN 1746-1405. doi:10.1002/capr.12699

  21. [29]

    Rakesh K. Maurya. Using AI Based Chatbot ChatGPT for Practicing Counseling Skills Through Role-Play . Journal of Creativity in Mental Health, 19 0 (4): 0 513--528, October 2024 b . ISSN 1540-1383. doi:10.1080/15401383.2023.2297857

  22. [30]

    PAIR : Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational Interviewing

    Do June Min, Ver \'o nica P \'e rez-Rosas , Kenneth Resnicow, and Rada Mihalcea. PAIR : Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational Interviewing . In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference o...

  23. [31]

    LLM Juries for Evaluation , February 2025

    Abby Morgan. LLM Juries for Evaluation , February 2025

  24. [32]

    Motivational interviewing treatment integrity coding manual 4.2

    TB Moyers, JK Manuel, and D Ernst. Motivational interviewing treatment integrity coding manual 4.2. 1., 2014. Unpublished manual. https://motivationalinterviewing. org/sites/default/files/miti4\_2. pdf, 2014

  25. [33]

    Enhancing communication and clinical reasoning in medical education: Building virtual patients with generative AI

    Lewis Potter and Chris Jefferies. Enhancing communication and clinical reasoning in medical education: Building virtual patients with generative AI . Future Healthcare Journal, 11: 0 100043, April 2024. ISSN 2514-6645. doi:10.1016/j.fhj.2024.100043

  26. [34]

    Interactive Agents : Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions , August 2024

    Huachuan Qiu and Zhenzhong Lan. Interactive Agents : Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions , August 2024

  27. [35]

    Rajaraman

    V. Rajaraman. From ELIZA to ChatGPT . Resonance, 28 0 (6): 0 889--905, June 2023. ISSN 0973-712X. doi:10.1007/s12045-023-1620-6

  28. [36]

    An AI-based virtual client for educational role-playing in the training of online counselors

    Eric Rudolph, Natalie Engert, and Jens Albrecht. An AI-based virtual client for educational role-playing in the training of online counselors. In Proceedings of the 16th International Conference on Computer Supported Education - Volume 2: CSEDU , pages 108--117. SciTePress / I...

  29. [37]

    Automated feedback generation in an intelligent tutoring system for counselor education

    Eric Rudolph, Hanna Seer, Carina Mothes, and Jens Albrecht. Automated feedback generation in an intelligent tutoring system for counselor education. In 2024 19th Conference on Computer Science and Intelligence Systems ( FedCSIS ) , pages 501--512, September 2024 b . doi:10.154...

  30. [38]

    Lin, Adam S

    Ashish Sharma, Inna W. Lin, Adam S. Miner, David C. Atkins, and Tim Althoff. Human-- AI collaboration enables more empathic conversations in text-based peer-to-peer mental health support. Nature Machine Intelligence, 5 0 (1): 0 46--57, January 2023. ISSN 2522-5839. doi:10.1038...

  31. [39]

    Expert system methodologies and applications---a decade review from 1995 to 2004

    Shu-Hsien Liao . Expert system methodologies and applications---a decade review from 1995 to 2004. Expert Systems with Applications, 28 0 (1): 0 93--103, January 2005. ISSN 0957-4174. doi:10.1016/j.eswa.2004.08.003

  32. [40]

    Semantic networks

    John F Sowa. Semantic networks. 2: 0 1493--1511

  33. [41]

    Stade, Shannon Wiltsey Stirman, Lyle H

    Elizabeth C. Stade, Shannon Wiltsey Stirman, Lyle H. Ungar, Cody L. Boland, H. Andrew Schwartz, David B. Yaden, Jo \ a o Sedoc, Robert J. DeRubeis, Robb Willer, and Johannes C. Eichstaedt. Large language models could change the future of behavioral healthcare: A proposal for r...

  34. [42]

    Bickmore

    Ian Steenstra, Farnaz Nouraei, Mehdi Arjmand, and Timothy W. Bickmore. Virtual Agents for Alcohol Use Counseling : Exploring LLM-Powered Motivational Interviewing . In Proceedings of the ACM International Conference on Intelligent Virtual Agents , pages 1--10, September 2024. ...

  35. [43]

    The Turbulent Past and Uncertain Future of AI : Is there a way out of AI 's boom-and-bust cycle? IEEE Spectrum, 58 0 (10): 0 26--31, October 2021

    Eliza Strickland. The Turbulent Past and Uncertain Future of AI : Is there a way out of AI 's boom-and-bust cycle? IEEE Spectrum, 58 0 (10): 0 26--31, October 2021. ISSN 1939-9340. doi:10.1109/MSPEC.2021.9563956

  36. [44]

    Tan, L.S

    C.F. Tan, L.S. Wahidin, S.N. Khalil, N. Tamaldin, J. Hu, and G.W.M. Rauterberg. The application of expert system: A review of research and applications. ARPN Journal of Engineering and Applied Sciences, 11 0 (4): 0 2448--2453, 2016. ISSN 1819-6608

  37. [45]

    LaMDA : Language Models for Dialog Applications , February 2022

    Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Q...

  38. [46]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need , August 2023

  39. [47]

    The Rise and Potential of Large Language Model Based Agents : A Survey , September 2023

    Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, W...

  40. [48]

    Wong, and Ruifeng Xu

    Ancheng Xu, Di Yang, Renhao Li, Jingwei Zhu, Minghuan Tan, Min Yang, Wanxin Qiu, Mingchen Ma, Haihong Wu, Bingyu Li, Feng Sha, Chengming Li, Xiping Hu, Qiang Qu, Derek F. Wong, and Ruifeng Xu. AutoCBT : An Autonomous Multi-agent Framework for Cognitive Behavioral Therapy in Ps...

  41. [49]

    CAMI : A Counselor Agent Supporting Motivational Interviewing through State Inference and Topic Exploration , February 2025

    Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Kit Phey Leng, Nicholas Gabriel Lim, Cameron Tan Shi Ern, and Ee-peng Lim. CAMI : A Counselor Agent Supporting Motivational Interviewing through State Inference and Topic Exploration , February 2025

  42. [50]

    Cultural Value Differences of LLMs : Prompt , Language , and Model Size , June 2024

    Qishuai Zhong, Yike Yun, and Aixin Sun. Cultural Value Differences of LLMs : Prompt , Language , and Model Size , June 2024

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.