Pith. sign in

REVIEW 4 major objections 5 minor 43 references

MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read Fully simulated CBT mentorship — LLM patient, trainee, and supervisor — reproduces real trainee competence only in native speech-to-speech; text inflates scores, and feedback helps five of seven models but degrades the two smallest.

desk verdict A large, transparent simulation of CBT mentorship cycles with a real feedback-scale effect, but the native-speech CTRS claim is confounded by a mentor-model swap and unvalidated LLM self-scoring — still worth refereeing. read the letter →

arxiv 2607.25667 v1 pith:DXJ3267B submitted 2026-07-28 cs.CL cs.AI

classification cs.CLcs.AI
keywords largelanguagemodelsspeech-to-speechinteractiondeliberatepracticecognitivebehaviouraltherapypsychotherapytrainingRatingScalemulti-agentsimulationmultimodaldialogue
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that large language models can run a complete cognitive behavioural therapy (CBT) mentorship cycle — a DSM-5-grounded simulated patient, a therapist-in-training, and an expert supervisor — producing 2,100 full training sessions. Simulated patients expressed disorder-congruent emotional profiles that trainee therapists mirrored, as in human counselling. Its sharpest claim is that the native speech-to-speech dyad received Cognitive Therapy Rating Scale (CTRS) scores close to the human reference mean and stayed below the threshold for competent CBT delivery, while text-based and re-synthesized-audio conditions were scored far above human levels — text fluency, the authors argue, creates an illusion of competence. Supervisor feedback improved diagnostic accuracy in five of seven model conditions but degraded it in the two smallest models, which tended to abandon correct diagnoses after feedback, and symptom identification accuracy rose with model size. A sympathetic reader would care because the result points toward scalable, consequence-free, feedback-rich clinical training in a field with a drastic shortage of expert supervisors.

What carries the argument

The load-bearing mechanism is the triadic mentorship cycle: a patient persona grounded in DSM-5-TR clinical case vignettes (one each for major depressive disorder, generalised anxiety disorder, and borderline personality disorder); a trainee-therapist persona calibrated item-by-item to the median competence scores of community clinicians learning CBT, drawn from a published dataset of 1,264 recorded sessions; and a supervisor persona that scores the session transcript on the Cognitive Therapy Rating Scale — the standard eleven-item, 0–66 measure of CBT competence — and returns structured feedback with a reflective question. The CTRS comparison against a human reference distribution is the in

What would settle it

Have blinded human CBT experts score samples of the native-audio and text session transcripts with the same rating scale the AI mentor used. If human raters do not find the text transcripts less competent than the native-audio transcripts — the opposite of what the AI mentor's scores showed — the central claim collapses. A second decisive test: run the mentorship cycle with different model families assigned to patient, trainee, and supervisor; if the native-audio-versus-text competence gap disappears when the scorer is a different model, the result is driven by self-consistency rather than by

Watch

Extended reading notes

Core claim

The central discovery, on the paper's own terms, is that a fully simulated multimodal mentorship loop can reproduce the measured competence profile of real CBT trainees — provided the clinical encounter happens through native speech-to-speech. The native-audio dyad scored 29.97 on the 66-point Cognitive Therapy Rating Scale (CTRS) against a human mean of 31.04 and, like humans, remained below the score-40 threshold conventionally taken to indicate competent CBT delivery; the text and re-synthesized-speech conditions scored between 41 and 56, with several items saturating near the maximum. The authors interpret this as evidence that native audio preserves hesitation, pacing, and prosody — par

Load-bearing premise

The central comparison assumes that the AI supervisor's competence ratings behave like human expert ratings, yet the human reference is reconstructed from published summary statistics and the supervisor shares its underlying model with the therapist it scores — so model self-consistency, not therapeutic skill, could be driving the differences between conditions.

Editorial extensions

If this is right

  • If the central claim holds, scalable deliberate practice for psychotherapy becomes feasible: the environment generated 2,100 complete patient–therapist–supervisor cycles automatically, giving trainees unlimited, consequence-free rehearsal before they face real patients.
  • Text-only and re-synthesized-audio training environments are not neutral: they systematically overestimate trainee competence on the CTRS, so any AI training tool must be validated with native audio or it risks teaching overconfidence.
  • Supervisory feedback is not uniformly beneficial: in the two smallest model conditions, feedback converted correct diagnoses into errors, so a safe training system must calibrate feedback to trainee capability or let trainees retain a justified judgment when feedback is weak.
  • Symptom identification accuracy rising steeply with model size means clinical grounding is a capacity property: small models perform near chance on symptom recognition and should not be deployed as unsupervised trainees.
  • The complete cycle offers a testbed for deciding when synthetic training is faithful, calibrated, and safe enough to be useful — the question the authors frame as the environment's purpose.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the CTRS gap between native audio and text suggests spoken interaction acts as a paralinguistic regularizer — real-time speech forces the model to operate under constraints that strip away the fluency inflating text-based self-assessment; a direct test would run the same model family in live audio versus transcribed text with identical prompts and human expert ratings.
  • Inference: because the same model usually played patient, trainee, and supervisor, the diagnostic-improvement results may partly reflect self-consistency rather than genuine pedagogical transfer; assigning different model families to each role would reveal how much of the learning effect is real.
  • Inference: the over-deference shown by the smallest models is a testable instance of a general instruction-following risk — adding an explicit prompt that feedback is evidence to weigh, not a command to change one's answer, should reduce the harmful-feedback rate, and this prediction can be checked in the same environment.
  • Inference: the human emotional benchmark was an English counselling corpus while simulations ran in Italian, a mismatch the authors flag; running matched-language human sessions, or measuring acoustic features such as pauses and pitch variance in the native-audio condition, would test whether paralinguistic fidelity is genuinely what drives the competence result.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces MyMentorLLM, a multimodal voice/text simulation environment for CBT deliberate practice, and reports 2,100 simulated mentorship cycles. Each cycle consists of a DSM-5-TR-grounded patient persona (MDD, GAD, or BPD), a trainee therapist persona calibrated to the CTRS performance of real community clinicians, and an expert mentor persona that scores the session on the CTRS and provides feedback. The authors compare seven model–modality conditions (native audio, synthesised audio, text) across three model families. They report three main findings: simulated patients show disorder-congruent emotional profiles that trainee therapists mirror in attenuated form; the native speech-to-speech condition receives CTRS scores close to the human reference while text and synthesised-speech conditions are scored much higher; and mentor feedback improves diagnostic accuracy in five of seven model conditions but harms the two smallest models. The paper is candid about limitations, including shared model biases, reconstructed human reference distributions, and the qualitative nature of the emotional benchmark.

Significance. If the central claims are validated, the paper would provide a large-scale, systematically varied testbed for simulated psychotherapy training, with a useful distinction between conversational fluency and clinical competence. The scale (2,100 sessions), the use of DSM-5-TR clinical cases for patient personas, the attempt to calibrate a trainee persona to empirical CTRS distributions, and the explicit reporting of limitations are strengths. The paper also proposes interpretable emotional-network analyses via EmoAtlas and connects them to a human counselling corpus. However, the main evaluation rests on LLM-generated CTRS scores that are not independently validated, and the pivotal native-speech condition is confounded with a change in scoring model. These issues currently prevent the paper from supporting its strongest claims, though they are addressable with additional experiments.

major comments (4)
  1. [Section 3.2 / Table 2] The central comparison treats mentor-assigned CTRS scores as measures of trainee competence. The mentor LLM is never validated against trained human raters, and in most conditions the same model instantiates patient, trainee, and mentor (Section 2.2). This makes the observed differences compatible with self-consistency, leniency, or prompt-calibration artifacts rather than clinical skill. The Limitations section acknowledges shared biases, but the Results and Abstract do not hedge. I would require an independent validation of LLM CTRS scoring on a sample of transcripts (e.g., two trained human raters) or, at minimum, re-scoring all conditions with a fixed external rater before the competence claim can be evaluated.
  2. [Section 2.2 / Section 3.2] The pivotal native-speech condition is confounded with a change of scorer: the Gemini-3.1 LA patient/trainee dyad is scored by Gemini-3.5 Flash, while every other condition uses the same model as mentor. The 'native speech close to human' result is therefore inseparable from a change in the scoring model. This needs a control: either keep the mentor fixed across all conditions, or re-score all transcripts with the same (ideally human) rater.
  3. [Section 2.4.2 / Table 2] The human CTRS reference is reconstructed from published summaries rather than session-level data, and the comparisons are described as descriptive. Despite this, the paper claims that native speech 'reproduces the competence profile of a real training session.' The observed mean difference (29.97 vs 31.04) is not tested for equivalence or difference, and no confidence intervals are provided. Please add a formal equivalence/non-inferiority analysis, or qualify the claim to state only that the LLM scores fall near the human range.
  4. [Section 3.1 / Section 2.4.1] The emotional-resonance claim is based on a qualitative comparison to the HOPE corpus, which is in English, pooled across counselling contexts, and not stratified by disorder, whereas the simulations were conducted in Italian. The text itself says this comparison is qualitative, but the Results phrase 'parallel those of real psychotherapy' overstates the evidence. Please either obtain a matched-language, disorder-stratified corpus, or soften the claim to reflect the qualitative nature of the benchmark.
minor comments (5)
  1. [Data/code availability] The statements 'Not applicable before publication' are not adequate for reproducibility. The authors should release the persona prompts, anonymised transcripts, and analysis scripts, or provide a clear plan for doing so.
  2. [Table 2] Per-item CTRS scores are reported as means without standard deviations or confidence intervals. Given the large sample sizes, including these would help assess the saturation effects visible in items such as Understanding and Interpersonal effectiveness.
  3. [Eq. (1)] The normalised gain g = (A_F - A_I)/A_F uses A_F in the denominator, which is unconventional and can be unstable when A_F is small. Consider reporting raw differences or Wilson confidence intervals for the accuracy changes.
  4. [Figure 3a] The pooled score distribution across all conditions is dominated by the high-scoring text and synthesised-speech conditions. It would be more informative to show per-condition distributions or overlay them.
  5. [Section 2.2 / Table 1] The label 'E2B' is used without definition in the text; write 'Gemma-4-E2B' consistently at first use and in figure legends.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main CTRS comparison has acknowledged validity limits (same-model scoring, reconstructed human reference), but no claimed result reduces to its inputs by construction.

full rationale

The paper's claims are empirical simulation outputs rather than derivations that become equal to their inputs by construction. In the CTRS analysis, the mentor LLM reads a transcript and returns item scores; this is an observed model output, not a fitted parameter designed to reproduce the human mean. The trainee prompt is admittedly calibrated to Goldberg et al.'s CTRS distributions, so the near-match in the Gemini-3.1 LA condition is partly a fidelity check of prompt adherence rather than an independent prediction, but no equation or identity forces the observed total, and the text/re-synthesised conditions did not match the same target. The paper itself flags the same-model design and the architecture/modality confound as limitations ('In most sessions, the same model instantiated patient, trainee and mentor, allowing shared biases to propagate across the cycle'; 'modality cannot be separated from model architecture and supervisory configuration'), which are validity threats, not circular reductions. Diagnostic and symptom claims are scored against external DSM-5-TR ground truth, and emotional analyses use external lexicons (EmoLex/EmoAtlas) plus the HOPE corpus as a qualitative benchmark, so the central evaluation is substantially anchored outside the simulation. Self-citations appear as background and methodological support, not as the load-bearing proof of any result. Consequently, no circular step meeting the required standard is identified.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The study relies on several domain assumptions about LLM persona stability, cross-lingual emotional analysis validity, the transferability of reconstructed human CTRS baselines, and the representativeness of three fixed clinical cases. No numeric free parameters are fitted: the trainee competence profile is derived from literature medians, and thresholds come from prior publications.

assumptions (5)
  • domain assumption LLM persona prompting yields stable, role-consistent behavior across 31-turn sessions
    The entire study depends on LLMs maintaining patient, trainee, and mentor personas across dialogues; prior work supports this, but it remains an assumption about simulation fidelity. Invoked throughout Section 2.3.
  • domain assumption EmoAtlas/EmoLex word-emotion mapping validly measures emotion in Italian psychotherapy transcripts
    Section 2.4.1 uses EmoAtlas with an English-derived lexicon applied to Italian text, with a pooled English HOPE benchmark; cross-language validity is assumed.
  • domain assumption CTRS distributions from Goldberg et al. (2020) are a valid human baseline for trainee competence
    Section 2.4.2 reconstructs the baseline from published summaries rather than raw session data; item-level medians also drive the trainee persona design, creating a potential circular dependency.
  • ad hoc to paper The same LLM family used as trainee and mentor does not introduce systematic self-consistency bias in CTRS ratings
    The Limitations section admits shared biases could propagate; the CTRS comparison assumes away this contamination, although the paper treats it as a limitation.
  • domain assumption Three DSM-5-TR case vignettes and fixed first sessions represent sufficient clinical heterogeneity
    Only one case per disorder (MDD, GAD, BPD) and fixed 31-turn first encounters are used; generalization to routine care is assumed. Discussed in Limitations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice." pith.science (2026). https://pith.science/paper/DXJ3267B

@misc{pith2026260725667,
  author       = {Pith},
  title        = {Pith review of: MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DXJ3267B}},
  note         = {Machine review of arXiv:2607.25667}
}
read the original abstract

Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a multimodal voice- and text-based simulation environment for deliberate practice, used to generate 2,100 complete Cognitive Behavioural Therapy (CBT) training sessions. Each session links a DSM-5-TR-grounded patient (with major depressive, generalised anxiety or borderline personality disorder), a therapist-in-training and an expert supervisor. As an initial implementation, we adopted CBT because its structured procedures and competency-based supervision facilitate standardised simulation and evaluation. Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy. Simulated patients expressed disorder-congruent emotional profiles, which trainee therapists mirrored as in real human counselling. The quality of supervision differed across LLMs: while most models overestimated trainees' competences, native speech-to-speech was closest to human scores. Supervisors' feedback led to better diagnoses in simulated psychotherapists in 5 out of 7 LLMs, and symptom identification accuracy increased with model size. This work shows that simulation of deliberate practice is possible for CBT training, although patient fidelity, calibration of supervisors, and harmful feedback should be evaluated together.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 1 canonical work pages

  1. [1]

    World Health Organization.World mental health report: Transforming mental health for all (World Health Organization, Geneva, 2022)

  2. [2]

    Rousmaniere, T.Deliberate practice for psychotherapists: A guide to improving clinical effectiveness(Routledge, 2024). 21

  3. [3]

    S., Aghazadeh Ardebili, A

    Franchino, E., Rizzi, R., De Duro, E. S., Aghazadeh Ardebili, A. & Stella, M. Digital shadows in mental health map how llms simulate depression, anxiety, and stress through language and psychometrics.PsyArXiv(2026)

  4. [4]

    Digital mental health: Role of artificial intelligence in psychotherapy.Annals of neurosciences32, 117–127 (2025)

    Bhatt, S. Digital mental health: Role of artificial intelligence in psychotherapy.Annals of neurosciences32, 117–127 (2025)

  5. [5]

    Carrillo, A.et al.Llms can persuade only psychologically susceptible humans on soci- etal issues, via trust in ai and emotional appeals, amid logical fallacies.arXiv preprint arXiv:2604.16935(2026)

  6. [6]

    & Stella, M

    Aghazadeh Ardebili, A. & Stella, M. Mapping how llms debate societal issues when shadow- ing human personality traits, sociodemographics and social media behavior.arXiv preprint arXiv:2604.27624(2026). URL https://arxiv.org/abs/2604.27624

  7. [7]

    Casoria, L., Neroni, P., Sabatucci, L., Augello, A. & Caggianese, G.Evaluating llms for synthetic personas generation: A comparative analysis of personality representation and cen- sorship effects.Proceedings of the 16th Biannual Conference of the Italian SIGCHI Chapter, 1–9 (2025)

  8. [8]

    P.et al.Out of one, many: Using language models to simulate human samples

    Argyle, L. P.et al.Out of one, many: Using language models to simulate human samples. Political Analysis31, 337–351 (2023)

Show all 43 references
  1. [9]

    S., Improta, R

    De Duro, E. S., Improta, R. & Stella, M. Introducing counsellme: A dataset of simulated men- tal health dialogues for comparing llms like haiku, llamantino and chatgpt against humans. Emerging Trends in Drugs, Addictions, and Health5, 100170 (2025)

  2. [10]

    & Collier, N

    Hu, T. & Collier, N. Quantifying the persona effect in llm simulations.arXiv preprint arXiv:2402.10811(2024)

  3. [11]

    & Dickerson, J

    Wang, A., Morgenstern, J. & Dickerson, J. P. Large language models that replace human par- ticipants can harmfully misportray and flatten identity groups.Nature Machine Intelligence 7, 400–411 (2025)

  4. [12]

    URL https://doi.org/10.1057/9781137496850 30

    Muntigl, P.Storytelling, Depression, and Psychotherapy, 577–596 (Palgrave Macmillan UK, London, 2016). URL https://doi.org/10.1057/9781137496850 30

  5. [13]

    M.Speech and text psychometrics: Identifying suicide risk factors with large language models and acoustic networks(Harvard University, 2024)

    Low, D. M.Speech and text psychometrics: Identifying suicide risk factors with large language models and acoustic networks(Harvard University, 2024). 22

  6. [14]

    & Vinciarelli, A

    Tao, F., Esposito, A. & Vinciarelli, A. The androids corpus: A new publicly available benchmark for speech based depression detection.Depression47, 11–9 (2023)

  7. [15]

    arXiv preprint arXiv:2402.01680(2024)

    Guo, T.et al.Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680(2024)

  8. [16]

    Diagnostic and statistical manual of mental disorders: DSM-5Vol

    American Psychiatric Association, D., American Psychiatric Association, D.et al. Diagnostic and statistical manual of mental disorders: DSM-5Vol. 5 (American psychiatric association Washington, DC, 2013)

  9. [17]

    & Gray, K

    Dillion, D., Tandon, N., Gu, Y. & Gray, K. Can ai language models replace human participants?Trends in Cognitive Sciences27, 597–600 (2023)

  10. [18]

    & Stella, M

    Haim, E. & Stella, M. Cognitive networks for knowledge modeling: A gentle introduction for data-and cognitive scientists.Wiley Interdisciplinary Reviews: Cognitive Science17, e70026 (2026)

  11. [19]

    S., Wulff, D

    Siew, C. S., Wulff, D. U., Beckage, N. M. & Kenett, Y. N. Cognitive network science: A review of research on cognition through the lens of network representations, processes, and dynamics.Complexity2019, 2108423 (2019)

  12. [20]

    Stella, M.et al.Cognitive modelling of concepts in the mental lexicon with multilayer networks: Insights, advancements, and future challenges.Psychonomic Bulletin & Review 31, 1981–2004 (2024)

  13. [21]

    Fatima, A., Li, Y., Hills, T. T. & Stella, M. Dasentimental: Detecting depression, anxiety, and stress in texts via emotional recall, cognitive networks, and machine learning.Big data and cognitive computing5, 77 (2021)

  14. [22]

    Semeraro, A.et al.Emoatlas: An emotional network analyzer of texts that merges psycho- logical lexicons, artificial intelligence, and network science.Behavior Research Methods57, 77 (2025)

  15. [23]

    inA general psychoevolutionary theory of emotion(edsTheories of emotion) Plutchik, R

    Plutchik, R. inA general psychoevolutionary theory of emotion(edsTheories of emotion) Plutchik, R. & Kellerman, H. 3–33 (Elsevier, 1980)

  16. [24]

    & Johnstone, T

    Al-Mosaiwi, M. & Johnstone, T. In an absolute state: Elevated use of absolutist words is a marker specific to anxiety, depression, and suicidal ideation.Clinical psychological science 6, 529–542 (2018). 23

  17. [25]

    Liu, D.et al.Detecting and measuring depression on social media using a machine learning approach: systematic review.JMIR Mental Health9, e27244 (2022)

  18. [26]

    W.DSM-5-TR Clinical Cases(American Psychiatric Association Publishing, Washington, DC, 2023)

    Barnhill, J. W.DSM-5-TR Clinical Cases(American Psychiatric Association Publishing, Washington, DC, 2023)

  19. [27]

    edn (American Psychiatric Association Publishing, Washington, DC, 2022)

    American Psychiatric Association.Diagnostic and statistical manual of mental disorders: DSM-5-TR5th, text rev. edn (American Psychiatric Association Publishing, Washington, DC, 2022)

  20. [28]

    Gemini 3.1 flash audio (flash live, tts) — model card (2026)

    Google DeepMind. Gemini 3.1 flash audio (flash live, tts) — model card (2026). URL https://storage.googleapis.com/deepmind-media/Model-Cards/ Gemini-3-1-Flash-Audio-Model-Card.pdf. Published March 2026, updated April 2026. Accessed 2026-07-27

  21. [29]

    Gemma-4-12b-it — model card (2026)

    Gemma Team, Google DeepMind. Gemma-4-12b-it — model card (2026). URL https: //huggingface.co/google/gemma-4-12B-it

  22. [30]

    Zhu, H.et al.Omnivoice: Towards omnilingual zero-shot text-to-speech with diffusion language models.arXiv preprint arXiv:2604.00688(2026)

  23. [31]

    Qwen3.6-35b-a3b-awq — model card (2026)

    Qwen Team, Alibaba. Qwen3.6-35b-a3b-awq — model card (2026). URL https:// huggingface.co/QuantTrio/Qwen3.6-35B-A3B-A WQ

  24. [32]

    Gemini 3.5 flash — model card (2026)

    Google DeepMind. Gemini 3.5 flash — model card (2026). URL https://storage.googleapis. com/deepmind-media/Model-Cards/Gemini-3-5-Flash-Model-Card.pdf. Published May

  25. [33]

    Gemma-4-e2b-it — model card (2026)

    Gemma Team, Google DeepMind. Gemma-4-e2b-it — model card (2026). URL https: //huggingface.co/google/gemma-4-E2B-it

  26. [34]

    Qwen3.5: Towards native multimodal agents (2026)

    Qwen Team, Alibaba. Qwen3.5: Towards native multimodal agents (2026). URL https: //huggingface.co/Qwen/Qwen3.5-9B

  27. [35]

    B.et al.The structure of competence: Evaluating the factor structure of the cognitive therapy rating scale.Behavior Therapy51, 113–122 (2020)

    Goldberg, S. B.et al.The structure of competence: Evaluating the factor structure of the cognitive therapy rating scale.Behavior Therapy51, 113–122 (2020)

  28. [36]

    A., Wolk, C

    Creed, T. A., Wolk, C. B., Feinberg, B., Evans, A. C. & Beck, A. T. Beyond the label: Relationship between community therapists’ self-report of a cognitive behavioral therapy orientation and observed skills.Administration and Policy in Mental Health and Mental 24 Health Servic...

  29. [37]

    Malhotra, G., Waheed, A., Srivastava, A., Akhtar, M. S. & Chakraborty, T.Speaker and time- aware joint contextual learning for dialogue-act classification in counselling conversations. Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining (WSD...

  30. [38]

    Mohammad, S. M. & Turney, P. D. Crowdsourcing a word–emotion association lexicon. Computational intelligence29, 436–465 (2013)

  31. [39]

    F.et al.Therapist competence ratings in relation to clinical outcome in cognitive therapy of depression.Journal of Consulting and Clinical Psychology67, 837–846 (1999)

    Shaw, B. F.et al.Therapist competence ratings in relation to clinical outcome in cognitive therapy of depression.Journal of Consulting and Clinical Psychology67, 837–846 (1999)

  32. [40]

    & Bowman, S

    Turpin, M., Michael, J., Perez, E. & Bowman, S. Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.Advances in Neural Information Processing Systems36, 74952–74965 (2023)

  33. [41]

    arXiv preprint arXiv:2404.13208(2024)

    Wallace, E.et al.The instruction hierarchy: Training llms to prioritize privileged instructions. arXiv preprint arXiv:2404.13208(2024)

  34. [42]

    Harrigian, K., Aguirre, C. A. & Dredze, M.Do models of mental health based on social media data generalize? Findings of the association for computational linguistics: EMNLP 2020, 3774–3788 (2020)

  35. [43]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature machine intelligence1, 206–215 (2019)

    Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature machine intelligence1, 206–215 (2019). 25 Supplementary Tables T able S1:Emotion z-scores by condition and disorder.EmoAtlas z-scores for each of t...

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.