REVIEW 4 cited by
The open-source VERA-MH benchmark's automated LLM judge agrees with clinician consensus on chatbot safety in suicide-risk conversations (chance-corrected reliability 0.81), supporting its use as a clinically valid evaluation tool.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 04:19 UTC pith:JMCQJUWD
load-bearing objection First human validation of VERA-MH, but the gold-standard clinicians are trained and selected by the same team, so the headline IRR proves rubric-application consistency more than independent clinical validity.
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that VERA-MH's automated LLM judge produces safety ratings that closely match a gold-standard clinical consensus. In 90 simulated conversations spanning four suicide risk levels and varied disclosure styles, clinicians agreed with one another on the five rubric dimensions with a chance-corrected inter-rater reliability of 0.77, and the LLM judge agreed with the clinician consensus at 0.81. Agreement was stable across several judge LLMs and across user risk levels, and clinicians generally found the simulated users realistic. The authors take this as evidence that VERA-MH—an open-source, fully automated evaluation—can serve as a clinically valid benchmark for detecting un
What carries the argument
The load-bearing machinery is the VERA-MH rubric plus the LLM-as-a-judge setup. The rubric defines five safety dimensions—Detects Potential Risk, Confirms Risk, Guides to Human Care, Supportive Conversation, and Follows AI Boundaries—each rated as Best Practice, Suboptimal but Low Potential for Harm, High Potential for Harm, or Not Relevant, with the most severe indicator present determining the rating. Both the human clinicians and the LLM judge apply the very same item text and instructions, which standardizes the comparison but also means the study measures the reliability of applying the rubric, not the rubric's independent validity. The judge is a general-purpose LLM prompted with the c
Load-bearing premise
The study's gold standard is not independent of the benchmark: the clinicians applied the exact VERA-MH rubric, were trained on it, and the most rubric-concordant counselors were selected for the rating phase, so the high agreement could reflect consistent application of the rubric rather than valid measurement of chatbot safety.
What would settle it
Ask a new panel of licensed clinicians—who have never seen the VERA-MH rubric—to rate the same 90 conversations for chatbot safety using their own professional judgment, then compare those unconstrained ratings to the LLM judge's VERA-MH ratings. If the agreement drops substantially (e.g., below 0.6 chance-corrected), the benchmark would be shown to measure rubric-consistency rather than clinical safety.
If this is right
- If VERA-MH is adopted, developers could run automated safety audits on chatbots before release, without needing a panel of human clinicians each time.
- Because each unsafe rating maps to specific rubric items, the benchmark gives developers actionable feedback on exactly which behavior was unsafe (e.g., failure to confirm risk or failure to guide to human care).
- The finding that the judge generalizes across several LLMs suggests the evaluation can be re-run as chatbot models evolve and new ones appear.
- The benchmark could be expanded beyond suicide risk to other mental health safety domains, such as psychosis or harm from others, which the authors identify as a future direction.
- A reliable automated benchmark could inform regulators, legislators, and the public about whether a particular AI chatbot is safe enough for mental health support.
Where Pith is reading between the lines
- The study compares the LLM judge to clinicians who were trained on—and selected for agreement with—the same rubric the judge uses. This makes the reported reliability a measure of consistent rubric application; it does not yet establish that the rubric captures clinician judgments of safety that are independent of the rubric. An independent validation with clinicians' unrestricted safety judgments
- Only 36.5% of simulated conversations actually matched their prompted disclosure level, and clinicians rated communication-style realism at a median of 3 on a 5-point scale. This suggests the simulation pipeline may not reliably produce the range of indirect and ambiguous disclosure styles the benchmark intends to cover, which is worth testing explicitly.
- Widespread adoption of such a benchmark could create a measurement loop: chatbot developers may optimize for the five rubric dimensions while other safety-relevant behaviors not captured by the rubric go unmeasured. Complementing VERA-MH with free-form clinician audits or outcome-based criteria (e.g., whether users later seek help) would guard against that.
- Since the paper reports lower LLM-clinician agreement when the provider chatbot was a specific Gemini model (IRR 0.71), a testable extension is to map VERA-MH reliability across a wider, more adversarial set of chatbots to see where the benchmark starts to fail.
Editorial analysis
A structured set of objections, weighed in public.
Circularity Check
The 'gold-standard' clinician consensus is generated by clinicians trained on and selected for concordance with the VERA-MH rubric; the LLM judge uses identical item text, so the 0.81 agreement supports rubric-application consistency, not independent clinical validity.
specific steps
-
self definitional
[Methods, Measures/Procedure and LLM judge evaluation (also Abstract)]
"Clinicians then used the VERA-MH rubric to independently rate each conversation for safety, and we used their ratings to assess how consistent individual clinicians were with one another. Next, drawing from the LLM-as-a-judge framework, an LLM judge used the same VERA-MH rubric to rate the same conversations... The system prompt ... consisted of the conversation to evaluate and the same item text and instructions used in the clinician rating form."
The clinical consensus called a 'gold-standard reference' is produced by raters applying the VERA-MH rubric, and the LLM judge is given the same item text and instructions. The 0.81 LLM-consensus IRR is therefore an inter-rater reliability statistic between two applications of the same instrument. It can show that GPT-4o applies the VERA-MH rubric consistently with calibrated clinicians, but it cannot validate the five rubric dimensions as measures of safety, because no external clinical criterion is used. The conclusion that these results 'support the clinical validity and reliability of VERA-MH' assumes the rubric is valid; the design only supports rubric-application consistency.
-
fitted input called prediction
[Supplemental Method, Clinician rater training and calibration]
"Six counselors/therapists began the training and calibration phase; the four evidencing the strongest concordance with the reference standard set during calibration moved to the independent rating phase with the two psychologists (six total raters)."
Clinicians were selected for strongest concordance with a reference standard before producing the ratings used to define the gold-standard consensus. This selection makes the published clinician-clinician IRR (0.77) and the consensus reference partly a function of the rubric/reference rather than an independent clinical judgment. The 'gold standard' is fitted to the instrument being validated and then used as the criterion for VERA-MH, so the validation loop is partially closed. No external safety criterion is reported; the reference standard itself is a doctoral psychologist applying the same rubric.
full rationale
The paper reports real inter-rater statistics: clinician-clinician α = 0.77 and LLM-consensus α = 0.81, with overlapping CIs across judge LLMs. These computations are not circular by themselves. The circularity is at the level of the criterion. Clinicians are Spring Health employees, trained in three rounds of practice coding on the VERA-MH rubric, and the four counselors who proceeded were selected for strongest concordance with a reference standard. Both humans and the LLM judge then applied the same 30-item VERA-MH form, with the LLM prompt using 'the same item text and instructions used in the clinician rating form.' The study therefore establishes that the LLM can apply the VERA-MH rubric similarly to calibrated clinicians, but it does not provide an external gold standard for chatbot safety. The abstract's inference that these results 'support the clinical validity and reliability of VERA-MH' conflates rubric-application reliability with construct validity. This is a partial, design-level circularity, not a fully forced mathematical equivalence; the agreement values could have been lower, and the stability and realism analyses add independent information. The paper's self-citations to the earlier VERA-MH concept paper are not load-bearing for the validity inference, so the score is 6 rather than higher. The limitation section asks for future validation of updated versions and generalizability, but it does not acknowledge that the current gold standard was generated by the same rubric, which is the central issue.
Axiom & Free-Parameter Ledger
axioms (5)
- ad hoc to paper The VERA-MH rubric is a valid operationalization of chatbot safety in suicide-risk conversations.
- domain assumption Clinician consensus, obtained by applying the VERA-MH rubric, is the gold-standard reference for chatbot safety.
- domain assumption The simulated LLM user-agent conversations are representative of real-world mental-health interactions.
- standard math Krippendorff's alpha is an appropriate chance-corrected agreement measure for these categorical ratings.
- ad hoc to paper The four counselors with strongest calibration concordance are representative raters for estimating clinician-clinician reliability.
read the original abstract
Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are safe. The field currently lacks a validated, automated benchmark for evaluating AI chatbot safety, particularly for users at risk of suicide. The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) evaluation was recently proposed to address this need. This human validation study examined the alignment of VERA-MH safety ratings with expert clinician judgments. We simulated conversations between large language model (LLM)-based users spanning a range of suicide risk levels and disclosure styles and general-purpose AI chatbots. Licensed mental health clinicians from Spring Health independently rated chatbot safety using the VERA-MH scoring rubric. An LLM-based evaluator ("judge") applied the same rubric to the same conversations. We examined agreement among clinicians, between clinician consensus and the LLM judge, and across different judge LLMs. Clinicians also rated user-agent realism, suicide risk, and disclosure. Clinicians showed strong agreement in safety ratings (chance-corrected inter-rater reliability [IRR] = 0.77), establishing a reliable clinical consensus reference. The LLM judge was strongly aligned with this consensus (IRR = 0.81), and ratings were stable across judge models and repeated evaluations. Ratings of user-agent realism and fidelity to intended suicide risk and disclosure styles were mixed. These findings support the reliability of VERA-MH as an open-source, fully automated benchmark for evaluating AI chatbot suicide risk detection and response. Because these results reflect an earlier version of the benchmark, future work should validate updated versions, assess generalizability and robustness, and expand VERA-MH to additional domains of AI safety in mental health.
Forward citations
Cited by 4 Pith papers
-
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
LLMs detect user distress equally with or without delusional framing but suppress safety interventions up to 4.5x more when distress is embedded in delusions.
-
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment whe...
-
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Using simulated psychiatric user profiles, the authors show that AI chatbots frequently produce 'concerning behavior' that accumulates over turns, and that superficially supportive responses can amplify vulnerability—...
-
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
A three-factor framework and eighteen aspirational behavioral directions link specific chatbot patterns to user risk factors and potential psychological harms across everyday, role-play, and support uses.
Reference graph
Works this paper leans on
-
[2]
Use of generative AI for mental health advice among US adolescents and young adults
McBain RK, Bozick R, Diliberti M, et al. Use of generative AI for mental health advice among US adolescents and young adults. JAMA Netw Open. 2025;8(11):e2542281. doi:10.1001/jamanetworkopen.2025.42281
arXiv 2025
-
[3]
Strengthening ChatGPT’s responses in sensitive conversations
OpenAI. Strengthening ChatGPT’s responses in sensitive conversations. October 27, 2025. Accessed February 3, 2026. https://openai.com/index/strengthening-chatgpt-responses-in- sensitive-conversations/
2025
-
[5]
For argument’s sake, show me how to harm myself: Jailbreaking LLMs in suicide and self-harm contexts
Schoene AM, Canca C. For argument’s sake, show me how to harm myself: Jailbreaking LLMs in suicide and self-harm contexts. In: Proceedings of the IEEE International Symposium on Technology and Society (ISTAS). 2025. doi:10.1109/ISTAS65609.2025.11269647
arXiv 2025
-
[6]
Iftikhar Z, Xiao A, Ransom S, Huang J, Suresh H. How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 2025;8(2):1311-1323. doi:10.1609/aies.v8i2.36632
-
[7]
Towards understanding sycophancy in language models
Sharma M, Tong M, Korbak T, et al. Towards understanding sycophancy in language models. arXiv [preprint]. Published October 20, 2023. arXiv:2310.13548. doi:10.48550/arXiv.2310.13548
-
[8]
Longitudinal study on social and emotional use of AI conversational agent
Chandra M, Hernandez J, Ramos G, Ershadi M, Bhattacharjee A, Amores J, et al. Longitudinal study on social and emotional use of AI conversational agent. arXiv [preprint]. 2025. doi:10.48550/arXiv.2504.14112
-
[10]
Moore J, Grabb D, Agnew W, Klyman K, Chancellor S, Ong DS, Haber N. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In: VERA-MH 16 Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2025:599-627. doi:10.1145/3715275.3732039
arXiv 2025
-
[11]
Incident 826: Character.ai chatbot allegedly influenced teen user toward suicide amid claims of missing guardrails
Atherton D. Incident 826: Character.ai chatbot allegedly influenced teen user toward suicide amid claims of missing guardrails. AI Incident Database. 2024. Accessed September 2025. https://incidentdatabase.ai/cite/826/
2024
-
[12]
Incident 1192: 16-year-old allegedly received suicide method guidance from ChatGPT before death
Atherton D. Incident 1192: 16-year-old allegedly received suicide method guidance from ChatGPT before death. AI Incident Database. 2025. Accessed September 2025. https://incidentdatabase.ai/cite/1192/
2025
-
[13]
Lawsuits blame ChatGPT for suicides and harmful delusions
Hill K. Lawsuits blame ChatGPT for suicides and harmful delusions. New York Times. November 6, 2025. Accessed February 3, 2026. https://www.nytimes.com/2025/11/06/technology/chatgpt-lawsuit-suicides-delusions.html
2025
-
[14]
VERA-MH concept paper: Validation of ethical and responsible AI in mental health
Belli L, Bentley K, Alexander W, Ward E, Hawrilenko M, Johnston K, et al. VERA-MH concept paper: Validation of ethical and responsible AI in mental health. arXiv [preprint]. October 17,
-
[15]
LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods
Li H, Dong Q, Chen J, Su H, Zhou Y, Ai Q, et al. LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods. arXiv [preprint]. 2024. doi:10.48550/arXiv.2412.05579
-
[17]
Open letter to the AI and technology industry: Protecting youth mental health and preventing suicide in the age of AI
The Jed Foundation. Open letter to the AI and technology industry: Protecting youth mental health and preventing suicide in the age of AI. September 17, 2025. Accessed February 3, 2026. https://jedfoundation.org/open-letter-to-the-ai-and-technology-industry/
2025
-
[19]
Managing suicidal risk: A collaborative approach
Jobes DA. Managing suicidal risk: A collaborative approach. 2nd ed. Guilford Press; 2016
2016
-
[20]
SAFE-T suicide assessment five- step evaluation and triage (PEP24-01-036)
Substance Abuse and Mental Health Services Administration. SAFE-T suicide assessment five- step evaluation and triage (PEP24-01-036). Published 2024. Accessed October 16, 2025. https://library.samhsa.gov/product/safe-t-suicide-assessment-five-step-evaluation-and- triage/pep24-01-036
2024
-
[21]
About QPR
QPR Institute. About QPR. Accessed February 3, 2026. https://qprinstitute.com/about-qpr/ VERA-MH 17
2026
-
[22]
Zero Suicide framework
Zero Suicide. Zero Suicide framework. Accessed February 3, 2026. https://zerosuicide.edc.org/zero-suicide-framework
2026
-
[24]
Can large language models identify implicit suicidal ideation? An empirical evaluation
Li T, Yang S, Wu J, Wei J, Hu L, Li M, et al. Can large language models identify implicit suicidal ideation? An empirical evaluation. arXiv [preprint]. 2025. arXiv:2502.17899v2. https://arxiv.org/abs/2502.17899v2
arXiv 2025
-
[26]
Content analysis in mass communication: Assessment and reporting of intercoder reliability
Lombard M, Snyder-Duch J, Bracken CC. Content analysis in mass communication: Assessment and reporting of intercoder reliability. Hum Commun Res. 2002;28(4):587-604. doi:10.1093/hcr/28.4.587
-
[27]
MindEval: Benchmarking language models on multi-turn mental health support
Pombal J, D’Eon M, Guerreiro NM, Martins PH, Farinhas A, Rei R. MindEval: Benchmarking language models on multi-turn mental health support. arXiv [preprint]. 2025. arXiv:2511.18491. doi:10.48550/arXiv.2511.18491
-
[28]
Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation
Eriksson M, Purificato E, Noroozian A, Vinagre J, Chaslot G, Gómez E, Fernández-Llorca D. Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation. arXiv [preprint]. 2025. arXiv:2502.06559
Pith/arXiv arXiv 2025
-
[29]
AI and the everything in the whole wide world benchmark
Raji D, Denton E, Bender EM, Hanna A, Paullada A. AI and the everything in the whole wide world benchmark. In: Vanschoren J, Yeung S, eds. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1. NeurIPS; 2021. Accessed February 3, 2026. https://datasets-benchmarks- proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed...
2021
-
[30]
Heterogeneity in suicide risk: Evidence from personalized dynamic models
Coppersmith DDL, Kleiman EM, Millner AJ, Wang SB, Arizmendi C, Bentley KH, et al. Heterogeneity in suicide risk: Evidence from personalized dynamic models. Behav Res Ther. 2024;180:104574. doi:10.1016/j.brat.2024.104574
arXiv 2024
-
[31]
Disclosure of suicidal ideation and behaviours: A systematic review and meta-analysis of prevalence
Hallford DJ, Rusanov D, Winestone B, Kaplan R, Fuller-Tyszkiewicz M, Melvin G. Disclosure of suicidal ideation and behaviours: A systematic review and meta-analysis of prevalence. Clin Psychol Rev. 2023;101:102272. doi:10.1016/j.cpr.2023.102272 VERA-MH 18
arXiv 2023
-
[32]
Richards JE, Whiteside U, Ludman EJ, Pabiniak C, Kirlin B, Hidalgo R, Simon G. Understanding why patients may not report suicidal ideation at a health care visit prior to a suicide attempt: A qualitative study. Psychiatr Serv. 2019;70(1):40-45. doi:10.1176/appi.ps.201800342
-
[33]
LLMs get lost in multi-turn conversation
Laban P, Hayashi H, Zhou Y, Neville J. LLMs get lost in multi-turn conversation. arXiv [preprint]. 2025. arXiv:2505.06120. doi:10.48550/arXiv.2505.06120
-
[34]
Helping people when they need it most
OpenAI. Helping people when they need it most. August 26, 2025. Accessed February 3, 2026. https://openai.com/index/helping-people-when-they-need-it-most/
2025
-
[35]
Keshavan M, Torous J, Yassin W. Do generative AI chatbots increase psychosis risk? World Psychiatry. 2026;25(1):150-151. doi:10.1002/wps.70017 VERA-MH 19 Supplemental Tables Supplemental Table S1 Overview of VERA-MH User-Agent Profiles Demographics Mental Health Background and Recent Stressors Suicidal Thoughts and Behaviors Communication Style and Respon...
-
[37]
Winslow B, Shreibati J, Perez J, Su HW, Young-Lin N, Hammerquist N, et al. A principle- based framework for the development and evaluation of large language models for health and wellness. arXiv [preprint]. 2025. arXiv:2512.08936. doi:10.48550/arXiv.2512.08936
-
[38]
Health advisory: Use of generative AI chatbots and wellness applications for mental health
American Psychological Association. Health advisory: Use of generative AI chatbots and wellness applications for mental health. Published 2025. Accessed February 3, 2026. https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory- chatbots-wellness-apps/
2025
-
[39]
Independent clinical evaluation of general-purpose LLM responses to signals of suicide risk
Judd N, Vaz A, Paeth K, Davis LI, Esherick M, Brand J, et al. Independent clinical evaluation of general-purpose LLM responses to signals of suicide risk. arXiv [preprint]. 2025. https://arxiv.org/pdf/2510.27521
arXiv 2025
-
[40]
Suicidal ideation detection in conversational AI: A compliance and implementation guide
ThroughLine. Suicidal ideation detection in conversational AI: A compliance and implementation guide. Published 2026. Accessed February 3, 2026. https://cdn.prod.website- files.com/693a4901d87d2a9fd3bf0d09/69448c563c28a5ee29113ca5_Throughline%20- %20Whitepaper.pdf
2026
-
[41]
Dazzi T, Gribble R, Wessely S, Fear NT. Does asking about suicide and related behaviours induce suicidal ideation? What is the evidence? Psychol Med. 2014;44(16):3361-3363. doi:10.1017/S0033291714001299
-
[42]
About QPR
QPR Institute. About QPR. Accessed February 3, 2026. https://qprinstitute.com/about-qpr/
2026
-
[43]
Making chatbots safe for suicidal patients
Frances A, Whiteside U. Making chatbots safe for suicidal patients. Psychiatric Times. November 18, 2025. Accessed February 3, 2026. https://www.psychiatrictimes.com/view/making-chatbots-safe-for-suicidal-patients
2025
-
[44]
Food and Drug Administration
U.S. Food and Drug Administration. Executive summary for the Digital Health Advisory Committee meeting: Generative artificial intelligence-enabled digital mental health medical devices. November 6, 2025. Accessed February 3, 2026. https://www.fda.gov/media/189391/download
2025
-
[45]
Open letter on AI safety
Now Matters Now. Open letter on AI safety. Published 2025. Accessed February 3, 2026. https://nowmattersnow.org/open-letter-on-ai-safety/
2025
-
[46]
The case against no-suicide contracts: The commitment to treatment statement as a practice alternative
Rudd MD, Mandrusiak M, Joiner TE. The case against no-suicide contracts: The commitment to treatment statement as a practice alternative. J Clin Psychol. 2006;62(2):243-
2006
-
[48]
McBain RK, Cantor JH, Zhang LA, Baker O, Zhang F, Burnett A, et al. Evaluation of alignment between large language models and expert clinicians in suicide risk assessment. Psychiatr Serv. Published online 2025. doi:10.1176/appi.ps.20250086
-
[49]
Content analysis: An introduction to its methodology
Krippendorff K. Content analysis: An introduction to its methodology. 4th ed. Sage Publications; 2018
2018
-
[251]
doi:10.1002/jclp.20227 VERA-MH 33
-
[2025]
arXiv:2510.15297. doi:10.48550/arXiv.2510.15297
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2510.15297
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.