Pith. sign in

REVIEW 4 cited by

The open-source VERA-MH benchmark's automated LLM judge agrees with clinician consensus on chatbot safety in suicide-risk conversations (chance-corrected reliability 0.81), supporting its use as a clinically valid evaluation tool.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 04:19 UTC pith:JMCQJUWD

load-bearing objection First human validation of VERA-MH, but the gold-standard clinicians are trained and selected by the same team, so the headline IRR proves rubric-application consistency more than independent clinical validity.

arxiv 2602.05088 v4 pith:JMCQJUWD submitted 2026-02-04 cs.AI

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

classification cs.AI
keywords suicide riskAI chatbot safetyLLM-as-a-judgeclinical validationinter-rater reliabilitymental health AIVERA-MHbenchmark evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Generative AI chatbots are now widely used for emotional support, and a growing number of users disclose suicidal thoughts to them, yet the field lacks an established way to automatically decide whether a chatbot responded safely. This paper tests the clinical validity of an open-source benchmark called VERA-MH, which scores chatbot conversations on five safety dimensions: risk detection, risk confirmation, guidance to human care, supportive conversation, and AI boundaries. Six licensed clinicians independently rated 90 simulated user-chatbot conversations using the VERA-MH rubric, and an LLM 'judge' rated the same conversations with the same rubric. The paper reports strong clinician-clinician agreement (inter-rater reliability 0.77) and strong LLM-clinician agreement (0.81), concluding that the automated evaluation is a clinically valid and reliable way to benchmark chatbot safety in suicide-risk situations.

Core claim

The central claim is that VERA-MH's automated LLM judge produces safety ratings that closely match a gold-standard clinical consensus. In 90 simulated conversations spanning four suicide risk levels and varied disclosure styles, clinicians agreed with one another on the five rubric dimensions with a chance-corrected inter-rater reliability of 0.77, and the LLM judge agreed with the clinician consensus at 0.81. Agreement was stable across several judge LLMs and across user risk levels, and clinicians generally found the simulated users realistic. The authors take this as evidence that VERA-MH—an open-source, fully automated evaluation—can serve as a clinically valid benchmark for detecting un

What carries the argument

The load-bearing machinery is the VERA-MH rubric plus the LLM-as-a-judge setup. The rubric defines five safety dimensions—Detects Potential Risk, Confirms Risk, Guides to Human Care, Supportive Conversation, and Follows AI Boundaries—each rated as Best Practice, Suboptimal but Low Potential for Harm, High Potential for Harm, or Not Relevant, with the most severe indicator present determining the rating. Both the human clinicians and the LLM judge apply the very same item text and instructions, which standardizes the comparison but also means the study measures the reliability of applying the rubric, not the rubric's independent validity. The judge is a general-purpose LLM prompted with the c

Load-bearing premise

The study's gold standard is not independent of the benchmark: the clinicians applied the exact VERA-MH rubric, were trained on it, and the most rubric-concordant counselors were selected for the rating phase, so the high agreement could reflect consistent application of the rubric rather than valid measurement of chatbot safety.

What would settle it

Ask a new panel of licensed clinicians—who have never seen the VERA-MH rubric—to rate the same 90 conversations for chatbot safety using their own professional judgment, then compare those unconstrained ratings to the LLM judge's VERA-MH ratings. If the agreement drops substantially (e.g., below 0.6 chance-corrected), the benchmark would be shown to measure rubric-consistency rather than clinical safety.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If VERA-MH is adopted, developers could run automated safety audits on chatbots before release, without needing a panel of human clinicians each time.
  • Because each unsafe rating maps to specific rubric items, the benchmark gives developers actionable feedback on exactly which behavior was unsafe (e.g., failure to confirm risk or failure to guide to human care).
  • The finding that the judge generalizes across several LLMs suggests the evaluation can be re-run as chatbot models evolve and new ones appear.
  • The benchmark could be expanded beyond suicide risk to other mental health safety domains, such as psychosis or harm from others, which the authors identify as a future direction.
  • A reliable automated benchmark could inform regulators, legislators, and the public about whether a particular AI chatbot is safe enough for mental health support.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The study compares the LLM judge to clinicians who were trained on—and selected for agreement with—the same rubric the judge uses. This makes the reported reliability a measure of consistent rubric application; it does not yet establish that the rubric captures clinician judgments of safety that are independent of the rubric. An independent validation with clinicians' unrestricted safety judgments
  • Only 36.5% of simulated conversations actually matched their prompted disclosure level, and clinicians rated communication-style realism at a median of 3 on a 5-point scale. This suggests the simulation pipeline may not reliably produce the range of indirect and ambiguous disclosure styles the benchmark intends to cover, which is worth testing explicitly.
  • Widespread adoption of such a benchmark could create a measurement loop: chatbot developers may optimize for the five rubric dimensions while other safety-relevant behaviors not captured by the rubric go unmeasured. Complementing VERA-MH with free-form clinician audits or outcome-based criteria (e.g., whether users later seek help) would guard against that.
  • Since the paper reports lower LLM-clinician agreement when the provider chatbot was a specific Gemini model (IRR 0.71), a testable extension is to map VERA-MH reliability across a wider, more adversarial set of chatbots to see where the benchmark starts to fail.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Circularity Check

2 steps flagged

The 'gold-standard' clinician consensus is generated by clinicians trained on and selected for concordance with the VERA-MH rubric; the LLM judge uses identical item text, so the 0.81 agreement supports rubric-application consistency, not independent clinical validity.

specific steps
  1. self definitional [Methods, Measures/Procedure and LLM judge evaluation (also Abstract)]
    "Clinicians then used the VERA-MH rubric to independently rate each conversation for safety, and we used their ratings to assess how consistent individual clinicians were with one another. Next, drawing from the LLM-as-a-judge framework, an LLM judge used the same VERA-MH rubric to rate the same conversations... The system prompt ... consisted of the conversation to evaluate and the same item text and instructions used in the clinician rating form."

    The clinical consensus called a 'gold-standard reference' is produced by raters applying the VERA-MH rubric, and the LLM judge is given the same item text and instructions. The 0.81 LLM-consensus IRR is therefore an inter-rater reliability statistic between two applications of the same instrument. It can show that GPT-4o applies the VERA-MH rubric consistently with calibrated clinicians, but it cannot validate the five rubric dimensions as measures of safety, because no external clinical criterion is used. The conclusion that these results 'support the clinical validity and reliability of VERA-MH' assumes the rubric is valid; the design only supports rubric-application consistency.

  2. fitted input called prediction [Supplemental Method, Clinician rater training and calibration]
    "Six counselors/therapists began the training and calibration phase; the four evidencing the strongest concordance with the reference standard set during calibration moved to the independent rating phase with the two psychologists (six total raters)."

    Clinicians were selected for strongest concordance with a reference standard before producing the ratings used to define the gold-standard consensus. This selection makes the published clinician-clinician IRR (0.77) and the consensus reference partly a function of the rubric/reference rather than an independent clinical judgment. The 'gold standard' is fitted to the instrument being validated and then used as the criterion for VERA-MH, so the validation loop is partially closed. No external safety criterion is reported; the reference standard itself is a doctoral psychologist applying the same rubric.

full rationale

The paper reports real inter-rater statistics: clinician-clinician α = 0.77 and LLM-consensus α = 0.81, with overlapping CIs across judge LLMs. These computations are not circular by themselves. The circularity is at the level of the criterion. Clinicians are Spring Health employees, trained in three rounds of practice coding on the VERA-MH rubric, and the four counselors who proceeded were selected for strongest concordance with a reference standard. Both humans and the LLM judge then applied the same 30-item VERA-MH form, with the LLM prompt using 'the same item text and instructions used in the clinician rating form.' The study therefore establishes that the LLM can apply the VERA-MH rubric similarly to calibrated clinicians, but it does not provide an external gold standard for chatbot safety. The abstract's inference that these results 'support the clinical validity and reliability of VERA-MH' conflates rubric-application reliability with construct validity. This is a partial, design-level circularity, not a fully forced mathematical equivalence; the agreement values could have been lower, and the stability and realism analyses add independent information. The paper's self-citations to the earlier VERA-MH concept paper are not load-bearing for the validity inference, so the score is 6 rather than higher. The limitation section asks for future validation of updated versions and generalizability, but it does not acknowledge that the current gold standard was generated by the same rubric, which is the central issue.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The central claim rests on the VERA-MH rubric as the operational definition of safety, on the assumption that clinician consensus is an independent gold standard, and on the representativeness of simulated conversations. The first two are weakened by organizational and training overlap; the third is only partially supported by realism ratings. There are no fitted free parameters in the statistical model, but the rubric itself, its severity ordering, and the 'most severe indicator' scoring rule are hand-made modeling choices.

axioms (5)
  • ad hoc to paper The VERA-MH rubric is a valid operationalization of chatbot safety in suicide-risk conversations.
    All safety ratings, from both clinicians and the LLM judge, are based on this rubric. The rubric was created by the authors; no external criterion is used to validate the rubric itself.
  • domain assumption Clinician consensus, obtained by applying the VERA-MH rubric, is the gold-standard reference for chatbot safety.
    The paper defines validity as alignment with clinicians, but the clinicians are Spring Health employees and were trained on the VERA-MH rubric, weakening the independence of this gold standard.
  • domain assumption The simulated LLM user-agent conversations are representative of real-world mental-health interactions.
    The study tests realism and finds mixed results (presentation median 4/5, communication median 3/5). The intended disclosure styles were not reliably conveyed, partially undermining this assumption.
  • standard math Krippendorff's alpha is an appropriate chance-corrected agreement measure for these categorical ratings.
    A standard measure in content analysis; appropriate for multiple raters and nominal categories.
  • ad hoc to paper The four counselors with strongest calibration concordance are representative raters for estimating clinician-clinician reliability.
    Six counselors began training and the four with highest agreement with the reference standard moved to independent rating, potentially biasing the clinician IRR upward.

pith-pipeline@v1.3.0-alltime-deepseek · 17010 in / 10319 out tokens · 106961 ms · 2026-08-03T04:19:59.929993+00:00 · methodology

0 comments
read the original abstract

Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental health is whether these tools are safe. The field currently lacks a validated, automated benchmark for evaluating AI chatbot safety, particularly for users at risk of suicide. The Validation of Ethical and Responsible AI in Mental Health (VERA-MH) evaluation was recently proposed to address this need. This human validation study examined the alignment of VERA-MH safety ratings with expert clinician judgments. We simulated conversations between large language model (LLM)-based users spanning a range of suicide risk levels and disclosure styles and general-purpose AI chatbots. Licensed mental health clinicians from Spring Health independently rated chatbot safety using the VERA-MH scoring rubric. An LLM-based evaluator ("judge") applied the same rubric to the same conversations. We examined agreement among clinicians, between clinician consensus and the LLM judge, and across different judge LLMs. Clinicians also rated user-agent realism, suicide risk, and disclosure. Clinicians showed strong agreement in safety ratings (chance-corrected inter-rater reliability [IRR] = 0.77), establishing a reliable clinical consensus reference. The LLM judge was strongly aligned with this consensus (IRR = 0.81), and ratings were stable across judge models and repeated evaluations. Ratings of user-agent realism and fidelity to intended suicide risk and disclosure styles were mixed. These findings support the reliability of VERA-MH as an open-source, fully automated benchmark for evaluating AI chatbot suicide risk detection and response. Because these results reflect an earlier version of the benchmark, future work should validate updated versions, assess generalizability and robustness, and expand VERA-MH to additional domains of AI safety in mental health.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

    cs.CL 2026-05 unverdicted novelty 6.0

    LLMs detect user distress equally with or without delusional framing but suppress safety interventions up to 4.5x more when distress is embedded in delusions.

  2. LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment

    cs.CY 2026-05 unverdicted novelty 6.0

    Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment whe...

  3. A clinically validated framework for auditing AI chatbot behavior in mental health interactions

    q-bio.NC 2026-02 conditional novelty 6.0

    Using simulated psychiatric user profiles, the authors show that AI chatbots frequently produce 'concerning behavior' that accumulates over turns, and that superficially supportive responses can amplify vulnerability—...

  4. Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being

    cs.AI 2026-07 conditional novelty 5.5

    A three-factor framework and eighteen aspirational behavioral directions link specific chatbot patterns to user risk factors and potential psychological harms across everyday, role-play, and support uses.

Reference graph

Works this paper leans on

42 extracted references · 6 canonical work pages · cited by 4 Pith papers · 1 internal anchor

  1. [2]

    Use of generative AI for mental health advice among US adolescents and young adults

    McBain RK, Bozick R, Diliberti M, et al. Use of generative AI for mental health advice among US adolescents and young adults. JAMA Netw Open. 2025;8(11):e2542281. doi:10.1001/jamanetworkopen.2025.42281

  2. [3]

    Strengthening ChatGPT’s responses in sensitive conversations

    OpenAI. Strengthening ChatGPT’s responses in sensitive conversations. October 27, 2025. Accessed February 3, 2026. https://openai.com/index/strengthening-chatgpt-responses-in- sensitive-conversations/

  3. [5]

    For argument’s sake, show me how to harm myself: Jailbreaking LLMs in suicide and self-harm contexts

    Schoene AM, Canca C. For argument’s sake, show me how to harm myself: Jailbreaking LLMs in suicide and self-harm contexts. In: Proceedings of the IEEE International Symposium on Technology and Society (ISTAS). 2025. doi:10.1109/ISTAS65609.2025.11269647

  4. [6]

    How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework

    Iftikhar Z, Xiao A, Ransom S, Huang J, Suresh H. How LLM counselors violate ethical standards in mental health practice: A practitioner-informed framework. In: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 2025;8(2):1311-1323. doi:10.1609/aies.v8i2.36632

  5. [7]

    Towards understanding sycophancy in language models

    Sharma M, Tong M, Korbak T, et al. Towards understanding sycophancy in language models. arXiv [preprint]. Published October 20, 2023. arXiv:2310.13548. doi:10.48550/arXiv.2310.13548

  6. [8]

    Longitudinal study on social and emotional use of AI conversational agent

    Chandra M, Hernandez J, Ramos G, Ershadi M, Bhattacharjee A, Amores J, et al. Longitudinal study on social and emotional use of AI conversational agent. arXiv [preprint]. 2025. doi:10.48550/arXiv.2504.14112

  7. [10]

    Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers

    Moore J, Grabb D, Agnew W, Klyman K, Chancellor S, Ong DS, Haber N. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In: VERA-MH 16 Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. ACM; 2025:599-627. doi:10.1145/3715275.3732039

  8. [11]

    Incident 826: Character.ai chatbot allegedly influenced teen user toward suicide amid claims of missing guardrails

    Atherton D. Incident 826: Character.ai chatbot allegedly influenced teen user toward suicide amid claims of missing guardrails. AI Incident Database. 2024. Accessed September 2025. https://incidentdatabase.ai/cite/826/

  9. [12]

    Incident 1192: 16-year-old allegedly received suicide method guidance from ChatGPT before death

    Atherton D. Incident 1192: 16-year-old allegedly received suicide method guidance from ChatGPT before death. AI Incident Database. 2025. Accessed September 2025. https://incidentdatabase.ai/cite/1192/

  10. [13]

    Lawsuits blame ChatGPT for suicides and harmful delusions

    Hill K. Lawsuits blame ChatGPT for suicides and harmful delusions. New York Times. November 6, 2025. Accessed February 3, 2026. https://www.nytimes.com/2025/11/06/technology/chatgpt-lawsuit-suicides-delusions.html

  11. [14]

    VERA-MH concept paper: Validation of ethical and responsible AI in mental health

    Belli L, Bentley K, Alexander W, Ward E, Hawrilenko M, Johnston K, et al. VERA-MH concept paper: Validation of ethical and responsible AI in mental health. arXiv [preprint]. October 17,

  12. [15]

    LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods

    Li H, Dong Q, Chen J, Su H, Zhou Y, Ai Q, et al. LLMs-as-judges: A comprehensive survey on LLM-based evaluation methods. arXiv [preprint]. 2024. doi:10.48550/arXiv.2412.05579

  13. [17]

    Open letter to the AI and technology industry: Protecting youth mental health and preventing suicide in the age of AI

    The Jed Foundation. Open letter to the AI and technology industry: Protecting youth mental health and preventing suicide in the age of AI. September 17, 2025. Accessed February 3, 2026. https://jedfoundation.org/open-letter-to-the-ai-and-technology-industry/

  14. [19]

    Managing suicidal risk: A collaborative approach

    Jobes DA. Managing suicidal risk: A collaborative approach. 2nd ed. Guilford Press; 2016

  15. [20]

    SAFE-T suicide assessment five- step evaluation and triage (PEP24-01-036)

    Substance Abuse and Mental Health Services Administration. SAFE-T suicide assessment five- step evaluation and triage (PEP24-01-036). Published 2024. Accessed October 16, 2025. https://library.samhsa.gov/product/safe-t-suicide-assessment-five-step-evaluation-and- triage/pep24-01-036

  16. [21]

    About QPR

    QPR Institute. About QPR. Accessed February 3, 2026. https://qprinstitute.com/about-qpr/ VERA-MH 17

  17. [22]

    Zero Suicide framework

    Zero Suicide. Zero Suicide framework. Accessed February 3, 2026. https://zerosuicide.edc.org/zero-suicide-framework

  18. [24]

    Can large language models identify implicit suicidal ideation? An empirical evaluation

    Li T, Yang S, Wu J, Wei J, Hu L, Li M, et al. Can large language models identify implicit suicidal ideation? An empirical evaluation. arXiv [preprint]. 2025. arXiv:2502.17899v2. https://arxiv.org/abs/2502.17899v2

  19. [26]

    Content analysis in mass communication: Assessment and reporting of intercoder reliability

    Lombard M, Snyder-Duch J, Bracken CC. Content analysis in mass communication: Assessment and reporting of intercoder reliability. Hum Commun Res. 2002;28(4):587-604. doi:10.1093/hcr/28.4.587

  20. [27]

    MindEval: Benchmarking language models on multi-turn mental health support

    Pombal J, D’Eon M, Guerreiro NM, Martins PH, Farinhas A, Rei R. MindEval: Benchmarking language models on multi-turn mental health support. arXiv [preprint]. 2025. arXiv:2511.18491. doi:10.48550/arXiv.2511.18491

  21. [28]

    Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation

    Eriksson M, Purificato E, Noroozian A, Vinagre J, Chaslot G, Gómez E, Fernández-Llorca D. Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation. arXiv [preprint]. 2025. arXiv:2502.06559

  22. [29]

    AI and the everything in the whole wide world benchmark

    Raji D, Denton E, Bender EM, Hanna A, Paullada A. AI and the everything in the whole wide world benchmark. In: Vanschoren J, Yeung S, eds. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1. NeurIPS; 2021. Accessed February 3, 2026. https://datasets-benchmarks- proceedings.neurips.cc/paper/2021/hash/084b6fbb10729ed...

  23. [30]

    Heterogeneity in suicide risk: Evidence from personalized dynamic models

    Coppersmith DDL, Kleiman EM, Millner AJ, Wang SB, Arizmendi C, Bentley KH, et al. Heterogeneity in suicide risk: Evidence from personalized dynamic models. Behav Res Ther. 2024;180:104574. doi:10.1016/j.brat.2024.104574

  24. [31]

    Disclosure of suicidal ideation and behaviours: A systematic review and meta-analysis of prevalence

    Hallford DJ, Rusanov D, Winestone B, Kaplan R, Fuller-Tyszkiewicz M, Melvin G. Disclosure of suicidal ideation and behaviours: A systematic review and meta-analysis of prevalence. Clin Psychol Rev. 2023;101:102272. doi:10.1016/j.cpr.2023.102272 VERA-MH 18

  25. [32]

    Understanding why patients may not report suicidal ideation at a health care visit prior to a suicide attempt: A qualitative study

    Richards JE, Whiteside U, Ludman EJ, Pabiniak C, Kirlin B, Hidalgo R, Simon G. Understanding why patients may not report suicidal ideation at a health care visit prior to a suicide attempt: A qualitative study. Psychiatr Serv. 2019;70(1):40-45. doi:10.1176/appi.ps.201800342

  26. [33]

    LLMs get lost in multi-turn conversation

    Laban P, Hayashi H, Zhou Y, Neville J. LLMs get lost in multi-turn conversation. arXiv [preprint]. 2025. arXiv:2505.06120. doi:10.48550/arXiv.2505.06120

  27. [34]

    Helping people when they need it most

    OpenAI. Helping people when they need it most. August 26, 2025. Accessed February 3, 2026. https://openai.com/index/helping-people-when-they-need-it-most/

  28. [35]

    High Potential for Harm

    Keshavan M, Torous J, Yassin W. Do generative AI chatbots increase psychosis risk? World Psychiatry. 2026;25(1):150-151. doi:10.1002/wps.70017 VERA-MH 19 Supplemental Tables Supplemental Table S1 Overview of VERA-MH User-Agent Profiles Demographics Mental Health Background and Recent Stressors Suicidal Thoughts and Behaviors Communication Style and Respon...

  29. [37]

    A principle- based framework for the development and evaluation of large language models for health and wellness

    Winslow B, Shreibati J, Perez J, Su HW, Young-Lin N, Hammerquist N, et al. A principle- based framework for the development and evaluation of large language models for health and wellness. arXiv [preprint]. 2025. arXiv:2512.08936. doi:10.48550/arXiv.2512.08936

  30. [38]

    Health advisory: Use of generative AI chatbots and wellness applications for mental health

    American Psychological Association. Health advisory: Use of generative AI chatbots and wellness applications for mental health. Published 2025. Accessed February 3, 2026. https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory- chatbots-wellness-apps/

  31. [39]

    Independent clinical evaluation of general-purpose LLM responses to signals of suicide risk

    Judd N, Vaz A, Paeth K, Davis LI, Esherick M, Brand J, et al. Independent clinical evaluation of general-purpose LLM responses to signals of suicide risk. arXiv [preprint]. 2025. https://arxiv.org/pdf/2510.27521

  32. [40]

    Suicidal ideation detection in conversational AI: A compliance and implementation guide

    ThroughLine. Suicidal ideation detection in conversational AI: A compliance and implementation guide. Published 2026. Accessed February 3, 2026. https://cdn.prod.website- files.com/693a4901d87d2a9fd3bf0d09/69448c563c28a5ee29113ca5_Throughline%20- %20Whitepaper.pdf

  33. [41]

    Does asking about suicide and related behaviours induce suicidal ideation? What is the evidence? Psychol Med

    Dazzi T, Gribble R, Wessely S, Fear NT. Does asking about suicide and related behaviours induce suicidal ideation? What is the evidence? Psychol Med. 2014;44(16):3361-3363. doi:10.1017/S0033291714001299

  34. [42]

    About QPR

    QPR Institute. About QPR. Accessed February 3, 2026. https://qprinstitute.com/about-qpr/

  35. [43]

    Making chatbots safe for suicidal patients

    Frances A, Whiteside U. Making chatbots safe for suicidal patients. Psychiatric Times. November 18, 2025. Accessed February 3, 2026. https://www.psychiatrictimes.com/view/making-chatbots-safe-for-suicidal-patients

  36. [44]

    Food and Drug Administration

    U.S. Food and Drug Administration. Executive summary for the Digital Health Advisory Committee meeting: Generative artificial intelligence-enabled digital mental health medical devices. November 6, 2025. Accessed February 3, 2026. https://www.fda.gov/media/189391/download

  37. [45]

    Open letter on AI safety

    Now Matters Now. Open letter on AI safety. Published 2025. Accessed February 3, 2026. https://nowmattersnow.org/open-letter-on-ai-safety/

  38. [46]

    The case against no-suicide contracts: The commitment to treatment statement as a practice alternative

    Rudd MD, Mandrusiak M, Joiner TE. The case against no-suicide contracts: The commitment to treatment statement as a practice alternative. J Clin Psychol. 2006;62(2):243-

  39. [48]

    Evaluation of alignment between large language models and expert clinicians in suicide risk assessment

    McBain RK, Cantor JH, Zhang LA, Baker O, Zhang F, Burnett A, et al. Evaluation of alignment between large language models and expert clinicians in suicide risk assessment. Psychiatr Serv. Published online 2025. doi:10.1176/appi.ps.20250086

  40. [49]

    Content analysis: An introduction to its methodology

    Krippendorff K. Content analysis: An introduction to its methodology. 4th ed. Sage Publications; 2018

  41. [251]

    doi:10.1002/jclp.20227 VERA-MH 33

  42. [2025]

    VERA-MH Concept Paper

    arXiv:2510.15297. doi:10.48550/arXiv.2510.15297