{"id":"bd352a54-2f48-4994-a740-68ea85bff990","arxiv_id":"2509.08839","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Across six LLMs responding to high-risk mental health prompts, none consistently met five clinician-defined crisis-response standards, with Claude scoring highest at 0.88 out of 1.","lead":"This paper tested how six popular chatbots respond to crisis-level mental health statements, scoring five clinician-defined safety behaviors. It finds most models sound caring but often fail to name the danger, give crisis resources, or invite the user to keep talking.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Composite safety score lacks criterion validity and the 'satisfactory standard' threshold is circular, so the conclusion that no LLM is clinically safe by default does not follow.","rationale":"The reader's verdict correctly identifies the coding framework's validity as the weakest assumption. The empirical results are only as good as the outcome measure: the five codes are reasonable surface behaviors, but the paper never demonstrates that they capture 'clinical safety' in a way that supports a universal negative. The Discussion's own caveat about AI-specific guidelines is a direct admission that the framework may not be appropriate. Additionally, the paper never specifies a threshold for 'satisfactory'; it moves from 'no model scored 1.0 on all codes' to 'no model is safe by default,' which is a non-sequitur. The data-collection inconsistency (68 prompts vs 30 per model) is a serious concern that also merits mention, but the construct validity issue is more load-bearing because it would remain even if the counts were corrected. A concrete test—comparing composite scores to independent clinician safety judgments—would settle whether the coding framework is a valid measure. I therefore agree with the reader's REJECT verdict and would not change it.","tokens_in":11869,"tokens_out":10559,"duration_ms":112064,"concrete_test":"Recruit a new panel of 3–5 licensed clinicians with crisis-intervention experience, blind to the study's coding framework. Present all 180 LLM responses (or a representative random subset) and ask each clinician to rate global clinical safety on a validated scale (e.g., safety assessment scale) and to independently identify which response features they consider essential for safety. Compute the correlation between mean clinician global safety ratings and the composite scores reported in Figure 6, and compare clinician-identified essential features to the five codes. If the correlation is weak or clinicians highlight behaviors not captured by the framework, the five-code composite is not a valid proxy for clinical safety, and the conclusion that no LLM is safe by default is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'no current general-purpose LLM can be considered clinically safe by default'—is a universal negative that depends entirely on the validity of the five-code composite score constructed in 'Developing a Clinically Grounded Coding Framework' and 'Coding and Scoring Procedure.' The paper assumes, without independent validation, that the five binary codes (explicit risk acknowledgment, empathy, help-seeking encouragement, specific resources, continuation invitation) are necessary and sufficient for 'minimum standards' of crisis response, and that an unweighted mean of these codes is a meaningful safety metric. The Discussion itself concedes that 'clinical guidelines explicitly tailored for AI-driven therapy might diverge from current standards' (Discussion, final paragraph before Conclusions), undermining the framework's relevance to LLMs. Even if the codes were valid, the threshold for 'satisfactory' is never defined: the Discussion states 'no model excelled across all behaviors' and the abstract concludes none meet satisfactory standards, but the only explicit criterion is that no model achieved perfect scores on all five codes. This is circular: calling a model 'not safe' because it fails an arbitrary perfect-score bar is a definitional artifact, not an empirical finding. A model could score 0.88 and still be reasonably safe in most crisis situations, while a model scoring 1.0 could still produce harmful advice not captured by these surface behaviors. Thus, the central claim requires external criterion validation, which is absent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates six commercial/open-weight LLMs (Claude, Gemini, Deepseek, ChatGPT, Grok, LLAMA) on their responses to one-shot high-risk mental-health disclosures. A five-code framework (explicit risk acknowledgment, empathy, help-seeking encouragement, specific resources, continuation invitation) was developed by licensed clinicians, and three expert raters coded model outputs with reported inter-rater agreement (Fleiss' κ = 0.775). Each code was scored as present/absent per response, averaged per model, and combined into an unweighted composite safety metric. The paper reports that Claude scores highest (0.88), Gemini and Deepseek intermediate, and Grok 3, ChatGPT, and LLAMA below 0.5. The authors conclude that none of the tested models can be considered clinically safe by default, and recommend human oversight and targeted fine-tuning.","tokens_in":12140,"tokens_out":4024,"duration_ms":47614,"significance":"If the empirical claims were supported, the paper would address a timely public-health question: whether general-purpose LLMs, already used informally for crisis support, meet minimal standards of safe crisis communication. Strengths include a clinician-derived coding framework, use of multiple expert raters with quantified agreement, concrete response examples per code, and an explicit acknowledgment of limitations (single-turn prompts, English-only, version drift). The comparative, behavior-level decomposition is useful for guiding future safety fine-tuning. However, the central conclusion depends on the validity and operationalization of the five-code composite and on a threshold for 'satisfactory clinical standards' that is not defined. The manuscript also contains important methodological inconsistencies (prompt counts, model version) that currently undermine the reliability of the reported scores.","major_comments":[{"comment":"The prompt inventory is internally inconsistent and load-bearing. The text lists six psychiatric-emergency domains with counts 13, 12, 10, 14, 10, and 9, which sum to 68 prompts per model, yet immediately states 'In total, 180 responses were evaluated (30 per model).' With six models, 30 per model implies 180 responses, but 68 prompts per model would imply 408. If only a subset of 30 prompts was used, the selection procedure is not described. All reported code averages and the composite scores depend on which prompts were actually administered, so this discrepancy must be resolved before the results can be interpreted.","section":"Method, Data collection"},{"comment":"The composite safety score is an unweighted mean of five binary codes, but the paper provides no justification for equal weighting or for the claim that these five codes are necessary and sufficient for 'minimum standards' of crisis response. The assumption that each code is always beneficial ('having is better') is asserted rather than validated; for example, an invitation to continue the conversation might be inappropriate or even risky in the absence of human follow-up, and provision of specific resources could be harmful if the referral is inaccurate. Furthermore, the paper never defines the cutoff for 'satisfactory clinical standards'; the only observable criterion appears to be a perfect score on all five codes. Thus the conclusion that 'no current general-purpose LLM can be considered clinically safe by default' follows from an arbitrary perfect-score threshold, not from an empiri","section":"Method, Coding and Scoring Procedure; Discussion and Conclusions"},{"comment":"The inter-rater reliability is computed on 'a shared set of LLM responses,' but the size and selection of this set are not reported. It is unclear whether all 180 (or 408) responses were double-coded or only a subset. Additionally, the scoring heuristic includes a 0.5 rule for equal numbers of 'Yes' and 'No' ratings, but with three raters, a tie is impossible (only 0, 1, 2, or 3 Yes ratings can occur). This suggests either a different number of raters than stated, missing ratings, or an error in the description. Please clarify the coding procedure and the actual distribution of rater counts.","section":"Method, Coding and Scoring Procedure"},{"comment":"The model versions are inconsistent. The abstract and all results/figures refer to 'Grok 3,' but the Method section states 'Grok (Model: Grok4).' Since the entire study is time-sensitive and model versions can change behavior, this discrepancy must be corrected. The same issue applies to the exact version of ChatGPT (gpt-4.1) and others; please standardize the naming throughout.","section":"Abstract; Method, Data collection; Results"},{"comment":"The comparative claims ('Claude outperformed all others,' 'Grok 3, ChatGPT, and LLAMA underperformed') are presented without any uncertainty quantification. The reported values are sample means from a small number of prompts (at most 30 per model, possibly 68), yet no confidence intervals, standard errors, or statistical tests are provided. Given the small sample and the coding-based measurement, the absence of inferential statistics makes it impossible to assess whether the observed differences are meaningful. This is particularly important for the domain-level comparisons, where some models differ by small margins.","section":"Results, Overall performance comparison"}],"minor_comments":[{"comment":"The text says 'five psychiatric-emergency domains' but then lists six (Threats to Self, Threats to Harm Others, Domestic Violence, Psychotic Symptoms, Inappropriate Behavior regarding Children, Dangerous Neglect of Dependent Adults). The number should be corrected, and the category labels should be reconciled with the later mention of 'suicidality, self-harm, domestic violence, psychosis, and exploitation.'","section":"Method, Data collection"},{"comment":"There are typographical inconsistencies: 'Deepsek' appears instead of 'Deepseek' in Figures 1 and in the text; 'showing' is misspelled as 'showing' in Table 2; and the reference to 'Janse van Rensburg & and van der Wath' has an extra 'and.' Please proofread.","section":"Results, Figures 1–6 and Tables"},{"comment":"The claim that 'safety is not an emergent property of scale' is presented as a conclusion, but the design does not manipulate scale or training approach; it is an interpretative speculation. Please mark it as such.","section":"Discussion and Conclusions"},{"comment":"The estimate of 'over 18 million psychotherapy/counseling conversations per month' for ChatGPT relies on a Similarweb usage figure and an extrapolation from Claude's proportion. This is presented without acknowledging the substantial uncertainty in these numbers. Please soften or add a caveat.","section":"Introduction, Literature Review"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the underlying idea of using clinician-derived codes to evaluate LLM crisis responses is valuable. However, the current manuscript contains several load-bearing methodological inconsistencies (prompt counts, model version, and coding details) and an undefined threshold for the central safety conclusion. These are fixable in principle, but they require more than minor edits. I would encourage the editor to request a revised version that either corrects the data collection description and re-analyzes accordingly, or substantially narrows the claims to a descriptive comparison of model behaviors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is one of the first side-by-side, clinician-coded comparisons of six named LLMs on high-risk mental-health disclosures, and the five-code framework is sensible and practically motivated. Second, the paper currently has internal inconsistencies—the prompt count doesn't add up, and the reported tie rule can't fire with three raters—so the numbers need verification before the 'no model is safe' claim can be trusted.\n\nWhat's good: the framework was built by four licensed therapists, the codes (risk acknowledgment, empathy, help-seeking, specific resources, continuation) are concrete and clinically meaningful, and inter-rater reliability (Fleiss' kappa 0.775) is respectable. The qualitative examples in the tables make the ranking tangible. The authors also openly list limitations: single-turn English prompts, six models at a point in time, no outcomes. That honesty counts.\n\nWhere it gets soft: the central claim—'no current general-purpose LLM can be considered clinically safe by default'—does not follow from the data. The paper never defines a satisfactory threshold; it shows no model scored perfectly across all five codes, then treats that as 'not safe.' But the codes are minimal necessary behaviors, not a sufficient safety test. A model that scores 1.0 on all five could still give harmful advice, and one scoring 0.88 could be reasonably safe in most crises. The Discussion itself concedes that AI-specific guidelines 'might diverge from current standards,' which undercuts the assumption that these five human-therapy codes are the right yardstick. So the conclusion is an overreach. The safer reading is the descriptive one: models vary, most empathize, few consistently give resources, and some (Grok 3, ChatGPT, LLAMA) fall below 0.5 on the composite.\n\nAlso, the data section is inconsistent: the six domains sum to 68 prompts, yet the paper reports 180 responses at 30 per model—a factor-of-two-plus discrepancy. The abstract says Grok 3, the method says Grok4. And the 0.5 tie rule is impossible with three raters. These aren't fatal to the qualitative finding, but they make the composite scores hard to trust. No data or code are shared, which doesn't help.\n\nBottom line: this is a useful descriptive snapshot for people thinking about chatbot safety, and it deserves a serious referee—but only after the numerical inconsistencies are fixed and the safety claim is toned down to what the data show. I'd send it to peer review with a request for major revision, not desk reject it.","headline":"A useful descriptive snapshot of six LLMs' crisis responses, but the headline claim that no model is clinically safe overreaches the data and the paper has internal inconsistencies that need fixing.","tokens_in":12668,"tokens_out":3318,"would_cite":false,"duration_ms":36759,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"No general-purpose LLM currently meets clinical safety standards for responding to mental-health crises, a comparative test of six chatbots finds.","keywords":["large language models","crisis intervention","ethics","mental health","clinical safety","high-risk disclosure","suicide risk","chatbot evaluation"],"falsifier":"A concrete check would rerun the evaluation as multi-turn conversations—five exchanges per crisis prompt, same five codes, rated by a fresh clinician panel. If most models recover by later turns and satisfy all five codes, the blanket 'none safe by default' claim would be overturned; in parallel, reconciling the stated 68 vetted prompts with the reported 180 scored responses would settle whether the sample is complete.","tokens_in":11768,"feed_emoji":"⚠️","tokens_out":13287,"duration_ms":139636,"temperature":0.7,"pith_summary":"This paper compares how six popular chatbots respond to one-shot crisis disclosures—suicidal intent, threats to harm others, domestic abuse, psychosis, child exploitation, and dangerous neglect—drafted and vetted by four licensed therapists. A clinician-built coding framework scored each response on five minimum-standard behaviors: naming the risk explicitly, showing empathy, encouraging help-seeking, providing specific resources, and inviting the user to continue talking. The paper's central claim is that none of the six models is clinically safe by default: Claude scored highest on the composite (0.88), Gemini and DeepSeek fell in the middle, and Grok 3, ChatGPT, and Llama scored below 0.5. The stakes the authors give are concrete: people already use general-purpose LLMs for mental-health support in numbers comparable to major health-system caseloads, yet the tools sit outside current medical-device and high-risk AI regulation. The authors caution that the absolute scores are tied to a small, single-turn, English-only snapshot.","feed_headline":"No chatbot meets clinical safety standards in mental-health crises","feed_subtitle":"Six popular LLMs scored on five crisis behaviors; none reached clinician-defined minimums.","key_machinery":"The load-bearing instrument is a five-code coding framework developed by four licensed clinicians, each code corresponding to one behavior treated as a minimum for a safe crisis response: explicit acknowledgment of risk, empathy, encouragement to seek help, provision of specific resources, and invitation to continue the conversation. Independent raters applied the codes to model outputs and reached substantial agreement (κ = 0.775). Each code was scored as present or absent and averaged across raters and prompts, and the overall safety score was the simple mean of the five code averages. This framework converts 'clinically safe' from a vague impression into five separately measurable behavio","core_discovery":"The paper's central claim, stated in the Discussion, is that 'no current general-purpose LLM can be considered clinically safe by default.' The supporting result is a comparative scorecard for six chatbots, built from five clinician-defined behaviors and scored 0–1 per behavior. The behaviors do not move together: most models expressed empathy in most responses, but empathy did not predict whether a model acknowledged the risk, gave a concrete resource, or invited the user to continue. Claude was the only model to acknowledge risk in every response and the strongest overall (composite 0.88); Grok 3 supplied no specific resources in any response, and ChatGPT combined a perfect empathy score w","pith_inferences":["The composite metric treats the five behaviors as equally important. A different weighting—say, explicit risk acknowledgment weighted above empathy—would probably shift the rank order even if the headline conclusion about no model being safe by default remains.","The methods text lists 68 vetted prompts by domain but reports 180 scored responses (30 per model), without saying how the 30 were selected; until that is clarified, the per-model averages are best read as indicative rather than a complete census of the prompt set.","A natural extension is to use these five codes as a fine-tuning reward signal: if a model fine-tuned directly on clinician-defined risk acknowledgment, resource provision, and continuation invitations moves above 0.9 on the same composite, it would confirm the paper's claim that safety is a designed and optimizable property, not an emergent one.","The same five-code lens could be applied to purpose-built mental-health chatbots, voice assistants, and automated crisis lines; if dedicated tools also fail on resource provision or continuation invitations, the paper's 'safety is deliberate design' message generalizes beyond general-purpose LLMs."],"forward_implications":["General-purpose LLMs should not be the sole responder when a user signals acute psychological risk; the paper is explicit that human oversight remains indispensable and that clinicians should treat these tools as adjuncts.","High empathy does not equal safety: ChatGPT's perfect empathy score came with a composite below 0.5, so screening tools that measure tone alone would miss practical-care deficits.","Safety behaviors are learnable and optimizable: specific models led on specific codes, and the paper identifies resource provision and continuation invitations as cheap, high-impact targets for fine-tuning.","Regulatory guidance lags behind deployment: because general-purpose chat software is not classed as a medical device or high-risk AI in the U.S. or EU frameworks the paper cites, standards for crisis-response quality must come from vendors and clinicians first.","The results are a snapshot, not a verdict for all time: the authors note model weights and safety layers change silently, and future evaluations should be longitudinal and non-English before clinical deployment."],"supporting_citations":[{"why":"Supplies the clinical risk-assessment standard the paper holds LLMs to.","marker":"(Fowler, 2012)"},{"why":"Defines the ethics-of-care expectations for therapists that the five codes operationalize.","marker":"(Pope & Vasquez, 2016)"},{"why":"The crisis-communication literature the coding framework was explicitly built from.","marker":"(Cole-King et al., 2013)"},{"why":"National survey showing 48.7% of LLM-using respondents with mental-health diagnoses used LLMs for support; establishes the public-safety stakes.","marker":"Rousmaniere, Zhang, Li, and Shah (2025)"},{"why":"Reports the share of Claude conversations involving therapy/counseling, used to estimate monthly quasi-clinical session volume.","marker":"(McCain et al., 2025)"},{"why":"Device-software guidance that the paper cites to show chat-based LLMs fall outside FDA medical-device oversight.","marker":"(Food & Drug Administration, 2023)"},{"why":"The EU AI Act provisions that the paper says leave general-purpose mental-health LLMs out of the high-risk category.","marker":"(European Union, 2024)"},{"why":"Prior vignette study finding GPT-4 suicide-risk assessments aligned with professionals; the main clinical-evaluation precedent this paper extends.","marker":"(Levkovich & Elyoseph, 2023)"}],"fun_headline_variants":["No chatbot passes clinician safety bar for mental-health crises","Empathy not enough: LLMs fail crisis-response safety checks","Six LLMs scored, none clinically safe for mental-health crises","LLMs miss safety essentials despite empathy in crisis replies","Even strongest LLM falls short of clinical crisis-safety standard"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The conclusion depends on accepting that five equally weighted clinician-defined behaviors—risk acknowledgment, empathy, help-seeking encouragement, specific resources, continuation invitation—are a valid and sufficient definition of a minimally safe crisis response, a premise the paper itself flags as open to revision by AI-specific guidelines.","fun_headline_variants_meta":{"raw":{"variants":["No chatbot passes clinician safety bar for mental-health crises","Empathy not enough: LLMs fail crisis-response safety checks","Six LLMs scored, none clinically safe for mental-health crises","LLMs miss safety essentials despite empathy in crisis replies","Even strongest LLM falls short of clinical crisis-safety standard"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000937,"raw_usage":{"total_tokens":3821,"prompt_tokens":698,"completion_tokens":3123,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":442,"completion_tokens_details":{"reasoning_tokens":3041}},"tokens_in":442,"tokens_out":3123,"duration_ms":22802,"temperature":1.0,"reasoning_tokens":3041,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:23:09.734334+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would rerun the evaluation as multi-turn conversations—five exchanges per crisis prompt, same five codes, rated by a fresh clinician panel. If most models recover by later turns and satisfy all five codes, the blanket 'none safe by default' claim would be overturned; in parallel, reconciling the stated 68 vetted prompts with the reported 180 scored responses would settle whether the sample is complete.","supporting_citations":[],"review_version":1}