Pith. sign in

REVIEW 3 major objections 6 minor 102 references

Even the strongest large language models still trail clinicians by 37 percentage points on complete psychiatric encounters, with mental-status assessment as the main bottleneck.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 10:29 UTC pith:V4B3HWYH

load-bearing objection Solid full-S.O.A.P. psychiatric benchmark with real multi-center EHRs and specialist-aligned judges; the 37-point gap is real enough to cite, with a modest metric-calibration caveat on mental-status coverage. the 3 major comments →

arxiv 2607.08257 v1 pith:V4B3HWYH submitted 2026-07-09 cs.AI

MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

classification cs.AI
keywords psychiatric clinical encounterslarge language modelsS.O.A.P. workflowstandardized patientselectronic health recordsdual-track evaluationMentalEvalmental status assessment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that success on isolated psychiatric tasks—dialogue, diagnosis, or treatment planning—does not show that language models can run a full clinical encounter. It builds MentalHospital, a virtual hospital that forces models through the entire S.O.A.P. workflow: interview a skill-augmented standardized patient drawn from real electronic health records, order exams, write notes, diagnose, and plan treatment. Encounters are scored two ways: objective recovery of EHR-derived clinical facts, and process quality judged by MentalEval, five specialist-trained scorers for empathy, professionalism, notes, diagnostic rigor, and treatment fit. Clinicians rate the environment as clinically plausible, and MentalEval tracks expert ratings closely. Under this protocol, even the best models lag medical trainees and experts by large margins on objective competence, especially when they must elicit and use mental-status findings.

Core claim

Full-process psychiatric competence is still far from clinician level: the strongest medical-specific model trails medical trainees by about 27 points and human experts by about 37 points on average objective metrics across interviewing, examination, notes, category diagnosis, and disorder diagnosis, with mental-status coverage as a recurring failure mode. Subjective process scores are mixed—models can match or exceed humans on expressed empathy and treatment appropriateness while remaining weaker on interviewing professionalism and note quality—so models are not substitutes, but may complement clinicians in affective support and treatment assistance.

What carries the argument

MentalHospital: an EHR-grounded S.O.A.P. simulation using skill-augmented standardized patients (role plus presentation skill plus topic-level memory skill) from 1,193 de-identified cases covering all major ICD-11 psychiatric categories and 76 disorders, paired with a dual-track evaluation protocol and MentalEval—five Qwen3-8B evaluators trained by rubric-grounded supervised fine-tuning then expert-guided preference optimization—to scale specialist judgment of process quality.

Load-bearing premise

That the de-identified EHR checkpoints and the skill-augmented patients form a faithful, non-leaking gold standard for what a doctor should recover and how a real psychiatric patient would present.

What would settle it

Re-run the same doctor agents on a held-out multi-center set with independent psychiatrist re-annotation of checkpoints and live standardized-patient sessions; if the LLM–clinician objective gap shrinks below roughly ten points or mental-status coverage ceases to be the dominant miss, the measured competence gap is an artifact of the current patient construction.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Psychiatric AI evaluation must move from single-turn or dialogue-only tests to full interview–exam–note–diagnosis–treatment episodes with EHR-grounded references.
  • Mental-status examination becomes a primary training and evaluation target, not an optional side skill.
  • Specialist-aligned judges (rubric SFT + expert preference) can replace general LLM-as-judge for scalable process scoring in psychiatry.
  • LLMs may be useful as empathic communication and treatment-drafting assistants while remaining unsuitable as autonomous clinicians under this protocol.
  • Controlled access to de-identified EHR-derived cases plus public environment code can become a standard for psychiatric agent benchmarks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If mental-status probing is the bottleneck, interview curricula that force explicit MSE checklists before diagnosis may close more of the gap than larger generic models alone.
  • The dual-track split implies future model cards should report objective evidence recovery and process quality separately rather than a single medical accuracy score.
  • Because comorbid and multi-center cases are already in the bank, the same scaffold could stress-test safety behaviors (self-harm, psychosis reinforcement) without inventing synthetic patients from scratch.
  • Clinician survey positivity for training suitability suggests the environment may first land as a trainee simulator even if model scores remain low.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces MentalHospital, an EHR-grounded virtual environment for evaluating LLM psychiatric clinical encounters under a full S.O.A.P. workflow, built from 1,193 de-identified multi-center cases spanning all major ICD-11 categories and 76 disorders. Skill-augmented standardized patients (representation + memory skills) and a hospital examination module force agents to elicit evidence rather than read the chart. Evaluation is dual-track: objective coverage against EHR-derived checkpoints and subjective process quality via MentalEval, five Qwen3-8B evaluators trained with rubric-grounded SFT then expert-guided DPO. A 22-clinician survey rates clinical fidelity at 3.88/5; MentalEval reaches average QWK 0.944 on held-out expert labels. Benchmarking of 12 LLMs against experts and trainees reports that the strongest model trails clinicians by 37.28 percentage points on objective metrics, with mental-status assessment as the principal bottleneck, while LLMs show complementary strengths in expressed empathy and treatment appropriateness.

Significance. If the environment and dual-track protocol hold under scrutiny, this is a substantial contribution to medical AI evaluation: it moves psychiatric LLM assessment from isolated dialogue/diagnosis tasks to complete, EHR-anchored encounters with both outcome and process measures. Strengths include multi-center real EHR grounding, explicit patient-construction ablations (Table 4), specialist-aligned evaluators with strong held-out QWK, multi-group human baselines (experts, trainees, crowdworkers), and a clear, falsifiable bottleneck claim on mental-status recovery. The resource-release protocol (code, rubrics, MentalEval weights public; raw EHRs controlled) is a responsible compromise. These elements make the work useful for training, benchmarking, and diagnosing where current LLMs fail in psychiatry, independent of any single headline number.

major comments (3)
  1. [§4, Eq. (13), Tables 2–3] §4, Eq. (13) and Tables 2–3: the headline 37.28 pp LLM–clinician gap is defined by Coverage(yc, Kc) with a hand-chosen semantic threshold τ=0.85 plus Multi-LLM adjudication for pairs below threshold. Clinicians and LLMs interact with the same memory-gated patients, but the paper does not report a calibration study of human vs. LLM utterance matching on identical elicited content (inter-rater agreement on k ⪯ yc, or re-scoring of clinician transcripts under the same automatic matcher). Table 3’s large CC–MS gaps for LLMs could therefore partly reflect phrasing/style sensitivity of the matcher rather than pure competence. A load-bearing fix is to (i) report human–LLM matching agreement on a shared probe set, (ii) ablate τ, and (iii) recompute the gap under exact/grounding-based coverage where available (Appendix K already logs patient grounding fields).
  2. [§2.2, Table 4, Appendix L] §2.2, Table 4, Appendix L: the claim that skill-augmented patients constitute a faithful, non-leaking gold standard rests on a small objective probe set (20 cases × 12 probes = 240 responses) and three-clinician subjective ratings. There is no quantitative check that de-identification (Appendix F) preserved psychopathological logic at the checkpoint level, nor a leakage audit showing that patients never disclose unasked future-stage evidence under adversarial doctor prompts. Because both objective coverage and the LLM–clinician gap are measured against these patients, expand the fidelity study (more cases, adversarial probes, inter-psychiatrist agreement on whether disclosed content matches the original EHR logic) or qualify the gap as conditional on the current patient construction.
  3. [§3.1, Table 5] §3.1 and Table 5: MentalEval’s cold-start SFT is supervised by a five-LLM judge ensemble that the paper itself reports as poorly aligned with specialists (LLM-as-a-Judge QWK 0.677, Acc. 0.225). Although expert-guided DPO on low-confidence sets raises average QWK to 0.944 on held-out cases, residual dependence on weak LLM judges for the bulk of SFT targets is a load-bearing design choice for scalable subjective scores. Report (a) the fraction of SFT data that survived consensus filtering vs. was rewritten/augmented, (b) agreement of SFT-only vs. SFT+DPO evaluators stratified by score extremity, and (c) whether clinician DPO preferences were collected independently of the models being ranked in Table 2.
minor comments (6)
  1. [Abstract / §1 / §4.5] Abstract and §4.5 report clinician fidelity as 3.88/5 while the introduction states 3.96/5; reconcile the two figures and state which sample (experts only vs. experts+trainees) each uses.
  2. [Table 2, §4.1] Table 2 lists Empathy scores where lower appears better for humans (experts 1.22) but higher for some LLMs; clarify whether the empathy rubric is inverted relative to other 1–5 dimensions or whether experts deliberately suppress affective language.
  3. [§2, Eq. (1)–(4)] Eq. (1) uses xc = {Kpat_c, Kexam_c} and Y*_c; later ˆyc uses different symbols for the same conceptual objects. A short notation table in §2 would reduce reader load.
  4. [Figure 3] Figure 3 confusion matrices are hard to read in grayscale; add numeric cell annotations and a shared color scale.
  5. [§6 / Appendix A] Appendix A Limitations correctly flags missing safety/adversarial evaluation (self-harm, psychosis reinforcement); a short pointer in the main-text conclusion would set expectations for deployment claims.
  6. [Table 2, §4] Several model names appear inconsistently (Deepseek-v4-Pro vs. DeepSeek-V4-Pro; Claude-Sonnet-4.6 vs. Claude-Sonnet-4-6). Normalize throughout.

Circularity Check

0 steps flagged

No load-bearing circularity: the 37.28 pp LLM–clinician gap is measured by independent EHR-checkpoint coverage, not by construction from the evaluators or fitted parameters.

full rationale

MentalHospital’s central empirical claim (Table 2, §4.1–4.2) is an objective competence gap obtained by running LLMs and human clinicians through the same S.O.A.P. episodes and scoring Coverage(yc, Kc) against EHR-derived checkpoints (Eq. 13, Appendix K). Those checkpoints are extracted from de-identified multi-center EHRs that pre-exist the benchmark; they are not fitted to model outputs, nor are they defined in terms of the MentalEval scores. MentalEval itself is trained on multi-LLM trajectories plus expert DPO preferences on held-out cases and is used only for the subjective track; the headline 37.28-point figure does not depend on it. Patient construction (representation + memory skills) is ablated for fidelity (Table 4) but does not redefine the reference targets. There are no self-definitional equations, no fitted parameters re-labeled as predictions, no uniqueness theorems imported from overlapping authors, and no ansatz smuggled via self-citation that forces the reported gap. Minor self-reference exists only in the ordinary sense that the authors built the environment they evaluate; that does not reduce the measured gap to an input by construction. The derivation is therefore self-contained against external EHR gold standards.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 2 invented entities

The central performance gap rests on treating de-identified multi-center EHRs as ground-truth clinical targets, S.O.A.P. as the canonical workflow, and skill-augmented patients plus MentalEval as faithful proxies for real patients and specialist judgment. Free parameters are mainly matching and consensus thresholds; invented entities are the environment and evaluators themselves.

free parameters (3)
  • semantic similarity threshold τ = 0.85
    Set to 0.85 for high-precision automatic checkpoint matching; pairs below threshold go to Multi-LLM adjudication (§4).
  • judge consensus k
    Minimum number of the five-LLM ensemble that must agree on a score before an SFT trajectory is retained (§3.1).
  • DPO β = 0.3
    Preference coefficient fixed at 0.3 for expert-guided DPO of MentalEval (Appendix J).
axioms (3)
  • domain assumption S.O.A.P. workflow is the appropriate complete clinical encounter structure for psychiatric evaluation
    Instantiated throughout §2 and Figure 1; taken from nursing/medical documentation literature without further justification.
  • domain assumption De-identified EHR checkpoints remain clinically complete and diagnostically valid gold standards after privacy rewriting
    Asserted via dual-psychiatrist review (κ>0.8) and privacy audit (Appendix F); load-bearing for all objective metrics.
  • ad hoc to paper Rubric-grounded SFT + expert DPO produces evaluators whose 1–5 scores can stand in for specialist judgment at scale
    Core of MentalEval (§3); validated only on the paper’s own held-out set.
invented entities (2)
  • MentalHospital environment (skill-augmented standardized patients + examination module) no independent evidence
    purpose: Provide interactive full S.O.A.P. psychiatric encounters grounded in real EHRs
    New simulation stack; independent evidence limited to internal clinician survey and ablations.
  • MentalEval (five Qwen3-8B domain-specific evaluators) no independent evidence
    purpose: Scale subjective clinical-process scoring beyond expensive human review
    Trained and validated only inside this paper’s data and rubrics.

pith-pipeline@v1.1.0-grok45 · 42458 in / 2676 out tokens · 29960 ms · 2026-07-10T10:29:20.233429+00:00 · methodology

0 comments
read the original abstract

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.

Figures

Figures reproduced from arXiv: 2607.08257 by Haoyang Zeng, Jiang Zhong, Jingwang Huang, KaiWen Wei, Xiao Sun, Yuanwei Zou, Yuming Yang, Yun Chen, Zhengxiao Wu.

Figure 1
Figure 1. Figure 1: Overview of MentalHospital. A virtual psychiatric evaluation environment for LLM-based [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of MentalEval construction. MentalEval is trained on MentalHospital interaction [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: MentalEval achieves the strongest expert alignment, with the highest QWK (0.926) and [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Clinical fidelity ratings for MentalHospital. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Representative cases in MentalHospital [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Age distribution in MentalHospital. Left: single-diagnosis cases; right: comorbid cases. [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Disorder distribution. Left: single-diagnosis cases; right: comorbid cases. [PITH_FULL_IMAGE:figures/full_fig_p016_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Specific Disorders List. A list of 76 disorders mentioned in the paper, including their corresponding codes and specific names. conflicts were excluded from benchmark construction. Disagreements between the two reviewers were resolved through expert consensus. EHR Content and Case Structure. Each finalized EHR contains three types of clinical infor￾mation. To standardize these fields, we use a structured e… view at source ↗
Figure 9
Figure 9. Figure 9: System prompt for structuring psychiatric clinical records into standardized JSON format. [PITH_FULL_IMAGE:figures/full_fig_p019_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Example of a original clinical record in MentalHospital. [PITH_FULL_IMAGE:figures/full_fig_p020_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Example of a structured psychiatric clinical record in MentalHospital. [PITH_FULL_IMAGE:figures/full_fig_p021_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: System prompt for privacy-preserving anonymization of psychiatric admission records. [PITH_FULL_IMAGE:figures/full_fig_p022_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: System prompt for the Clinical History Collection stage in the psychiatric clinical [PITH_FULL_IMAGE:figures/full_fig_p023_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: System prompt for auxiliary examination recommendation in the psychiatric clinical [PITH_FULL_IMAGE:figures/full_fig_p024_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: System prompt for clinical record generation in the psychiatric clinical workflow. [PITH_FULL_IMAGE:figures/full_fig_p025_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: System prompt for evidence-based definitive diagnosis in the psychiatric clinical workflow. [PITH_FULL_IMAGE:figures/full_fig_p026_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: System prompt for treatment planning in the psychiatric clinical workflow. [PITH_FULL_IMAGE:figures/full_fig_p027_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Overview of standardized patient construction in MentalHospital. De-identified psychiatric [PITH_FULL_IMAGE:figures/full_fig_p027_18.png] view at source ↗
Figure 19
Figure 19. Figure 19: System prompt for standardized psychiatric patient simulation in MentalHospital. [PITH_FULL_IMAGE:figures/full_fig_p029_19.png] view at source ↗
Figure 20
Figure 20. Figure 20: Patient skill specification for mental-status-guided response generation and retrieval-based [PITH_FULL_IMAGE:figures/full_fig_p030_20.png] view at source ↗
Figure 21
Figure 21. Figure 21: Example of retrievable event memories for standardized patient simulation, covering [PITH_FULL_IMAGE:figures/full_fig_p031_21.png] view at source ↗
Figure 22
Figure 22. Figure 22: System prompt for auxiliary examination result retrieval in MentalHospital. [PITH_FULL_IMAGE:figures/full_fig_p032_22.png] view at source ↗
Figure 23
Figure 23. Figure 23: Control prompt for strict rubric-based scoring and standardized evaluation output. [PITH_FULL_IMAGE:figures/full_fig_p033_23.png] view at source ↗
Figure 24
Figure 24. Figure 24: Evaluation prompt for communication empathy, illustrating the task definition, scoring [PITH_FULL_IMAGE:figures/full_fig_p034_24.png] view at source ↗
Figure 25
Figure 25. Figure 25: Evaluation prompt for interviewing professionalism, defining level-specific criteria and [PITH_FULL_IMAGE:figures/full_fig_p035_25.png] view at source ↗
Figure 26
Figure 26. Figure 26: Evaluation prompt for clinical note quality, outlining ceiling rules and level-specific [PITH_FULL_IMAGE:figures/full_fig_p036_26.png] view at source ↗
Figure 27
Figure 27. Figure 27: Evaluation prompt for diagnostic rigor, specifying scoring restrictions and level-specific [PITH_FULL_IMAGE:figures/full_fig_p037_27.png] view at source ↗
Figure 28
Figure 28. Figure 28: Evaluation prompt for treatment appropriateness, defining level-specific criteria and [PITH_FULL_IMAGE:figures/full_fig_p038_28.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

102 extracted references · 102 canonical work pages · 22 internal anchors

  1. [1]

    Ellen E Lee, John Torous, Munmun De Choudhury, Colin A Depp, Sarah A Graham, Ho-Cheol Kim, Martin P Paulus, John H Krystal, and Dilip V Jeste. Artificial intelligence for mental health care: clinical applications, barriers, facilitators, and artificial wisdom.Biological Psychiatry: Cognitive Neuroscience and Neuroimaging, 6(9):856–864, 2021

  2. [2]

    Use of generative artificial intelligence (ai) in psychiatry and mental health care: a systematic review

    Sara Kolding, Robert M Lundin, Lasse Hansen, and Søren Dinesen Østergaard. Use of generative artificial intelligence (ai) in psychiatry and mental health care: a systematic review. Acta Neuropsychiatrica, 37:e37, 2025

  3. [3]

    HealthBench: Evaluating Large Language Models Towards Improved Human Health

    Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero- Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al. Healthbench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025

  4. [4]

    Cpsycoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling

    Chenhao Zhang, Renhao Li, Minghuan Tan, Min Yang, Jingwei Zhu, Di Yang, Jiahao Zhao, Guancheng Ye, Chengming Li, and Xiping Hu. Cpsycoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling. InFindings of the Association for Computational Linguistics: ACL 2024, pages 13947–13966, 2024

  5. [5]

    Omind: Framework for knowledge grounded finetuning and multi-turn dialogue benchmark for mental health llms

    Suraj Racha, Prashant Harish Joshi, Utkarsh Maurya, Nitin Yadav, Mridul Sharma, Ananya Kunisetty, Saranya Darisipudi, Nirmal Punjabi, and Ganesh Ramakrishnan. Omind: Framework for knowledge grounded finetuning and multi-turn dialogue benchmark for mental health llms. arXiv preprint arXiv:2603.25105, 2026

  6. [6]

    Interactive evaluation for medical llms via task-oriented dialogue system

    Ruoyu Liu, Kui Xue, Xiaofan Zhang, and Shaoting Zhang. Interactive evaluation for medical llms via task-oriented dialogue system. InProceedings of the 31st International Conference on Computational Linguistics, pages 4871–4896, 2025

  7. [7]

    Rashmi Patel, Soon Nan Wee, Rajagopalan Ramaswamy, Simran Thadani, Guruprabha Gu- ruswamy, Ruchir Garg, Nathan Calvanese, Matthew Valko, A Rush, M Rentería, et al. Neuroblu: A natural language processing (nlp) electronic health record (ehr) data analytic tool to generate real-world evidence in mental healthcare.European Psychiatry, 65(S1):S99–S100, 2022

  8. [8]

    Mentalseek- dx: Towards progressive hypothetico-deductive reasoning for real-world psychiatric diagnosis

    Xiao Sun, Yuming Yang, Junnan Zhu, Jiang Zhong, Xinyu Zhou, and Kaiwen Wei. Mentalseek- dx: Towards progressive hypothetico-deductive reasoning for real-world psychiatric diagnosis. arXiv preprint arXiv:2602.03340, 2026

  9. [9]

    Cbt-bench: Evaluating large language models on assisting cognitive behavior therapy

    Mian Zhang, Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C Chiu, Shaun M Eack, Fei Fang, William Yang Wang, and Zhiyu Chen. Cbt-bench: Evaluating large language models on assisting cognitive behavior therapy. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Tech...

  10. [10]

    The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025

    Monica Agrawal, Irene Y Chen, Freya Gulamali, and Shalmali Joshi. The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025

  11. [11]

    Counselbench: a large-scale expert evaluation and adversarial benchmark of large language models in mental health counseling.arXiv e-prints, pages arXiv–2506, 2025

    Yahan Li, Jifan Yao, John Bosco S Bunyi, Adam C Frank, Angel Hwang, and Ruishan Liu. Counselbench: a large-scale expert evaluation and adversarial benchmark of large language models in mental health counseling.arXiv e-prints, pages arXiv–2506, 2025

  12. [12]

    AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

    Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. arXiv preprint arXiv:2405.07960, 2024

  13. [13]

    Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator

    Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator. InProceedings of the 31st International Conference on Computational Linguistics, pages 10183–10213, 2025

  14. [14]

    OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

    Peigen Liu, Rui Ding, Yuren Mao, Ziyan Jiang, Yuxiang Ye, Yunjun Gao, Ying Zhang, Renjie Sun, Longbin Lai, and Zhengping Qian. Openhospital: A thing-in-itself arena for evolving and benchmarking llm-based collective intelligence.arXiv preprint arXiv:2603.14771, 2026. 10

  15. [15]

    Medagentbench: a virtual ehr environment to benchmark medical llm agents

    Yixing Jiang, Kameron C Black, Gloria Geng, Danny Park, James Zou, Andrew Y Ng, and Jonathan H Chen. Medagentbench: a virtual ehr environment to benchmark medical llm agents. Nejm Ai, 2(9):AIdbp2500144, 2025

  16. [16]

    PhD thesis, Royal College of Surgeons in Ireland, 2015

    Joseph Donohoe.Implementing an education programme and SOAP notes framework to improve nursing documentation. PhD thesis, Royal College of Surgeons in Ireland, 2015

  17. [17]

    Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026

    Lecheng Gong, Weimin Fang, Ting Yang, Dongjie Tao, Chunxiao Guo, Peng Wei, Bo Xie, Jinqun Guan, Zixiao Chen, Fang Shi, et al. Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026

  18. [18]

    LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis

    Shihao Xu, Tiancheng Zhou, Jiatong Ma, Yanli Ding, Yiming Yan, Ming Xiao, Guoyi Li, Haiyang Geng, Yunyun Han, Jianhua Chen, et al. Lingxidiagbench: A multi-agent framework for benchmarking llms in chinese psychiatric consultation and diagnosis.arXiv preprint arXiv:2602.09379, 2026

  19. [19]

    PSYCHE: A Multi-faceted Patient Simulation Framework for Evaluation of Psychiatric Assessment Conversational Agents

    Jingoo Lee, Kyungho Lim, Young-Chul Jung, and Byung-Hoon Kim. Psyche: A multi-faceted patient simulation framework for evaluation of psychiatric assessment conversational agents. arXiv preprint arXiv:2501.01594, 2025

  20. [20]

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

    Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models.arXiv preprint arXiv:2512.02556, 2025

  21. [21]

    Kimi K2: Open Agentic Intelligence

    Kimi Team, Yifan Bai, Yiping Bao, Y Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, et al. Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025

  22. [22]

    OpenAI GPT-5 System Card

    Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card.arXiv preprint arXiv:2601.03267, 2025

  23. [23]

    Constitutional AI: Harmlessness from AI Feedback

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073, 2022

  24. [24]

    Gemma: Open Models Based on Gemini Research and Technology

    Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology.arXiv preprint arXiv:2403.08295, 2024

  25. [25]

    The Llama 3 Herd of Models

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

  26. [26]

    Qwen3 Technical Report

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025

  27. [27]

    gpt-oss-120b & gpt-oss-20b Model Card

    Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card.arXiv preprint arXiv:2508.10925, 2025

  28. [28]

    MedGemma Technical Report

    Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, et al. Medgemma technical report.arXiv preprint arXiv:2507.05201, 2025

  29. [29]

    Baichuan-M2: Scaling Medical Capability with Large Verifier System

    Chengfeng Dou, Chong Liu, Fan Yang, Fei Li, Jiyuan Jia, Mingyang Chen, Qiang Ju, Shuai Wang, Shunya Dang, Tianpeng Li, et al. Baichuan-m2: Scaling medical capability with large verifier system.arXiv preprint arXiv:2509.02208, 2025

  30. [30]

    Toward expert-level medical question answering with large language models.Nature medicine, 31(3):943–950, 2025

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. Toward expert-level medical question answering with large language models.Nature medicine, 31(3):943–950, 2025. 11

  31. [31]

    Towards generalist biomedical ai.Nejm Ai, 1(3):AIoa2300138, 2024

    Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. Towards generalist biomedical ai.Nejm Ai, 1(3):AIoa2300138, 2024

  32. [32]

    Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024

    Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, et al. Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024

  33. [33]

    MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

    Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025

  34. [34]

    LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation

    Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha, Binze Hu, Yuhui Wang, Chenhao Huang, Shichun Liu, Jingqi Tong, Changhao Jiang, et al. Llmeval-med: a real-world clinical benchmark for medical llms with physician validation.arXiv preprint arXiv:2506.04078, 2025

  35. [35]

    A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains.npj Digital Medicine, 2025

    Shirui Wang, Zhihui Tang, Huaxia Yang, Qiuhong Gong, Tiantian Gu, Hongyang Ma, Yongxin Wang, Wubin Sun, Zeliang Lian, Kehang Mao, et al. A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains.npj Digital Medicine, 2025

  36. [36]

    ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

    Nikita Mehandru, Niloufar Golchini, David Bamman, Travis Zack, Melanie F Molina, and Ahmed Alaa. Er-reason: A benchmark dataset for llm-based clinical reasoning in the emergency room.arXiv preprint arXiv:2505.22919, 2025

  37. [37]

    Livemedbench: A contamination-free medical benchmark for llms with automated rubric evaluation.arXiv preprint arXiv:2602.10367, 2026

    Zhiling Yan, Dingjie Song, Zhe Fang, Yisheng Ji, Xiang Li, Quanzheng Li, and Lichao Sun. Livemedbench: A contamination-free medical benchmark for llms with automated rubric evaluation.arXiv preprint arXiv:2602.10367, 2026

  38. [38]

    ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

    Xiang Zheng, Han Li, Wenjie Luo, Weiqi Zhai, Yiyuan Li, Chuanmiao Yan, Tianyi Tang, Yubo Ma, Kexin Yang, Dayiheng Liu, et al. Clinconsensus: A consensus-based benchmark for evaluating chinese medical llms across difficulty levels.arXiv preprint arXiv:2603.02097, 2026

  39. [39]

    Mentalchat16k: A benchmark dataset for conversational mental health assistance

    Jia Xu, Tianyi Wei, Bojian Hou, Patryk Orzechowski, Shu Yang, Ruochen Jin, Rachael Paulbeck, Joost Wagenaar, George Demiris, and Li Shen. Mentalchat16k: A benchmark dataset for conversational mental health assistance. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 2, pages 5367–5378, 2025

  40. [40]

    PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice

    Shuyu Liu, Ruoxi Wang, Ling Zhang, Xuequan Zhu, Rui Yang, Xinzhu Zhou, Fei Wu, Zhi Yang, Cheng Jin, and Gang Wang. Psychbench: A comprehensive and professional benchmark for evaluating the performance of llm-assisted psychiatric clinical practice.arXiv preprint arXiv:2503.01903, 2025

  41. [41]

    Psychia- trybench: A multi-task benchmark for llms in psychiatry.arXiv preprint arXiv:2509.09711, 2025

    Aya E Fouda, Abdelrahamn A Hassan, Radwa J Hanafy, and Mohammed E Fouda. Psychia- trybench: A multi-task benchmark for llms in psychiatry.arXiv preprint arXiv:2509.09711, 2025

  42. [42]

    MentalBench: A DSM-Grounded Benchmark for Evaluating Psychiatric Diagnostic Capability of Large Language Models

    Hoyun Song, Migyeong Kang, Jisu Shin, Jihyun Kim, Chanbi Park, Hangyeol Yoo, Jihyun An, Alice Oh, Jinyoung Han, and KyungTae Lim. Mentalbench: A benchmark for evaluating psychiatric diagnostic capability of large language models.arXiv preprint arXiv:2602.12871, 2026

  43. [43]

    Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning.Advances in Neural Information Processing Systems, 37:28858–28888, 2024

    Shuyue S Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S Ilgen, Emma Pierson, Pang W Koh, and Yulia Tsvetkov. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning.Advances in Neural Information Processing Systems, 37:28858–28888, 2024

  44. [44]

    Dischar- gesim: A simulation benchmark for educational doctor–patient communication at discharge

    Zonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon, Soie Kwon, and Hong Yu. Dischar- gesim: A simulation benchmark for educational doctor–patient communication at discharge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 10783–10809, 2025. 12

  45. [45]

    Mindeval: Benchmarking language models on multi-turn mental health support

    José Pombal, Maya D’Eon, Nuno M Guerreiro, Pedro Henrique Martins, António Farinhas, and Ricardo Rei. Mindeval: Benchmarking language models on multi-turn mental health support. arXiv preprint arXiv:2511.18491, 2025

  46. [46]

    3mdbench: Medical multimodal multi-agent dialogue benchmark

    Ivan Sviridov, Amina Miftakhova, Tereshchenko Artemiy Vladimirovich, Galina Zubkova, Pavel Blinov, and Andrey Savchenko. 3mdbench: Medical multimodal multi-agent dialogue benchmark. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26625–26665, 2025

  47. [47]

    Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025

    Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025

  48. [48]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002

  49. [49]

    Rouge: A package for automatic evaluation of summaries

    Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. InText summarization branches out, pages 74–81, 2004

  50. [50]

    Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

    Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. InProceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005

  51. [51]

    BERTScore: Evaluating Text Generation with BERT

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019

  52. [52]

    Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

  53. [53]

    Clinical large language model evaluation by expert review (clever): Framework development and validation

    Veysel Kocaman, Mustafa Aytu˘g Kaya, Andrei Marian Feier, and David Talby. Clinical large language model evaluation by expert review (clever): Framework development and validation. JMIR AI, 4(1):e72153, 2025

  54. [54]

    Automating expert-level medical reasoning evaluation of large language models.npj Digital Medicine, 2025

    Shuang Zhou, Wenya Xie, Jiaxi Li, Zaifu Zhan, Meijia Song, Han Yang, Cheyenna Espinoza, Lindsay Welton, Xinnie Mai, Yanwei Jin, et al. Automating expert-level medical reasoning evaluation of large language models.npj Digital Medicine, 2025

  55. [55]

    Evaluating clinical ai summaries with large language models as judges.npj Digital Medicine, 8(1):640, 2025

    Emma Croxford, Yanjun Gao, Elliot First, Nicholas Pellegrino, Miranda Schnier, John Caskey, Madeline Oguss, Graham Wills, Guanhua Chen, Dmitriy Dligach, et al. Evaluating clinical ai summaries with large language models as judges.npj Digital Medicine, 8(1):640, 2025

  56. [56]

    Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists

    Yukyung Lee, Joonghoon Kim, Jaehee Kim, Hyowon Cho, Jaewook Kang, Pilsung Kang, and Najoung Kim. Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15782–15809, 2025

  57. [57]

    Who judges the judge? evaluating llm-as-a-judge for french medical open-ended qa

    Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils, Richard Dufour, and Benoit Favre. Who judges the judge? evaluating llm-as-a-judge for french medical open-ended qa. InProceedings of the 1st Workshop on Linguistic Analysis for Health (HeaLing 2026), pages 142–157, 2026

  58. [58]

    Human evaluators vs

    Gwydion Williams, Samuel Rutunda, Floris Nzabakira, and Bilal A Mateen. Human evaluators vs. llm-as-a-judge: Toward scalable, real-time evaluation of genai in global health.medRxiv, pages 2025–10, 2025

  59. [59]

    6A20 Schizophrenia

    Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge.The Innovation, 2024. 13 A Limitations MentalHospital currently focuses on text-based psychiatric interaction. Although we explored generative digital humans for richer patient presentation, current ...

  60. [60]

    Do not add, infer, or speculate

    Strictly follow the original text. Do not add, infer, or speculate

  61. [61]

    Retain negative information

  62. [62]

    Keep the wording as close to the original text as possible, with only necessary standardization

  63. [63]

    Remove duplicate array items, but do not merge independent factual units

  64. [64]

    Missing fields must be output as empty arrays, empty objects, or empty strings

  65. [65]

    Do not output explanations or Markdown

    Output only valid JSON. Do not output explanations or Markdown

  66. [66]

    id": "1000769

    Do not output“‘json,“‘, or any other code-fence markers. Input: <INPUT_JSON> Figure 9: System prompt for structuring psychiatric clinical records into standardized JSON format. Standardized De-identification Pipeline.Each raw EHR was converted into a benchmark case through a standardized on-site processing pipeline. All automated processing was conducted ...

  67. [67]

    You should independently generate the auxiliary examinations that need to be requested based on the current case information

  68. [68]

    If you determine that no auxiliary examination is currently necessary, output an empty array:[]

  69. [69]

    Each item you output must be a specific, standardized, and clearly defined examina- tion name, preferably using commonly accepted clinical medical terminology

  70. [70]

    Do not output overly broad, vague, or non-actionable category names, such as imaging examination,” laboratory examination,” or blood test.”

    The examination items you output must be sufficiently specific so that each item can be clearly mapped to an actual clinical examination. Do not output overly broad, vague, or non-actionable category names, such as imaging examination,” laboratory examination,” or blood test.”

  71. [71]

    Only output auxiliary examinations that are genuinely necessary for the current diagnostic process, and avoid irrelevant, redundant, or clearly duplicated items

  72. [72]

    The final result must be enclosed by [BEGIN_EXAMINATIONS] and [END_EXAMINATIONS]

  73. [73]

    Complete blood count

    The content between[BEGIN_EXAMINATIONS] and[END_EXAMINATIONS] must be, and must only be, a JSON array. Do not add any explanations, comments, or other content. Example output: [BEGIN_EXAMINATIONS] [ "Complete blood count", "Thyroid function tests", "Brain MRI", "Electroencephalography" ] [END_EXAMINATIONS] If no auxiliary examination is currently required...

  74. [74]

    Your output must be the main body of a clinical record

  75. [75]

    Do not output JSON, Markdown, explanatory notes, or any other additional content

    The clinical record must be plain text. Do not output JSON, Markdown, explanatory notes, or any other additional content

  76. [76]

    You must enclose the entire clinical record using the following markers: [BEGIN_CLINICAL_FORMULATION] Main body of the clinical record [END_CLINICAL_FORMULATION]

  77. [77]

    6A20 Schizophrenia

    The clinical record should be concise, clear, and complete. Avoid verbose descrip- tions and irrelevant elaboration while preserving all key information. After you output [END_CLINICAL_FORMULATION], the workflow will proceed to the next stage. Figure 15: System prompt for clinical record generation in the psychiatric clinical workflow. descriptions are no...

  78. [78]

    You must first provide the diagnostic rationale and supporting evidence to explain your diagnostic decision-making process

  79. [79]

    After completing the diagnostic rationale, you must output the final diagnostic results in list format

  80. [80]

    You may select only one or more diagnoses from the candidate diagnosis list above, including comorbid diagnoses where applicable

Showing first 80 references.