REVIEW 3 major objections 6 minor 102 references
Even the strongest large language models still trail clinicians by 37 percentage points on complete psychiatric encounters, with mental-status assessment as the main bottleneck.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 10:29 UTC pith:V4B3HWYH
load-bearing objection Solid full-S.O.A.P. psychiatric benchmark with real multi-center EHRs and specialist-aligned judges; the 37-point gap is real enough to cite, with a modest metric-calibration caveat on mental-status coverage. the 3 major comments →
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Full-process psychiatric competence is still far from clinician level: the strongest medical-specific model trails medical trainees by about 27 points and human experts by about 37 points on average objective metrics across interviewing, examination, notes, category diagnosis, and disorder diagnosis, with mental-status coverage as a recurring failure mode. Subjective process scores are mixed—models can match or exceed humans on expressed empathy and treatment appropriateness while remaining weaker on interviewing professionalism and note quality—so models are not substitutes, but may complement clinicians in affective support and treatment assistance.
What carries the argument
MentalHospital: an EHR-grounded S.O.A.P. simulation using skill-augmented standardized patients (role plus presentation skill plus topic-level memory skill) from 1,193 de-identified cases covering all major ICD-11 psychiatric categories and 76 disorders, paired with a dual-track evaluation protocol and MentalEval—five Qwen3-8B evaluators trained by rubric-grounded supervised fine-tuning then expert-guided preference optimization—to scale specialist judgment of process quality.
Load-bearing premise
That the de-identified EHR checkpoints and the skill-augmented patients form a faithful, non-leaking gold standard for what a doctor should recover and how a real psychiatric patient would present.
What would settle it
Re-run the same doctor agents on a held-out multi-center set with independent psychiatrist re-annotation of checkpoints and live standardized-patient sessions; if the LLM–clinician objective gap shrinks below roughly ten points or mental-status coverage ceases to be the dominant miss, the measured competence gap is an artifact of the current patient construction.
If this is right
- Psychiatric AI evaluation must move from single-turn or dialogue-only tests to full interview–exam–note–diagnosis–treatment episodes with EHR-grounded references.
- Mental-status examination becomes a primary training and evaluation target, not an optional side skill.
- Specialist-aligned judges (rubric SFT + expert preference) can replace general LLM-as-judge for scalable process scoring in psychiatry.
- LLMs may be useful as empathic communication and treatment-drafting assistants while remaining unsuitable as autonomous clinicians under this protocol.
- Controlled access to de-identified EHR-derived cases plus public environment code can become a standard for psychiatric agent benchmarks.
Where Pith is reading between the lines
- If mental-status probing is the bottleneck, interview curricula that force explicit MSE checklists before diagnosis may close more of the gap than larger generic models alone.
- The dual-track split implies future model cards should report objective evidence recovery and process quality separately rather than a single medical accuracy score.
- Because comorbid and multi-center cases are already in the bank, the same scaffold could stress-test safety behaviors (self-harm, psychosis reinforcement) without inventing synthetic patients from scratch.
- Clinician survey positivity for training suitability suggests the environment may first land as a trainee simulator even if model scores remain low.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MentalHospital, an EHR-grounded virtual environment for evaluating LLM psychiatric clinical encounters under a full S.O.A.P. workflow, built from 1,193 de-identified multi-center cases spanning all major ICD-11 categories and 76 disorders. Skill-augmented standardized patients (representation + memory skills) and a hospital examination module force agents to elicit evidence rather than read the chart. Evaluation is dual-track: objective coverage against EHR-derived checkpoints and subjective process quality via MentalEval, five Qwen3-8B evaluators trained with rubric-grounded SFT then expert-guided DPO. A 22-clinician survey rates clinical fidelity at 3.88/5; MentalEval reaches average QWK 0.944 on held-out expert labels. Benchmarking of 12 LLMs against experts and trainees reports that the strongest model trails clinicians by 37.28 percentage points on objective metrics, with mental-status assessment as the principal bottleneck, while LLMs show complementary strengths in expressed empathy and treatment appropriateness.
Significance. If the environment and dual-track protocol hold under scrutiny, this is a substantial contribution to medical AI evaluation: it moves psychiatric LLM assessment from isolated dialogue/diagnosis tasks to complete, EHR-anchored encounters with both outcome and process measures. Strengths include multi-center real EHR grounding, explicit patient-construction ablations (Table 4), specialist-aligned evaluators with strong held-out QWK, multi-group human baselines (experts, trainees, crowdworkers), and a clear, falsifiable bottleneck claim on mental-status recovery. The resource-release protocol (code, rubrics, MentalEval weights public; raw EHRs controlled) is a responsible compromise. These elements make the work useful for training, benchmarking, and diagnosing where current LLMs fail in psychiatry, independent of any single headline number.
major comments (3)
- [§4, Eq. (13), Tables 2–3] §4, Eq. (13) and Tables 2–3: the headline 37.28 pp LLM–clinician gap is defined by Coverage(yc, Kc) with a hand-chosen semantic threshold τ=0.85 plus Multi-LLM adjudication for pairs below threshold. Clinicians and LLMs interact with the same memory-gated patients, but the paper does not report a calibration study of human vs. LLM utterance matching on identical elicited content (inter-rater agreement on k ⪯ yc, or re-scoring of clinician transcripts under the same automatic matcher). Table 3’s large CC–MS gaps for LLMs could therefore partly reflect phrasing/style sensitivity of the matcher rather than pure competence. A load-bearing fix is to (i) report human–LLM matching agreement on a shared probe set, (ii) ablate τ, and (iii) recompute the gap under exact/grounding-based coverage where available (Appendix K already logs patient grounding fields).
- [§2.2, Table 4, Appendix L] §2.2, Table 4, Appendix L: the claim that skill-augmented patients constitute a faithful, non-leaking gold standard rests on a small objective probe set (20 cases × 12 probes = 240 responses) and three-clinician subjective ratings. There is no quantitative check that de-identification (Appendix F) preserved psychopathological logic at the checkpoint level, nor a leakage audit showing that patients never disclose unasked future-stage evidence under adversarial doctor prompts. Because both objective coverage and the LLM–clinician gap are measured against these patients, expand the fidelity study (more cases, adversarial probes, inter-psychiatrist agreement on whether disclosed content matches the original EHR logic) or qualify the gap as conditional on the current patient construction.
- [§3.1, Table 5] §3.1 and Table 5: MentalEval’s cold-start SFT is supervised by a five-LLM judge ensemble that the paper itself reports as poorly aligned with specialists (LLM-as-a-Judge QWK 0.677, Acc. 0.225). Although expert-guided DPO on low-confidence sets raises average QWK to 0.944 on held-out cases, residual dependence on weak LLM judges for the bulk of SFT targets is a load-bearing design choice for scalable subjective scores. Report (a) the fraction of SFT data that survived consensus filtering vs. was rewritten/augmented, (b) agreement of SFT-only vs. SFT+DPO evaluators stratified by score extremity, and (c) whether clinician DPO preferences were collected independently of the models being ranked in Table 2.
minor comments (6)
- [Abstract / §1 / §4.5] Abstract and §4.5 report clinician fidelity as 3.88/5 while the introduction states 3.96/5; reconcile the two figures and state which sample (experts only vs. experts+trainees) each uses.
- [Table 2, §4.1] Table 2 lists Empathy scores where lower appears better for humans (experts 1.22) but higher for some LLMs; clarify whether the empathy rubric is inverted relative to other 1–5 dimensions or whether experts deliberately suppress affective language.
- [§2, Eq. (1)–(4)] Eq. (1) uses xc = {Kpat_c, Kexam_c} and Y*_c; later ˆyc uses different symbols for the same conceptual objects. A short notation table in §2 would reduce reader load.
- [Figure 3] Figure 3 confusion matrices are hard to read in grayscale; add numeric cell annotations and a shared color scale.
- [§6 / Appendix A] Appendix A Limitations correctly flags missing safety/adversarial evaluation (self-harm, psychosis reinforcement); a short pointer in the main-text conclusion would set expectations for deployment claims.
- [Table 2, §4] Several model names appear inconsistently (Deepseek-v4-Pro vs. DeepSeek-V4-Pro; Claude-Sonnet-4.6 vs. Claude-Sonnet-4-6). Normalize throughout.
Circularity Check
No load-bearing circularity: the 37.28 pp LLM–clinician gap is measured by independent EHR-checkpoint coverage, not by construction from the evaluators or fitted parameters.
full rationale
MentalHospital’s central empirical claim (Table 2, §4.1–4.2) is an objective competence gap obtained by running LLMs and human clinicians through the same S.O.A.P. episodes and scoring Coverage(yc, Kc) against EHR-derived checkpoints (Eq. 13, Appendix K). Those checkpoints are extracted from de-identified multi-center EHRs that pre-exist the benchmark; they are not fitted to model outputs, nor are they defined in terms of the MentalEval scores. MentalEval itself is trained on multi-LLM trajectories plus expert DPO preferences on held-out cases and is used only for the subjective track; the headline 37.28-point figure does not depend on it. Patient construction (representation + memory skills) is ablated for fidelity (Table 4) but does not redefine the reference targets. There are no self-definitional equations, no fitted parameters re-labeled as predictions, no uniqueness theorems imported from overlapping authors, and no ansatz smuggled via self-citation that forces the reported gap. Minor self-reference exists only in the ordinary sense that the authors built the environment they evaluate; that does not reduce the measured gap to an input by construction. The derivation is therefore self-contained against external EHR gold standards.
Axiom & Free-Parameter Ledger
free parameters (3)
- semantic similarity threshold τ =
0.85
- judge consensus k
- DPO β =
0.3
axioms (3)
- domain assumption S.O.A.P. workflow is the appropriate complete clinical encounter structure for psychiatric evaluation
- domain assumption De-identified EHR checkpoints remain clinically complete and diagnostically valid gold standards after privacy rewriting
- ad hoc to paper Rubric-grounded SFT + expert DPO produces evaluators whose 1–5 scores can stand in for specialist judgment at scale
invented entities (2)
-
MentalHospital environment (skill-augmented standardized patients + examination module)
no independent evidence
-
MentalEval (five Qwen3-8B domain-specific evaluators)
no independent evidence
read the original abstract
Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce $\textbf{MentalHospital}$, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop $\textbf{MentalEval}$, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.
Figures
Reference graph
Works this paper leans on
-
[1]
Ellen E Lee, John Torous, Munmun De Choudhury, Colin A Depp, Sarah A Graham, Ho-Cheol Kim, Martin P Paulus, John H Krystal, and Dilip V Jeste. Artificial intelligence for mental health care: clinical applications, barriers, facilitators, and artificial wisdom.Biological Psychiatry: Cognitive Neuroscience and Neuroimaging, 6(9):856–864, 2021
work page 2021
-
[2]
Sara Kolding, Robert M Lundin, Lasse Hansen, and Søren Dinesen Østergaard. Use of generative artificial intelligence (ai) in psychiatry and mental health care: a systematic review. Acta Neuropsychiatrica, 37:e37, 2025
work page 2025
-
[3]
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Rahul K Arora, Jason Wei, Rebecca Soskin Hicks, Preston Bowman, Joaquin Quiñonero- Candela, Foivos Tsimpourlas, Michael Sharman, Meghan Shah, Andrea Vallone, Alex Beutel, et al. Healthbench: Evaluating large language models towards improved human health.arXiv preprint arXiv:2505.08775, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[4]
Chenhao Zhang, Renhao Li, Minghuan Tan, Min Yang, Jingwei Zhu, Di Yang, Jiahao Zhao, Guancheng Ye, Chengming Li, and Xiping Hu. Cpsycoun: A report-based multi-turn dialogue reconstruction and evaluation framework for chinese psychological counseling. InFindings of the Association for Computational Linguistics: ACL 2024, pages 13947–13966, 2024
work page 2024
-
[5]
Suraj Racha, Prashant Harish Joshi, Utkarsh Maurya, Nitin Yadav, Mridul Sharma, Ananya Kunisetty, Saranya Darisipudi, Nirmal Punjabi, and Ganesh Ramakrishnan. Omind: Framework for knowledge grounded finetuning and multi-turn dialogue benchmark for mental health llms. arXiv preprint arXiv:2603.25105, 2026
-
[6]
Interactive evaluation for medical llms via task-oriented dialogue system
Ruoyu Liu, Kui Xue, Xiaofan Zhang, and Shaoting Zhang. Interactive evaluation for medical llms via task-oriented dialogue system. InProceedings of the 31st International Conference on Computational Linguistics, pages 4871–4896, 2025
work page 2025
-
[7]
Rashmi Patel, Soon Nan Wee, Rajagopalan Ramaswamy, Simran Thadani, Guruprabha Gu- ruswamy, Ruchir Garg, Nathan Calvanese, Matthew Valko, A Rush, M Rentería, et al. Neuroblu: A natural language processing (nlp) electronic health record (ehr) data analytic tool to generate real-world evidence in mental healthcare.European Psychiatry, 65(S1):S99–S100, 2022
work page 2022
-
[8]
Xiao Sun, Yuming Yang, Junnan Zhu, Jiang Zhong, Xinyu Zhou, and Kaiwen Wei. Mentalseek- dx: Towards progressive hypothetico-deductive reasoning for real-world psychiatric diagnosis. arXiv preprint arXiv:2602.03340, 2026
-
[9]
Cbt-bench: Evaluating large language models on assisting cognitive behavior therapy
Mian Zhang, Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C Chiu, Shaun M Eack, Fei Fang, William Yang Wang, and Zhiyu Chen. Cbt-bench: Evaluating large language models on assisting cognitive behavior therapy. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Tech...
work page 2025
-
[10]
The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025
Monica Agrawal, Irene Y Chen, Freya Gulamali, and Shalmali Joshi. The evaluation illusion of large language models in medicine.npj Digital Medicine, 8(1):600, 2025
work page 2025
-
[11]
Yahan Li, Jifan Yao, John Bosco S Bunyi, Adam C Frank, Angel Hwang, and Ruishan Liu. Counselbench: a large-scale expert evaluation and adversarial benchmark of large language models in mental health counseling.arXiv e-prints, pages arXiv–2506, 2025
work page 2025
-
[12]
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor. Agentclinic: a multimodal agent benchmark to evaluate ai in simulated clinical environments. arXiv preprint arXiv:2405.07960, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[13]
Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator
Zhihao Fan, Lai Wei, Jialong Tang, Wei Chen, Wang Siyuan, Zhongyu Wei, and Fei Huang. Ai hospital: Benchmarking large language models in a multi-agent medical interaction simulator. InProceedings of the 31st International Conference on Computational Linguistics, pages 10183–10213, 2025
work page 2025
-
[14]
Peigen Liu, Rui Ding, Yuren Mao, Ziyan Jiang, Yuxiang Ye, Yunjun Gao, Ying Zhang, Renjie Sun, Longbin Lai, and Zhengping Qian. Openhospital: A thing-in-itself arena for evolving and benchmarking llm-based collective intelligence.arXiv preprint arXiv:2603.14771, 2026. 10
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[15]
Medagentbench: a virtual ehr environment to benchmark medical llm agents
Yixing Jiang, Kameron C Black, Gloria Geng, Danny Park, James Zou, Andrew Y Ng, and Jonathan H Chen. Medagentbench: a virtual ehr environment to benchmark medical llm agents. Nejm Ai, 2(9):AIdbp2500144, 2025
work page 2025
-
[16]
PhD thesis, Royal College of Surgeons in Ireland, 2015
Joseph Donohoe.Implementing an education programme and SOAP notes framework to improve nursing documentation. PhD thesis, Royal College of Surgeons in Ireland, 2015
work page 2015
-
[17]
Lecheng Gong, Weimin Fang, Ting Yang, Dongjie Tao, Chunxiao Guo, Peng Wei, Bo Xie, Jinqun Guan, Zixiao Chen, Fang Shi, et al. Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models.arXiv preprint arXiv:2601.03023, 2026
-
[18]
Shihao Xu, Tiancheng Zhou, Jiatong Ma, Yanli Ding, Yiming Yan, Ming Xiao, Guoyi Li, Haiyang Geng, Yunyun Han, Jianhua Chen, et al. Lingxidiagbench: A multi-agent framework for benchmarking llms in chinese psychiatric consultation and diagnosis.arXiv preprint arXiv:2602.09379, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[19]
Jingoo Lee, Kyungho Lim, Young-Chul Jung, and Byung-Hoon Kim. Psyche: A multi-faceted patient simulation framework for evaluation of psychiatric assessment conversational agents. arXiv preprint arXiv:2501.01594, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[20]
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, et al. Deepseek-v3. 2: Pushing the frontier of open large language models.arXiv preprint arXiv:2512.02556, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[21]
Kimi K2: Open Agentic Intelligence
Kimi Team, Yifan Bai, Yiping Bao, Y Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, et al. Kimi k2: Open agentic intelligence.arXiv preprint arXiv:2507.20534, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[22]
Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card.arXiv preprint arXiv:2601.03267, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[23]
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073, 2022
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[24]
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, et al. Gemma: Open models based on gemini research and technology.arXiv preprint arXiv:2403.08295, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[25]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[26]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[27]
gpt-oss-120b & gpt-oss-20b Model Card
Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card.arXiv preprint arXiv:2508.10925, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[28]
Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cían Hughes, Charles Lau, et al. Medgemma technical report.arXiv preprint arXiv:2507.05201, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[29]
Baichuan-M2: Scaling Medical Capability with Large Verifier System
Chengfeng Dou, Chong Liu, Fan Yang, Fei Li, Jiyuan Jia, Mingyang Chen, Qiang Ju, Shuai Wang, Shunya Dang, Tianpeng Li, et al. Baichuan-m2: Scaling medical capability with large verifier system.arXiv preprint arXiv:2509.02208, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[30]
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. Toward expert-level medical question answering with large language models.Nature medicine, 31(3):943–950, 2025. 11
work page 2025
-
[31]
Towards generalist biomedical ai.Nejm Ai, 1(3):AIoa2300138, 2024
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Charles Lau, Ryutaro Tanno, Ira Ktena, et al. Towards generalist biomedical ai.Nejm Ai, 1(3):AIoa2300138, 2024
work page 2024
-
[32]
Pengcheng Chen, Jin Ye, Guoan Wang, Yanjun Li, Zhongying Deng, Wei Li, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, et al. Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical ai.Advances in Neural Information Processing Systems, 37:94327–94427, 2024
work page 2024
-
[33]
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes, Alyssa Unell, Michael Wornow, Juan M Banda, Nikesh Kotecha, Timothy Keyes, Yifan Mai, Mert Oez, et al. Medhelm: Holistic evaluation of large language models for medical tasks.arXiv preprint arXiv:2505.23802, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[34]
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
Ming Zhang, Yujiong Shen, Zelin Li, Huayu Sha, Binze Hu, Yuhui Wang, Chenhao Huang, Shichun Liu, Jingqi Tong, Changhao Jiang, et al. Llmeval-med: a real-world clinical benchmark for medical llms with physician validation.arXiv preprint arXiv:2506.04078, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[35]
Shirui Wang, Zhihui Tang, Huaxia Yang, Qiuhong Gong, Tiantian Gu, Hongyang Ma, Yongxin Wang, Wubin Sun, Zeliang Lian, Kehang Mao, et al. A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains.npj Digital Medicine, 2025
work page 2025
-
[36]
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
Nikita Mehandru, Niloufar Golchini, David Bamman, Travis Zack, Melanie F Molina, and Ahmed Alaa. Er-reason: A benchmark dataset for llm-based clinical reasoning in the emergency room.arXiv preprint arXiv:2505.22919, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[37]
Zhiling Yan, Dingjie Song, Zhe Fang, Yisheng Ji, Xiang Li, Quanzheng Li, and Lichao Sun. Livemedbench: A contamination-free medical benchmark for llms with automated rubric evaluation.arXiv preprint arXiv:2602.10367, 2026
-
[38]
Xiang Zheng, Han Li, Wenjie Luo, Weiqi Zhai, Yiyuan Li, Chuanmiao Yan, Tianyi Tang, Yubo Ma, Kexin Yang, Dayiheng Liu, et al. Clinconsensus: A consensus-based benchmark for evaluating chinese medical llms across difficulty levels.arXiv preprint arXiv:2603.02097, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[39]
Mentalchat16k: A benchmark dataset for conversational mental health assistance
Jia Xu, Tianyi Wei, Bojian Hou, Patryk Orzechowski, Shu Yang, Ruochen Jin, Rachael Paulbeck, Joost Wagenaar, George Demiris, and Li Shen. Mentalchat16k: A benchmark dataset for conversational mental health assistance. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V . 2, pages 5367–5378, 2025
work page 2025
-
[40]
Shuyu Liu, Ruoxi Wang, Ling Zhang, Xuequan Zhu, Rui Yang, Xinzhu Zhou, Fei Wu, Zhi Yang, Cheng Jin, and Gang Wang. Psychbench: A comprehensive and professional benchmark for evaluating the performance of llm-assisted psychiatric clinical practice.arXiv preprint arXiv:2503.01903, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[41]
Aya E Fouda, Abdelrahamn A Hassan, Radwa J Hanafy, and Mohammed E Fouda. Psychia- trybench: A multi-task benchmark for llms in psychiatry.arXiv preprint arXiv:2509.09711, 2025
-
[42]
Hoyun Song, Migyeong Kang, Jisu Shin, Jihyun Kim, Chanbi Park, Hangyeol Yoo, Jihyun An, Alice Oh, Jinyoung Han, and KyungTae Lim. Mentalbench: A benchmark for evaluating psychiatric diagnostic capability of large language models.arXiv preprint arXiv:2602.12871, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[43]
Shuyue S Li, Vidhisha Balachandran, Shangbin Feng, Jonathan S Ilgen, Emma Pierson, Pang W Koh, and Yulia Tsvetkov. Mediq: Question-asking llms and a benchmark for reliable interactive clinical reasoning.Advances in Neural Information Processing Systems, 37:28858–28888, 2024
work page 2024
-
[44]
Dischar- gesim: A simulation benchmark for educational doctor–patient communication at discharge
Zonghai Yao, Michael Sun, Won Seok Jang, Sunjae Kwon, Soie Kwon, and Hong Yu. Dischar- gesim: A simulation benchmark for educational doctor–patient communication at discharge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 10783–10809, 2025. 12
work page 2025
-
[45]
Mindeval: Benchmarking language models on multi-turn mental health support
José Pombal, Maya D’Eon, Nuno M Guerreiro, Pedro Henrique Martins, António Farinhas, and Ricardo Rei. Mindeval: Benchmarking language models on multi-turn mental health support. arXiv preprint arXiv:2511.18491, 2025
-
[46]
3mdbench: Medical multimodal multi-agent dialogue benchmark
Ivan Sviridov, Amina Miftakhova, Tereshchenko Artemiy Vladimirovich, Galina Zubkova, Pavel Blinov, and Andrey Savchenko. 3mdbench: Medical multimodal multi-agent dialogue benchmark. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 26625–26665, 2025
work page 2025
-
[47]
Mohammad Almansoori, Komal Kumar, and Hisham Cholakkal. Self-evolving multi-agent simulations for realistic clinical interactions.arXiv preprint arXiv:2503.22678, 2025
-
[48]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th annual meeting of the Association for Computational Linguistics, pages 311–318, 2002
work page 2002
-
[49]
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. InText summarization branches out, pages 74–81, 2004
work page 2004
-
[50]
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. InProceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pages 65–72, 2005
work page 2005
-
[51]
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019
work page internal anchor Pith review Pith/arXiv arXiv 1904
-
[52]
Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023
work page 2023
-
[53]
Veysel Kocaman, Mustafa Aytu˘g Kaya, Andrei Marian Feier, and David Talby. Clinical large language model evaluation by expert review (clever): Framework development and validation. JMIR AI, 4(1):e72153, 2025
work page 2025
-
[54]
Shuang Zhou, Wenya Xie, Jiaxi Li, Zaifu Zhan, Meijia Song, Han Yang, Cheyenna Espinoza, Lindsay Welton, Xinnie Mai, Yanwei Jin, et al. Automating expert-level medical reasoning evaluation of large language models.npj Digital Medicine, 2025
work page 2025
-
[55]
Emma Croxford, Yanjun Gao, Elliot First, Nicholas Pellegrino, Miranda Schnier, John Caskey, Madeline Oguss, Graham Wills, Guanhua Chen, Dmitriy Dligach, et al. Evaluating clinical ai summaries with large language models as judges.npj Digital Medicine, 8(1):640, 2025
work page 2025
-
[56]
Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists
Yukyung Lee, Joonghoon Kim, Jaehee Kim, Hyowon Cho, Jaewook Kang, Pilsung Kang, and Najoung Kim. Checkeval: A reliable llm-as-a-judge framework for evaluating text generation using checklists. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15782–15809, 2025
work page 2025
-
[57]
Who judges the judge? evaluating llm-as-a-judge for french medical open-ended qa
Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils, Richard Dufour, and Benoit Favre. Who judges the judge? evaluating llm-as-a-judge for french medical open-ended qa. InProceedings of the 1st Workshop on Linguistic Analysis for Health (HeaLing 2026), pages 142–157, 2026
work page 2026
-
[58]
Gwydion Williams, Samuel Rutunda, Floris Nzabakira, and Bilal A Mateen. Human evaluators vs. llm-as-a-judge: Toward scalable, real-time evaluation of genai in global health.medRxiv, pages 2025–10, 2025
work page 2025
-
[59]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. A survey on llm-as-a-judge.The Innovation, 2024. 13 A Limitations MentalHospital currently focuses on text-based psychiatric interaction. Although we explored generative digital humans for richer patient presentation, current ...
work page 2024
-
[60]
Do not add, infer, or speculate
Strictly follow the original text. Do not add, infer, or speculate
-
[61]
Retain negative information
-
[62]
Keep the wording as close to the original text as possible, with only necessary standardization
-
[63]
Remove duplicate array items, but do not merge independent factual units
-
[64]
Missing fields must be output as empty arrays, empty objects, or empty strings
-
[65]
Do not output explanations or Markdown
Output only valid JSON. Do not output explanations or Markdown
-
[66]
Do not output“‘json,“‘, or any other code-fence markers. Input: <INPUT_JSON> Figure 9: System prompt for structuring psychiatric clinical records into standardized JSON format. Standardized De-identification Pipeline.Each raw EHR was converted into a benchmark case through a standardized on-site processing pipeline. All automated processing was conducted ...
-
[67]
You should independently generate the auxiliary examinations that need to be requested based on the current case information
-
[68]
If you determine that no auxiliary examination is currently necessary, output an empty array:[]
-
[69]
Each item you output must be a specific, standardized, and clearly defined examina- tion name, preferably using commonly accepted clinical medical terminology
-
[70]
The examination items you output must be sufficiently specific so that each item can be clearly mapped to an actual clinical examination. Do not output overly broad, vague, or non-actionable category names, such as imaging examination,” laboratory examination,” or blood test.”
-
[71]
Only output auxiliary examinations that are genuinely necessary for the current diagnostic process, and avoid irrelevant, redundant, or clearly duplicated items
-
[72]
The final result must be enclosed by [BEGIN_EXAMINATIONS] and [END_EXAMINATIONS]
-
[73]
The content between[BEGIN_EXAMINATIONS] and[END_EXAMINATIONS] must be, and must only be, a JSON array. Do not add any explanations, comments, or other content. Example output: [BEGIN_EXAMINATIONS] [ "Complete blood count", "Thyroid function tests", "Brain MRI", "Electroencephalography" ] [END_EXAMINATIONS] If no auxiliary examination is currently required...
-
[74]
Your output must be the main body of a clinical record
-
[75]
Do not output JSON, Markdown, explanatory notes, or any other additional content
The clinical record must be plain text. Do not output JSON, Markdown, explanatory notes, or any other additional content
-
[76]
You must enclose the entire clinical record using the following markers: [BEGIN_CLINICAL_FORMULATION] Main body of the clinical record [END_CLINICAL_FORMULATION]
-
[77]
The clinical record should be concise, clear, and complete. Avoid verbose descrip- tions and irrelevant elaboration while preserving all key information. After you output [END_CLINICAL_FORMULATION], the workflow will proceed to the next stage. Figure 15: System prompt for clinical record generation in the psychiatric clinical workflow. descriptions are no...
-
[78]
You must first provide the diagnostic rationale and supporting evidence to explain your diagnostic decision-making process
-
[79]
After completing the diagnostic rationale, you must output the final diagnostic results in list format
-
[80]
You may select only one or more diagnoses from the candidate diagnosis list above, including comorbid diagnoses where applicable
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.