{"id":"666477e2-7bcd-4b82-9517-7420a1eff54f","arxiv_id":"2607.08257","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"EHR-derived standardized patients and dual-track evaluation reveal LLMs trail clinicians by 37.28 points on full psychiatric encounters, with mental-status assessment the main bottleneck.","lead":"MentalHospital is an EHR-grounded virtual hospital that forces LLMs through full psychiatric S.O.A.P. encounters with skill-augmented standardized patients. It shows even the best models still lag clinicians by ~37 points on objective competence, especially mental-status assessment.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The 37.28-point LLM–clinician gap rests on a coverage metric that may systematically under-count LLM mental-status recovery relative to clinicians.","rationale":"The reader correctly flags patient faithfulness and de-identification as a soft spot, but the more load-bearing risk for the specific 37.28 pp claim is the objective coverage metric itself under skill-gated disclosure. Patient construction (Table 4) and clinician survey fidelity (3.88/5) already provide partial support that the environment is not grossly leaking or unfaithful; the remaining uncertainty is whether Coverage with τ=0.85 treats clinician and LLM language equivalently on mental-status items—the very component the paper identifies as the bottleneck. A human re-annotation of checkpoint presence on matched transcripts would settle whether the gap is competence or matching. That keeps the verdict CONDITIONAL (privacy + hand-set thresholds remain) without elevating correctness risk beyond the reader’s low assessment. No stronger internal inconsistency is present; the dual-track design and MentalEval QWK of 0.944 are solid supporting evidence.","tokens_in":38303,"tokens_out":598,"duration_ms":6640,"concrete_test":"On a stratified sample of ≥50 cases, have three psychiatrists independently mark which EHR mental-status checkpoints are present in (a) clinician transcripts and (b) the strongest LLM transcripts; recompute MS coverage with this human gold and with τ lowered to 0.70. If the LLM–clinician MS gap shrinks by >10 pp under either re-scoring, the 37.28 pp headline is partly metric-driven and should be qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline claim (Table 2, §4.1–4.2) is that even the strongest LLM trails clinicians by 37.28 pp in objective psychiatric competence, with mental-status assessment as the key bottleneck. Objective interviewing and note scores are Coverage(yc, Kc) = (1/|Kc|) Σ I[k ⪯ yc] with a conservative semantic threshold τ = 0.85 plus Multi-LLM adjudication for pairs below threshold (§4, Eq. 13; Appendix K). Clinicians interact with the same skill-augmented patients (representation + memory skills, §2.2, Table 4) whose disclosure is gated by topic-level memory and presentation rendering. If clinicians elicit more complete or more easily matched mental-status language while LLMs recover the same underlying facts in less checkpoint-aligned phrasing, the coverage gap is inflated by matching style rather than by true clinical competence. The paper’s own Table 3 shows large CC–MS gaps for LLMs, but does not report inter-rater agreement or human–LLM matching calibration on the same elicited utterances, so the 37.28 pp figure may partly measure metric sensitivity rather than pure competence difference.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces MentalHospital, an EHR-grounded virtual environment for evaluating LLM psychiatric clinical encounters under a full S.O.A.P. workflow, built from 1,193 de-identified multi-center cases spanning all major ICD-11 categories and 76 disorders. Skill-augmented standardized patients (representation + memory skills) and a hospital examination module force agents to elicit evidence rather than read the chart. Evaluation is dual-track: objective coverage against EHR-derived checkpoints and subjective process quality via MentalEval, five Qwen3-8B evaluators trained with rubric-grounded SFT then expert-guided DPO. A 22-clinician survey rates clinical fidelity at 3.88/5; MentalEval reaches average QWK 0.944 on held-out expert labels. Benchmarking of 12 LLMs against experts and trainees reports that the strongest model trails clinicians by 37.28 percentage points on objective metrics, with mental-status assessment as the principal bottleneck, while LLMs show complementary strengths in expressed empathy and treatment appropriateness.","tokens_in":38667,"tokens_out":1555,"duration_ms":23217,"significance":"If the environment and dual-track protocol hold under scrutiny, this is a substantial contribution to medical AI evaluation: it moves psychiatric LLM assessment from isolated dialogue/diagnosis tasks to complete, EHR-anchored encounters with both outcome and process measures. Strengths include multi-center real EHR grounding, explicit patient-construction ablations (Table 4), specialist-aligned evaluators with strong held-out QWK, multi-group human baselines (experts, trainees, crowdworkers), and a clear, falsifiable bottleneck claim on mental-status recovery. The resource-release protocol (code, rubrics, MentalEval weights public; raw EHRs controlled) is a responsible compromise. These elements make the work useful for training, benchmarking, and diagnosing where current LLMs fail in psychiatry, independent of any single headline number.","major_comments":[{"comment":"§4, Eq. (13) and Tables 2–3: the headline 37.28 pp LLM–clinician gap is defined by Coverage(yc, Kc) with a hand-chosen semantic threshold τ=0.85 plus Multi-LLM adjudication for pairs below threshold. Clinicians and LLMs interact with the same memory-gated patients, but the paper does not report a calibration study of human vs. LLM utterance matching on identical elicited content (inter-rater agreement on k ⪯ yc, or re-scoring of clinician transcripts under the same automatic matcher). Table 3’s large CC–MS gaps for LLMs could therefore partly reflect phrasing/style sensitivity of the matcher rather than pure competence. A load-bearing fix is to (i) report human–LLM matching agreement on a shared probe set, (ii) ablate τ, and (iii) recompute the gap under exact/grounding-based coverage where available (Appendix K already logs patient grounding fields).","section":"§4, Eq. (13), Tables 2–3"},{"comment":"§2.2, Table 4, Appendix L: the claim that skill-augmented patients constitute a faithful, non-leaking gold standard rests on a small objective probe set (20 cases × 12 probes = 240 responses) and three-clinician subjective ratings. There is no quantitative check that de-identification (Appendix F) preserved psychopathological logic at the checkpoint level, nor a leakage audit showing that patients never disclose unasked future-stage evidence under adversarial doctor prompts. Because both objective coverage and the LLM–clinician gap are measured against these patients, expand the fidelity study (more cases, adversarial probes, inter-psychiatrist agreement on whether disclosed content matches the original EHR logic) or qualify the gap as conditional on the current patient construction.","section":"§2.2, Table 4, Appendix L"},{"comment":"§3.1 and Table 5: MentalEval’s cold-start SFT is supervised by a five-LLM judge ensemble that the paper itself reports as poorly aligned with specialists (LLM-as-a-Judge QWK 0.677, Acc. 0.225). Although expert-guided DPO on low-confidence sets raises average QWK to 0.944 on held-out cases, residual dependence on weak LLM judges for the bulk of SFT targets is a load-bearing design choice for scalable subjective scores. Report (a) the fraction of SFT data that survived consensus filtering vs. was rewritten/augmented, (b) agreement of SFT-only vs. SFT+DPO evaluators stratified by score extremity, and (c) whether clinician DPO preferences were collected independently of the models being ranked in Table 2.","section":"§3.1, Table 5"}],"minor_comments":[{"comment":"Abstract and §4.5 report clinician fidelity as 3.88/5 while the introduction states 3.96/5; reconcile the two figures and state which sample (experts only vs. experts+trainees) each uses.","section":"Abstract / §1 / §4.5"},{"comment":"Table 2 lists Empathy scores where lower appears better for humans (experts 1.22) but higher for some LLMs; clarify whether the empathy rubric is inverted relative to other 1–5 dimensions or whether experts deliberately suppress affective language.","section":"Table 2, §4.1"},{"comment":"Eq. (1) uses xc = {Kpat_c, Kexam_c} and Y*_c; later ˆyc uses different symbols for the same conceptual objects. A short notation table in §2 would reduce reader load.","section":"§2, Eq. (1)–(4)"},{"comment":"Figure 3 confusion matrices are hard to read in grayscale; add numeric cell annotations and a shared color scale.","section":"Figure 3"},{"comment":"Appendix A Limitations correctly flags missing safety/adversarial evaluation (self-harm, psychosis reinforcement); a short pointer in the main-text conclusion would set expectations for deployment claims.","section":"§6 / Appendix A"},{"comment":"Several model names appear inconsistently (Deepseek-v4-Pro vs. DeepSeek-V4-Pro; Claude-Sonnet-4.6 vs. Claude-Sonnet-4-6). Normalize throughout.","section":"Table 2, §4"}],"recommendation":"major_revision","confidential_remarks":"Strong systems/evaluation paper for a medical-AI or general ML venue if the coverage-metric calibration and patient-fidelity expansions are done carefully. The 37.28 pp number is currently over-precise relative to the free parameters (τ, k, β) and matching procedure; editors may want authors to soft-pedal the exact percentage until calibration is shown. Novelty relative to AgentClinic / AI Hospital is real on the psychiatry + full S.O.A.P. + specialist evaluators axes. No integrity red flags; controlled EHR release is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first environment I have seen that actually runs a complete psychiatric encounter—interview, exams, note, diagnosis, treatment—on multi-center EHR cases rather than synthetic dialogues. That is the real advance. They built skill-augmented patients (representation + memory) from 1,193 de-identified records covering 76 ICD-11 disorders, dual-track scoring, and MentalEval (five Qwen3-8B judges trained with rubric SFT then clinician DPO). Clinician survey fidelity is 3.88/5; MentalEval hits average QWK 0.944 on held-out expert labels. The headline number—strongest LLM still 37 pp behind experts on objective metrics, with mental-status as the clear weak spot—is backed by Table 2, the CC/MS split in Table 3, and the patient ablations in Table 4.\n\nWhat they did well: the construction is careful (on-prem de-id, dual psychiatrist review, κ>0.8), the evaluation is not pure LLM-as-judge, and they report complementary strengths (LLMs higher on expressed empathy and treatment appropriateness, lower on professionalism and notes). The release plan is honest about privacy: code, rubrics, and MentalEval weights public; raw EHRs controlled access only.\n\nSoft spots, in proportion. The coverage metric uses τ=0.85 plus Multi-LLM adjudication; the stress-test worry that this under-counts LLM mental-status recovery relative to clinicians is plausible but not fatal. They do not report a human–LLM matching calibration study on the same utterances, so part of the 37 pp could be style sensitivity. That is a real but secondary caveat, not a circularity problem—objective checkpoints are still EHR-anchored and MentalEval is preference-aligned on independent clinician choices. Free parameters (τ, consensus k, DPO β) are disclosed. No public raw data is the main reproducibility tax, which they acknowledge.\n\nThis is for people building or evaluating medical LLMs and for psychiatric education researchers. It deserves a serious referee. I would bring it to reading group and cite the gap and the environment when I next write about clinical LLM evaluation.","headline":"Solid full-S.O.A.P. psychiatric benchmark with real multi-center EHRs and specialist-aligned judges; the 37-point gap is real enough to cite, with a modest metric-calibration caveat on mental-status coverage.","tokens_in":39248,"tokens_out":549,"would_cite":true,"duration_ms":7297,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Even the strongest large language models still trail clinicians by 37 percentage points on complete psychiatric encounters, with mental-status assessment as the main bottleneck.","keywords":["psychiatric clinical encounters","large language models","S.O.A.P. workflow","standardized patients","electronic health records","dual-track evaluation","MentalEval","mental status assessment"],"falsifier":"Re-run the same doctor agents on a held-out multi-center set with independent psychiatrist re-annotation of checkpoints and live standardized-patient sessions; if the LLM–clinician objective gap shrinks below roughly ten points or mental-status coverage ceases to be the dominant miss, the measured competence gap is an artifact of the current patient construction.","tokens_in":39214,"feed_emoji":"🏥","tokens_out":1007,"duration_ms":15234,"temperature":0.7,"pith_summary":"This paper argues that success on isolated psychiatric tasks—dialogue, diagnosis, or treatment planning—does not show that language models can run a full clinical encounter. It builds MentalHospital, a virtual hospital that forces models through the entire S.O.A.P. workflow: interview a skill-augmented standardized patient drawn from real electronic health records, order exams, write notes, diagnose, and plan treatment. Encounters are scored two ways: objective recovery of EHR-derived clinical facts, and process quality judged by MentalEval, five specialist-trained scorers for empathy, professionalism, notes, diagnostic rigor, and treatment fit. Clinicians rate the environment as clinically plausible, and MentalEval tracks expert ratings closely. Under this protocol, even the best models lag medical trainees and experts by large margins on objective competence, especially when they must elicit and use mental-status findings.","feed_headline":"Strongest LLMs trail clinicians by 37 points in psychiatry","feed_subtitle":"A full hospital simulation shows mental-status exams are the main failure mode","key_machinery":"MentalHospital: an EHR-grounded S.O.A.P. simulation using skill-augmented standardized patients (role plus presentation skill plus topic-level memory skill) from 1,193 de-identified cases covering all major ICD-11 psychiatric categories and 76 disorders, paired with a dual-track evaluation protocol and MentalEval—five Qwen3-8B evaluators trained by rubric-grounded supervised fine-tuning then expert-guided preference optimization—to scale specialist judgment of process quality.","core_discovery":"Full-process psychiatric competence is still far from clinician level: the strongest medical-specific model trails medical trainees by about 27 points and human experts by about 37 points on average objective metrics across interviewing, examination, notes, category diagnosis, and disorder diagnosis, with mental-status coverage as a recurring failure mode. Subjective process scores are mixed—models can match or exceed humans on expressed empathy and treatment appropriateness while remaining weaker on interviewing professionalism and note quality—so models are not substitutes, but may complement clinicians in affective support and treatment assistance.","pith_inferences":["If mental-status probing is the bottleneck, interview curricula that force explicit MSE checklists before diagnosis may close more of the gap than larger generic models alone.","The dual-track split implies future model cards should report objective evidence recovery and process quality separately rather than a single medical accuracy score.","Because comorbid and multi-center cases are already in the bank, the same scaffold could stress-test safety behaviors (self-harm, psychosis reinforcement) without inventing synthetic patients from scratch.","Clinician survey positivity for training suitability suggests the environment may first land as a trainee simulator even if model scores remain low."],"forward_implications":["Psychiatric AI evaluation must move from single-turn or dialogue-only tests to full interview–exam–note–diagnosis–treatment episodes with EHR-grounded references.","Mental-status examination becomes a primary training and evaluation target, not an optional side skill.","Specialist-aligned judges (rubric SFT + expert preference) can replace general LLM-as-judge for scalable process scoring in psychiatry.","LLMs may be useful as empathic communication and treatment-drafting assistants while remaining unsuitable as autonomous clinicians under this protocol.","Controlled access to de-identified EHR-derived cases plus public environment code can become a standard for psychiatric agent benchmarks."],"fun_headline_variants":["Strongest LLMs trail clinicians by 37 points in psychiatry encounters","Mental-status exams emerge as main bottleneck for LLMs in psych sim","Top LLMs lag experts 37 points on full SOAP psychiatric competence","Virtual hospital shows models 37 points behind on objective psych metrics","Full clinical encounters leave strongest LLMs far short of clinician level"],"cache_read_input_tokens":32896,"weakest_assumption_plain":"That the de-identified EHR checkpoints and the skill-augmented patients form a faithful, non-leaking gold standard for what a doctor should recover and how a real psychiatric patient would present.","fun_headline_variants_meta":{"raw":{"variants":["Strongest LLMs trail clinicians by 37 points in psychiatry encounters","Mental-status exams emerge as main bottleneck for LLMs in psych sim","Top LLMs lag experts 37 points on full SOAP psychiatric competence","Virtual hospital shows models 37 points behind on objective psych metrics","Full clinical encounters leave strongest LLMs far short of clinician level"]},"model":"grok-4.5","effort":"low","cost_usd":0.004352,"raw_usage":{"total_tokens":1336,"prompt_tokens":820,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":43520000,"prompt_tokens_details":{"text_tokens":820,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":443,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":820,"tokens_out":73,"duration_ms":4566,"temperature":1.0,"reasoning_tokens":443,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T10:29:20.233429+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same doctor agents on a held-out multi-center set with independent psychiatrist re-annotation of checkpoints and live standardized-patient sessions; if the LLM–clinician objective gap shrinks below roughly ten points or mental-status coverage ceases to be the dominant miss, the measured competence gap is an artifact of the current patient construction.","supporting_citations":[],"review_version":1}