{"id":"7afb606e-b2b3-481b-8959-6375c6d3269e","arxiv_id":"2505.14107","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"DiagnosisArena, a 1,113-case benchmark from top journals, shows state-of-the-art LLMs achieve at most 51% top-1 diagnostic accuracy, far below clinical-level competence.","lead":"This paper introduces DiagnosisArena, a benchmark of 1,113 real clinical case reports from top medical journals, and finds that even the best reasoning models score only 51% top-1 accuracy. It suggests that current AI systems are far from reliable diagnostic competence, and that multiple-choice tests overstate their abilities.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No human physician baseline: the claim that models are 'far from professional-level diagnostic competence' is not established, and the unvalidated GPT-4o judge is secondary to this missing comparison.","rationale":"The reader's verdict correctly flags the unvalidated GPT-4o judge and the absence of a human baseline. My stress-test identifies the missing human baseline as the more load-bearing concern: every top-1 and top-5 number in Figure 3 is a measurement of model behavior, but the headline claim is comparative ('far from professional-level'). A perfectly calibrated automatic judge would still not tell us what a professional physician scores on this filtered hard-case subset. In fact, the construction pipeline in Section 3.2 is designed to retain cases that are hard for LLMs: cases solvable by three frontier models are removed, and cases where DeepSeek-R1 cannot reach consensus in 8 samples are excluded. Therefore DiagnosisArena is not a representative sample of typical clinical encounters; it is a curated set of unusually challenging cases. On such a set, expert physicians might also score well below the 90%-plus accuracies seen on exam-style benchmarks, and possibly near or below o3's 51.12%. If so, the conclusion that models are 'far from professional-level' would be unsupported, even though the low absolute accuracy could still justify caution about autonomous deployment. The concrete test I propose - a blinded physician baseline on a sample of DiagnosisArena cases - would settle this directly. The reader's conditional verdict remains appropriate: the benchmark is valuable, but the central comparative claim needs this missing evidence before it can be stated as established.","tokens_in":18239,"tokens_out":5707,"duration_ms":100804,"concrete_test":"Recruit 5-10 board-certified physicians and give each a stratified random sample of 200 DiagnosisArena cases (same case information, physical examination, and diagnostic tests) with instructions to list their top-5 diagnoses in descending confidence, mirroring the LLM prompt in Appendix B. Score their responses with the same GPT-4o judge and, on a 50-case subset, have independent physician raters score them. If physician top-1 accuracy is significantly above o3's 51.12% (e.g., >75%), the 'far from professional-level' claim is supported; if physician accuracy is at or below o3's, the headline claim should be revised or qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and Section 4.3 that current LLMs are 'far from professional-level diagnostic competence' requires knowing what professional-level performance is on DiagnosisArena, but the paper never reports human expert accuracy on the benchmark. Board-certified physicians appear only in the data-curation pipeline (Section 3.2), where they help exclude ambiguous or missing-information cases; they are not used as an evaluation baseline. This matters especially because the construction pipeline deliberately removes cases solvable by Baichuan-M1, DeepSeek-V3, or GPT-4o, and retains only cases where DeepSeek-R1 reaches consensus in 8 samples (Section 3.2). The resulting benchmark is a hard-case subset, not a representative sample of clinical encounters, so expert human performance might be comparable to or even lower than o3's 51.12%. Without that comparison, the headline 'far from professional-level' is not established. The GPT-4o judge issue is real but secondary: it can shift the exact model numbers, whereas the missing professional reference point is what the central claim depends on.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. The paper gives us a new benchmark—1,113 structured cases from real top-tier journal case reports, with a documented pipeline and a leakage analysis—and that is genuinely valuable. The open-ended evaluation and the MCQ comparison are the right design choices. The finding that o3 scores just over 51% and most models are below 20% is striking and worth knowing, even if the absolute numbers are provisional.\n\nThe soft spot is exactly what the stress-test note flags: there is no human physician baseline. The construction pipeline deliberately filters out cases that several LLMs can solve, so DiagnosisArena is a hard-case subset, not a sample of typical clinical encounters. Without a group of board-certified physicians taking the same test, the claim that models are 'far from professional-level diagnostic competence' isn't actually established. A competent clinician might also score around 50% on this subset. I'd soften that claim in the abstract.\n\nThe unvalidated GPT-4o judge is a real but secondary issue. It could move the numbers, but it wouldn't change the overall picture of difficulty. The bigger fix is the baseline.\n\nI also want to give credit for the leakage analysis, which is easy to skip and hard to do well, and for releasing the benchmark and tools. The pipeline is described in enough detail to reproduce, which is more than many benchmark papers do.\n\nNet: this is a solid resource for anyone working on medical LLM evaluation. The headline numbers should be read as provisional until a human baseline and judge validation are added. I'd send it to peer review—the dataset is useful enough that referees should see it—but I'd ask the authors to add those two things before acceptance.","headline":"A genuinely useful hard-case benchmark for LLM diagnosis, but the 'far from professional-level' headline needs a human baseline it doesn't have.","tokens_in":19000,"tokens_out":2240,"would_cite":true,"duration_ms":22955,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:40:11.601217+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}