REVIEW 4 major objections 6 minor 4 cited by
Superhuman performance of a large language model on the reasoning tasks of a physician
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The o1 language model outperformed board-certified physicians on diagnostic and management reasoning across five benchmarks and real emergency-room second opinions.
desk verdict Strongest o1 clinical reasoning evaluation so far, but 'superhuman' overstates what the data show, especially in the two-physician ER study. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluation rests on three instruments: the Bond score, a 0-5 scale rating whether a differential list contains the exact or nearly exact diagnosis; the R-IDEA scale, a validated 10-point measure of documented clinical reasoning; and a battery of case sets with historical physician and GPT-4 controls from earlier studies. A blinded adjudication protocol, in which two attending physicians scored AI and human differentials without knowing the source, carries the emergency room comparison, with a locally hosted language model used only to standardize formatting. Statistical comparisons use mixed-effects models with random intercepts for case and participant.
What would settle it
A prospective, pre-registered trial that presents the same cases, prompts, and scoring rubrics to o1 and to a concurrently enrolled panel of board-certified physicians; if physicians match or exceed the model's Bond and R-IDEA scores, the 'superhuman' claim would be refuted.
Extended reading notes
Core claim
The central claim, stated on the paper's own terms, is that o1-preview and o1 display superhuman diagnostic and management reasoning. On 143 clinicopathological conference cases, the model included the correct diagnosis in 78.3% of differentials and selected a correct or helpful next test in 98.5% of scored cases. On structured reasoning documentation, it achieved a perfect R-IDEA score in 78 of 80 responses, above both attending and resident physicians from a prior study. In the emergency department study, its proportion of differentials with the exact or very close diagnosis was 65.8% at triage, 69.6% at physician evaluation, and 79.7% at admission, surpassing two attending physicians at each stage. The authors conclude that the model is ready for prospective trials.
Load-bearing premise
The superhuman claim rests on the comparability of historical human and GPT-4 baselines that were collected in earlier studies with different prompts, case sets, and scoring conditions, with no concurrently enrolled physicians in most experiments.
Editorial extensions
If this is right
- If the model's diagnostic edge holds, prospective randomized trials in emergency triage become the immediate next step, because that is where the largest measured gap appears.
- AI-generated second opinions could be tested as a safety layer at admission and ICU-transfer decisions, where the model scored highest in absolute terms.
- Clinical reasoning benchmarks will need concurrent physician baselines rather than historical controls, since prior model generations and doctors were evaluated on different case sets.
- Documentation quality, as measured by R-IDEA, may change what residency training and note assessment expect from augmented physicians.
- Deployment attention should shift from standalone diagnosis to human-computer interaction design, since the study measured model-only performance, not physician-plus-model teams.
Reading between the lines
- The paper's own Figure 1 caption notes that prior models and o1 were not evaluated on the same case sets, and several comparisons use historical controls; a reader should treat the 'superhuman' label as conditional on baseline comparability until a same-case, same-prompt head-to-head trial is run.
- If the model's edge is genuinely largest when information is scarce, the most promising near-term use is decision support at nurse triage, not replacement of specialist consultation, where the model's absolute advantage narrows.
- Benchmark saturation may soon force the field to move from case-based accuracy to process measures such as calibration, uncertainty expression, and the value of a second opinion in changing management.
- A testable extension would be to run the same physician adjudicators on both human and AI differentials for the same patients with concurrent enrollment, and to measure whether the AI's advantage changes with case rarity or patient complexity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates OpenAI o1-preview and o1 against physician baselines in six settings: NEJM clinicopathological conference (CPC) differential diagnosis and test selection, NEJM Healer diagnostic reasoning documentation, Grey Matters management cases, Landmark diagnostic cases, probabilistic reasoning, and a real-world emergency department (ED) second-opinion study. Performance is scored by physician raters using the Bond score, the R-IDEA scale, and rubrics from prior studies. The authors claim that the LLM displayed 'superhuman diagnostic and reasoning abilities' in all experiments, based largely on comparisons with historical controls from the same author group and on a small ED study against two attending physicians.
Significance. The study is valuable for its scale and design features: 143 NEJM CPCs adjudicated by physicians, validated psychometric instruments (Bond score, R-IDEA), a memorization sensitivity analysis around the model's pretraining cutoff, and a blinded real-world second-opinion protocol. If the central claim were fully supported, the result would be an important benchmark for LLM diagnostic reasoning. The strongest findings—78.3% inclusion of the correct CPC diagnosis, 87.5% correct next-test selection, and 78/80 perfect R-IDEA scores—are genuinely notable. However, the paper's headline claim is not yet supported by the reported analyses: several experiments show parity or non-significant differences, and the real-world component lacks any inferential statistics.
major comments (4)
- [Landmark Diagnostic Cases / Figure 5B] The abstract and Discussion claim 'superhuman performance in every experiment,' but the Landmark Diagnostic Cases analysis shows o1-preview performed comparably to GPT-4 (4.4% higher, 95% CI, -19.0% to 27.7%; p=0.7) and its advantage over physicians with GPT-4 (p=0.076) and physicians with conventional resources (p=0.055) did not reach statistical significance. Figure 5B is labeled 'Ns: not statistically significant.' This directly contradicts the unqualified 'superhuman' claim. The authors should replace the global claim with a per-experiment quantitative summary of effect sizes, confidence intervals, and significance, and qualify the conclusion accordingly.
- [Emergency Room Cases / Figure 7] The real-world ED second-opinion experiment is the cornerstone of the 'superhuman' claim, but it compares o1 and GPT-4o with exactly two board-certified internal medicine physicians, not emergency physicians, and reports only raw proportions (e.g., 65.8% vs 54.4% and 48.1% at triage) with no confidence intervals, p-values, effect sizes, or a mixed-effects model treating physician as a random effect. With 79 cases and a binary Bond 4/5 outcome, these differences may be within sampling variability. The Discussion concedes the experiment 'is best thought of as a proof-of-concept,' which is in tension with the abstract's unqualified real-world claim. The authors should add appropriate inferential statistics, report CIs for every comparison, and temper the real-world conclusion.
- [Statistical Analysis and Figure 1] Several key comparisons use historical controls from prior studies by the same author group, with different prompts, case sets, scoring conditions, and no concurrently collected human data. Figure 1's caption explicitly states that 'the set of CPCs each model was evaluated on are not the same,' yet the barplot juxtaposes those models' accuracies without adjustment. The same comparability problem applies to the GPT-4, attending, and resident baselines in Figures 2, 4, and 5. This is load-bearing for the physician-baseline claims. The authors should either include a concurrent human baseline collected under identical conditions, or label all human comparisons as historical and provide a formal sensitivity analysis assessing cross-study comparability.
- [Methods, Blinded Physician Evaluation of Emergency Room Cases] The handling of missing responses is not justified: two responses left unanswered by Physician 1 and one by o1 are assigned a Bond score of 0. This imputation could differentially penalize the human physicians if unanswered responses reflect process constraints rather than diagnostic failure. The text should report the number of missing responses per arm and analyze the data with and without this imputation, or use a more principled missing-data approach.
minor comments (6)
- [Abstract] The abstract says 'We conduct five experiments' and then refers to 'all experiments—both vignettes and emergency room second opinions,' which implies at least six distinct evaluations; the enumeration should be made consistent.
- [Table 3] Table 3 carries a footnote '*: p <= 0.05' but reports no p-values for the probabilistic reasoning comparisons; the footnote should be removed or the corresponding significance tests added.
- [NEJM CPC test selection] The inter-rater agreement for the test-selection outcome is κ=0.28, which is low; although the authors attribute this to class imbalance, the low reliability of a primary outcome in that experiment should be discussed explicitly as a limitation.
- [Introduction] The phrase 'has been seen an aspirational goal post' is ungrammatical and should read 'has been seen as an aspirational goal post' or similar; the same section also contains the typo 'computed based diagnostic systems' in the Methods for Landmark cases.
- [Emergency Room Cases / Methods] The two human comparators are described as 'board-certified physicians' and 'expert attending physicians,' but they were internal medicine attendings reviewing transcribed EHR touchpoints outside the live clinical workflow; the description should state their specialty and task constraints explicitly.
- [Figure 4B] The figure legend reports a total sample size of 70 with 18 responses from attending physicians, GPT-4, and o1-preview, and 16 responses from residents; the text does not explain why the resident count is 16 rather than 18 after excluding the two cases without cannot-miss diagnoses, so the discrepancy should be clarified.
Circularity Check
No significant circularity: the paper is an empirical benchmarking study with no derivation-level reduction of outputs to inputs.
full rationale
This paper makes no formal derivation or first-principles claim; it is an empirical benchmarking study. The central result is a set of accuracy comparisons between o1/o1-preview, GPT-4/GPT-4o, and human physicians across six tasks. Model outputs are scored by physician adjudicators against externally defined rubrics (Bond Score, R-IDEA, expert-consensus management rubrics, and literature-based probability reference ranges). The human and GPT-4 baselines are historical controls taken from prior studies, several by overlapping author groups (refs 12, 15, 17, 18, 29), but those citations supply measured outcome data rather than defining the outcome in terms of the current model. No parameter is fitted and later relabeled as a prediction; no benchmark is constructed from the model's own outputs; and no uniqueness or permissibility theorem is imported from the authors' prior work. The emergency-room comparison does use only two attending physicians as the human baseline and the paper itself calls that experiment a proof-of-concept, but that is a statistical-power and generalizability limitation, not circularity. The self-citations are ordinary references to prior data collections and do not make the present claim equivalent to its inputs by construction. Accordingly, the correct finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The Bond score, a 0-5 rating of differential diagnosis lists, is a valid proxy for diagnostic reasoning quality.
- domain assumption Historical human and GPT-4 baselines from prior studies are comparable to o1-preview performance even though case sets and prompts sometimes differ.
- domain assumption The final diagnosis determined by two physicians from chart review is the correct ground truth for the emergency room cases.
- domain assumption o1-preview did not memorize the NEJM CPC answers despite a training cutoff before some cases; the before-and-after cutoff sensitivity analysis is evidence of this.
Cite this review
Pith. "Pith review of Superhuman performance of a large language model on the reasoning tasks of a physician." pith.science (2026). https://pith.science/paper/5426EBKH
@misc{pith2026241210849,
author = {Pith},
title = {Pith review of: Superhuman performance of a large language model on the reasoning tasks of a physician},
year = {2026},
howpublished = {\url{https://pith.science/paper/5426EBKH}},
note = {Machine review of arXiv:2412.10849}
}
read the original abstract
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials.
Forward citations
Cited by 4 Pith papers
-
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
MedAgentBench is a FHIR-based virtual EHR benchmark with 300 physician-written tasks and 100 patient profiles, where the best LLM agent (Claude 3.5 Sonnet v2) succeeds on 69.67% of tasks.
-
A global log for medical AI
MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.
-
Teaching large language models to reason like expert diagnosticians
An LLM agent and a 10-task benchmark built from 7,102 NEJM clinicopathologic cases push medical AI evaluation beyond final diagnosis accuracy.
-
Towards medical AI misalignment: a preliminary study
A custom role-playing prompt called the Goofy Game made four major LLMs produce plausible but incorrect medical recommendations.
Reference graph
Works this paper leans on
-
[1]
R. S. Ledley, L. B. Lusted, Reasoning foundations of medical diagnosis; symbolic logic, probability, and value theory aid our understanding of how physicians reason. Science 130 , 9–21 (1959)
work page 1959
-
[2]
K. Brodman, A. J. Erdmann Jr, I. Lorge, C. P. Gershenson, H. G. Wolff, The Cornell Medical Index-Health Questionnaire. III. The evaluation of emotional disturbances. J. Clin. Psychol. 8 , 119–124 (1952)
work page 1952
-
[3]
F. T. de Dombal, D. J. Leaper, J. R. Staniland, A. P. McCann, J. C. Horrocks, Computer-aided diagnosis of acute abdominal pain. Br. Med. J. 2 , 9–13 (1972)
work page 1972
-
[4]
E. H. Shortliffe, Mycin: A knowledge-based computer program applied to infectious diseases. Proc. Annu. Symp. Comput. Appl. Med. Care , 66–69 (1977)
work page 1977
-
[5]
E. B. Ing, M. Balas, G. Nassrallah, D. DeAngelis, N. Nijhawan, The Isabel differential diagnosis generator for orbital diagnosis. Ophthal. Plast. Reconstr. Surg. 39 , 461–464 (2023)
work page 2023
-
[6]
E. L. Burkett, B. R. Todd, A novel use of an electronic differential diagnosis generator in the emergency department setting. Cureus 15 , e34211 (2023)
work page 2023
-
[7]
H. Nori, N. Usuyama, N. King, S. McKinney, X. Fernandes, S. Zhang, E. Horvitz, From Medprompt to o1: Exploration of run-time strategies for medical challenge problems and beyond. arXiv (2024)
work page 2024
-
[8]
https://openai.com/index/openai-o1-system-card/
OpenAI o1 System Card. https://openai.com/index/openai-o1-system-card/
Show all 29 references
-
[9]
P. Lee, S. Bubeck, J. Petro, Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 388 , 1233–1239 (2023)
2023
-
[10]
Bubeck, V
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, Y. Zhang, Sparks of Artificial General Intelligence: Early experiments with GPT-4, arXiv [cs.CL] (2023). http://arxiv.org/abs/2303.12712
2023 arXiv
-
[11]
Johri, J
S. Johri, J. Jeong, B. A. Tran, D. I. Schlessinger, S. Wongvibulsin, L. A. Barnes, H.-Y. Zhou, Z. R. Cai, E. M. Van Allen, D. Kim, R. Daneshjou, P. Rajpurkar, An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 31 , 77–86 (2025)
2025
-
[12]
Kanjee, B
Z. Kanjee, B. Crowe, A. Rodman, Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330 , 78–80 (2023)
2023
-
[13]
Rodman, L
A. Rodman, L. Zwaan, A. Olson, A. K. Manrai, When it comes to benchmarks, humans are the only way. NEJM AI 2 , AIe2500143 (2025)
2025
-
[14]
R.-E. E. Abdulnour, A. S. Parsons, D. Muller, J. Drazen, E. J. Rubin, J. Rencic, Deliberate practice at the virtual bedside to improve clinical reasoning. N. Engl. J. Med. 386 , 1946–1947 (2022)
2022
-
[15]
Cabral, D
S. Cabral, D. Restrepo, Z. Kanjee, P. Wilson, B. Crowe, R.-E. Abdulnour, A. Rodman, Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians. JAMA Intern. Med. 184 , 581–583 (2024)
2024
-
[16]
Schaye, L
V. Schaye, L. Miller, D. Kudlowitz, J. Chun, J. Burk-Rafel, P. Cocks, B. Guzman, Y. Aphinyanaphongs, M. Marin, Development of a Clinical Reasoning Documentation Assessment Tool for Resident and Fellow Admission Notes: a Shared Mental Model for Feedback. J. Gen. Intern. Med. 37...
2022
-
[17]
E. Goh, R. Gallo, E. Strong, Y. Weng, H. Kerman, J. Freed, J. A. Cool, Z. Kanjee, K. P. Lane, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, J. Hom, J. H. Chen, A. Rodman, Large language model influence on management reasoning: A randomized controll...
2024 doi
-
[18]
E. Goh, R. Gallo, J. Hom, E. Strong, Y. Weng, H. Kerman, J. A. Cool, Z. Kanjee, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, A. Rodman, J. H. Chen, Large language model influence on diagnostic reasoning: A randomized clinical trial: A randomized c...
2024
-
[19]
E. S. Berner, G. D. Webster, A. A. Shugerman, J. R. Jackson, J. Algina, A. L. Baker, E. V. Ball, C. G. Cobbs, V. W. Dennis, E. P. Frenkel, Performance of four computer-based diagnostic systems. N. Engl. J. Med. 330 , 1792–1796 (1994)
1994
-
[20]
D. J. Morgan, L. Pineles, J. Owczarzak, L. Magder, L. Scherer, J. P. Brown, C. Pfeiffer, C. Terndrup, L. Leykum, D. Feldstein, A. Foy, D. Stevens, C. Koch, M. Masnick, S. Weisenberg, D. Korenstein, Accuracy of practitioner estimates of probability of diagnosis before and after...
2021
-
[21]
R. M. Ratwani, D. W. Bates, D. C. Classen, Patient safety and artificial intelligence in clinical care. JAMA Health Forum 5 , e235514 (2024)
2024
-
[22]
Q. Jin, F. Chen, Y. Zhou, Z. Xu, J. M. Cheung, R. Chen, R. M. Summers, J. F. Rousseau, P. Ni, M. J. Landsman, S. L. Baxter, S. J. Al’Aref, Y. Li, A. Chen, J. A. Brejt, M. F. Chiang, Y. Peng, Z. Lu, Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicin...
2024
-
[23]
D. E. Newman-Toker, S. M. Peterson, S. Badihian, A. Hassoon, N. Nassery, D. Parizadeh, L. M. Wilson, Y. Jia, R. Omron, S. Tharmarajah, Others, Diagnostic errors in the emergency department: a systematic review. (2023)
2023
-
[24]
A. D. Auerbach, T. M. Lee, C. C. Hubbard, S. R. Ranji, K. Raffel, G. Valdes, J. Boscardin, A. K. Dalal, A. Harris, E. Flynn, J. L. Schnipper, UPSIDE Research Group, Diagnostic errors in hospitalized adults who died or were transferred to intensive care. JAMA Intern. Med. 184 ,...
2024
-
[25]
Goldszmidt, J
M. Goldszmidt, J. P. Minda, G. Bordage, Developing a unified list of physicians’ reasoning tasks during clinical encounters. Acad. Med. 88 , 390–394 (2013)
2013
-
[26]
W. F. Bond, L. M. Schwartz, K. R. Weaver, D. Levick, M. Giuliano, M. L. Graber, Differential diagnosis generators: an evaluation of currently available computer programs. J. Gen. Intern. Med. 27 , 213–219 (2012)
2012
-
[27]
McDuff, M
D. McDuff, M. Schaekermann, T. Tu, A. Palepu, A. Wang, J. Garrison, K. Singhal, Y. Sharma, S. Azizi, K. Kulkarni, L. Hou, Y. Cheng, Y. Liu, S. Sara Mahdavi, S. Prakash, A. Pathak, C. Semturs, S. Patel, D. R. Webster, E. Dominowska, J. Gottweis, J. Barral, K. Chou, G. S. Corrad...
2023 arXiv
-
[28]
Fritz, A
P. Fritz, A. Kleinhans, R. Raoufi, A. Sediqi, N. Schmid, S. Schricker, M. Schanz, C. Fritz-Kuisle, P. Dalquen, H. Firooz, G. Stauch, M. D. Alscher, Evaluation of medical decision support systems (DDX generators) using real medical cases of varying complexity and origin. BMC Me...
2022
-
[29]
Rodman, T
A. Rodman, T. A. Buckley, A. K. Manrai, D. J. Morgan, Artificial intelligence vs clinician performance in estimating probabilities of diagnoses before and after testing. JAMA Netw. Open 6 , e2347075 (2023)
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.