Pith. sign in

REVIEW 4 major objections 6 minor 4 cited by

Superhuman performance of a large language model on the reasoning tasks of a physician

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The o1 language model outperformed board-certified physicians on diagnostic and management reasoning across five benchmarks and real emergency-room second opinions.

desk verdict Strongest o1 clinical reasoning evaluation so far, but 'superhuman' overstates what the data show, especially in the two-physician ER study. read the letter →

arxiv 2412.10849 v3 pith:5426EBKH submitted 2024-12-14 cs.AI cs.CL

classification cs.AIcs.CL
keywords largelanguagemodelsclinicalreasoningdifferentialdiagnosismanagementemergencydepartmentphysiciancomparisondiagnosticbenchmarkso1model
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to test whether a current large language model, the o1 series, has reached or passed the diagnostic and management reasoning level of board-certified physicians. It compares the model with hundreds of physicians across five vignette experiments covering differential diagnosis, documented reasoning, triage differentials, probability estimates, and management choices, then adds a blinded comparison on real emergency department cases at three decision points. The authors claim the model outperformed physicians in every experiment, with the largest margin at the earliest triage step where information is most limited. If true, this would mean the long-standing goal of computer systems that can reason through complex clinical cases has been met, and the next step is prospective trials of AI-assisted care.

What carries the argument

The evaluation rests on three instruments: the Bond score, a 0-5 scale rating whether a differential list contains the exact or nearly exact diagnosis; the R-IDEA scale, a validated 10-point measure of documented clinical reasoning; and a battery of case sets with historical physician and GPT-4 controls from earlier studies. A blinded adjudication protocol, in which two attending physicians scored AI and human differentials without knowing the source, carries the emergency room comparison, with a locally hosted language model used only to standardize formatting. Statistical comparisons use mixed-effects models with random intercepts for case and participant.

What would settle it

A prospective, pre-registered trial that presents the same cases, prompts, and scoring rubrics to o1 and to a concurrently enrolled panel of board-certified physicians; if physicians match or exceed the model's Bond and R-IDEA scores, the 'superhuman' claim would be refuted.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that o1-preview and o1 display superhuman diagnostic and management reasoning. On 143 clinicopathological conference cases, the model included the correct diagnosis in 78.3% of differentials and selected a correct or helpful next test in 98.5% of scored cases. On structured reasoning documentation, it achieved a perfect R-IDEA score in 78 of 80 responses, above both attending and resident physicians from a prior study. In the emergency department study, its proportion of differentials with the exact or very close diagnosis was 65.8% at triage, 69.6% at physician evaluation, and 79.7% at admission, surpassing two attending physicians at each stage. The authors conclude that the model is ready for prospective trials.

Load-bearing premise

The superhuman claim rests on the comparability of historical human and GPT-4 baselines that were collected in earlier studies with different prompts, case sets, and scoring conditions, with no concurrently enrolled physicians in most experiments.

Editorial extensions

If this is right

  • If the model's diagnostic edge holds, prospective randomized trials in emergency triage become the immediate next step, because that is where the largest measured gap appears.
  • AI-generated second opinions could be tested as a safety layer at admission and ICU-transfer decisions, where the model scored highest in absolute terms.
  • Clinical reasoning benchmarks will need concurrent physician baselines rather than historical controls, since prior model generations and doctors were evaluated on different case sets.
  • Documentation quality, as measured by R-IDEA, may change what residency training and note assessment expect from augmented physicians.
  • Deployment attention should shift from standalone diagnosis to human-computer interaction design, since the study measured model-only performance, not physician-plus-model teams.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own Figure 1 caption notes that prior models and o1 were not evaluated on the same case sets, and several comparisons use historical controls; a reader should treat the 'superhuman' label as conditional on baseline comparability until a same-case, same-prompt head-to-head trial is run.
  • If the model's edge is genuinely largest when information is scarce, the most promising near-term use is decision support at nurse triage, not replacement of specialist consultation, where the model's absolute advantage narrows.
  • Benchmark saturation may soon force the field to move from case-based accuracy to process measures such as calibration, uncertainty expression, and the value of a second opinion in changing management.
  • A testable extension would be to run the same physician adjudicators on both human and AI differentials for the same patients with concurrent enrollment, and to measure whether the AI's advantage changes with case rarity or patient complexity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript evaluates OpenAI o1-preview and o1 against physician baselines in six settings: NEJM clinicopathological conference (CPC) differential diagnosis and test selection, NEJM Healer diagnostic reasoning documentation, Grey Matters management cases, Landmark diagnostic cases, probabilistic reasoning, and a real-world emergency department (ED) second-opinion study. Performance is scored by physician raters using the Bond score, the R-IDEA scale, and rubrics from prior studies. The authors claim that the LLM displayed 'superhuman diagnostic and reasoning abilities' in all experiments, based largely on comparisons with historical controls from the same author group and on a small ED study against two attending physicians.

Significance. The study is valuable for its scale and design features: 143 NEJM CPCs adjudicated by physicians, validated psychometric instruments (Bond score, R-IDEA), a memorization sensitivity analysis around the model's pretraining cutoff, and a blinded real-world second-opinion protocol. If the central claim were fully supported, the result would be an important benchmark for LLM diagnostic reasoning. The strongest findings—78.3% inclusion of the correct CPC diagnosis, 87.5% correct next-test selection, and 78/80 perfect R-IDEA scores—are genuinely notable. However, the paper's headline claim is not yet supported by the reported analyses: several experiments show parity or non-significant differences, and the real-world component lacks any inferential statistics.

major comments (4)
  1. [Landmark Diagnostic Cases / Figure 5B] The abstract and Discussion claim 'superhuman performance in every experiment,' but the Landmark Diagnostic Cases analysis shows o1-preview performed comparably to GPT-4 (4.4% higher, 95% CI, -19.0% to 27.7%; p=0.7) and its advantage over physicians with GPT-4 (p=0.076) and physicians with conventional resources (p=0.055) did not reach statistical significance. Figure 5B is labeled 'Ns: not statistically significant.' This directly contradicts the unqualified 'superhuman' claim. The authors should replace the global claim with a per-experiment quantitative summary of effect sizes, confidence intervals, and significance, and qualify the conclusion accordingly.
  2. [Emergency Room Cases / Figure 7] The real-world ED second-opinion experiment is the cornerstone of the 'superhuman' claim, but it compares o1 and GPT-4o with exactly two board-certified internal medicine physicians, not emergency physicians, and reports only raw proportions (e.g., 65.8% vs 54.4% and 48.1% at triage) with no confidence intervals, p-values, effect sizes, or a mixed-effects model treating physician as a random effect. With 79 cases and a binary Bond 4/5 outcome, these differences may be within sampling variability. The Discussion concedes the experiment 'is best thought of as a proof-of-concept,' which is in tension with the abstract's unqualified real-world claim. The authors should add appropriate inferential statistics, report CIs for every comparison, and temper the real-world conclusion.
  3. [Statistical Analysis and Figure 1] Several key comparisons use historical controls from prior studies by the same author group, with different prompts, case sets, scoring conditions, and no concurrently collected human data. Figure 1's caption explicitly states that 'the set of CPCs each model was evaluated on are not the same,' yet the barplot juxtaposes those models' accuracies without adjustment. The same comparability problem applies to the GPT-4, attending, and resident baselines in Figures 2, 4, and 5. This is load-bearing for the physician-baseline claims. The authors should either include a concurrent human baseline collected under identical conditions, or label all human comparisons as historical and provide a formal sensitivity analysis assessing cross-study comparability.
  4. [Methods, Blinded Physician Evaluation of Emergency Room Cases] The handling of missing responses is not justified: two responses left unanswered by Physician 1 and one by o1 are assigned a Bond score of 0. This imputation could differentially penalize the human physicians if unanswered responses reflect process constraints rather than diagnostic failure. The text should report the number of missing responses per arm and analyze the data with and without this imputation, or use a more principled missing-data approach.
minor comments (6)
  1. [Abstract] The abstract says 'We conduct five experiments' and then refers to 'all experiments—both vignettes and emergency room second opinions,' which implies at least six distinct evaluations; the enumeration should be made consistent.
  2. [Table 3] Table 3 carries a footnote '*: p <= 0.05' but reports no p-values for the probabilistic reasoning comparisons; the footnote should be removed or the corresponding significance tests added.
  3. [NEJM CPC test selection] The inter-rater agreement for the test-selection outcome is κ=0.28, which is low; although the authors attribute this to class imbalance, the low reliability of a primary outcome in that experiment should be discussed explicitly as a limitation.
  4. [Introduction] The phrase 'has been seen an aspirational goal post' is ungrammatical and should read 'has been seen as an aspirational goal post' or similar; the same section also contains the typo 'computed based diagnostic systems' in the Methods for Landmark cases.
  5. [Emergency Room Cases / Methods] The two human comparators are described as 'board-certified physicians' and 'expert attending physicians,' but they were internal medicine attendings reviewing transcribed EHR touchpoints outside the live clinical workflow; the description should state their specialty and task constraints explicitly.
  6. [Figure 4B] The figure legend reports a total sample size of 70 with 18 responses from attending physicians, GPT-4, and o1-preview, and 16 responses from residents; the text does not explain why the resident count is 16 rather than 18 after excluding the two cases without cannot-miss diagnoses, so the discrepancy should be clarified.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical benchmarking study with no derivation-level reduction of outputs to inputs.

full rationale

This paper makes no formal derivation or first-principles claim; it is an empirical benchmarking study. The central result is a set of accuracy comparisons between o1/o1-preview, GPT-4/GPT-4o, and human physicians across six tasks. Model outputs are scored by physician adjudicators against externally defined rubrics (Bond Score, R-IDEA, expert-consensus management rubrics, and literature-based probability reference ranges). The human and GPT-4 baselines are historical controls taken from prior studies, several by overlapping author groups (refs 12, 15, 17, 18, 29), but those citations supply measured outcome data rather than defining the outcome in terms of the current model. No parameter is fitted and later relabeled as a prediction; no benchmark is constructed from the model's own outputs; and no uniqueness or permissibility theorem is imported from the authors' prior work. The emergency-room comparison does use only two attending physicians as the human baseline and the paper itself calls that experiment a proof-of-concept, but that is a statistical-power and generalizability limitation, not circularity. The self-citations are ordinary references to prior data collections and do not make the present claim equivalent to its inputs by construction. Accordingly, the correct finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new entities or fitted parameters. Its load-bearing assumptions are about measurement validity, comparability of historical baselines, ground truth determination, and the absence of training-data contamination.

assumptions (4)
  • domain assumption The Bond score, a 0-5 rating of differential diagnosis lists, is a valid proxy for diagnostic reasoning quality.
    Used as the primary outcome in the NEJM CPC and emergency room experiments; if this scale does not capture clinical reasoning, the comparisons do not measure what the title claims.
  • domain assumption Historical human and GPT-4 baselines from prior studies are comparable to o1-preview performance even though case sets and prompts sometimes differ.
    Used for head-to-head claims in the CPC, Healer, Grey Matters, Landmark, and probabilistic reasoning experiments. Figure 1's caption explicitly states the set of CPCs each model was evaluated on is not the same.
  • domain assumption The final diagnosis determined by two physicians from chart review is the correct ground truth for the emergency room cases.
    Used to score all emergency room differentials; no autopsy, adjudicated outcome panel, or independent diagnostic verification is reported.
  • domain assumption o1-preview did not memorize the NEJM CPC answers despite a training cutoff before some cases; the before-and-after cutoff sensitivity analysis is evidence of this.
    The comparison between cases before and after the cutoff (79.8% vs 73.5%, p=0.59) is underpowered and cannot establish the absence of memorization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Superhuman performance of a large language model on the reasoning tasks of a physician." pith.science (2026). https://pith.science/paper/5426EBKH

@misc{pith2026241210849,
  author       = {Pith},
  title        = {Pith review of: Superhuman performance of a large language model on the reasoning tasks of a physician},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5426EBKH}},
  note         = {Machine review of arXiv:2412.10849}
}
read the original abstract

A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. Here, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases against a baseline of hundreds of physicians. We conduct five experiments to measure clinical reasoning across differential diagnosis generation, display of diagnostic reasoning, triage differential diagnosis, probabilistic reasoning, and management reasoning, all adjudicated by physician experts with validated psychometrics. We then report a real-world study comparing human expert and AI second opinions in randomly-selected patients in the emergency room of a major tertiary academic medical center in Boston, MA. We compared LLMs and board-certified physicians at three predefined diagnostic touchpoints: triage in the emergency room, initial evaluation by a physician, and admission to the hospital or intensive care unit. In all experiments--both vignettes and emergency room second opinions--the LLM displayed superhuman diagnostic and reasoning abilities, as well as continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have achieved superhuman performance on general medical diagnostic and management reasoning, fulfilling the vision put forth by Ledley and Lusted, and motivating the urgent need for prospective trials.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

    cs.LG 2025-01 conditional novelty 7.0 of 10

    MedAgentBench is a FHIR-based virtual EHR benchmark with 300 physician-written tasks and 100 patient profiles, where the best LLM agent (Claude 3.5 Sonnet v2) succeeds on 69.67% of tasks.

  2. A global log for medical AI

    cs.AI 2025-10 conditional novelty 6.0 of 10

    MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.

  3. Teaching large language models to reason like expert diagnosticians

    cs.AI 2025-09 conditional novelty 6.0 of 10

    An LLM agent and a 10-task benchmark built from 7,102 NEJM clinicopathologic cases push medical AI evaluation beyond final diagnosis accuracy.

  4. Towards medical AI misalignment: a preliminary study

    cs.CY 2025-05 conditional novelty 5.0 of 10

    A custom role-playing prompt called the Goofy Game made four major LLMs produce plausible but incorrect medical recommendations.

Reference graph

Works this paper leans on

29 extracted references · 27 canonical work pages · cited by 4 Pith papers

  1. [1]

    R. S. Ledley, L. B. Lusted, Reasoning foundations of medical diagnosis; symbolic logic, probability, and value theory aid our understanding of how physicians reason. Science 130 , 9–21 (1959)

  2. [2]

    Brodman, A

    K. Brodman, A. J. Erdmann Jr, I. Lorge, C. P. Gershenson, H. G. Wolff, The Cornell Medical Index-Health Questionnaire. III. The evaluation of emotional disturbances. J. Clin. Psychol. 8 , 119–124 (1952)

  3. [3]

    F. T. de Dombal, D. J. Leaper, J. R. Staniland, A. P. McCann, J. C. Horrocks, Computer-aided diagnosis of acute abdominal pain. Br. Med. J. 2 , 9–13 (1972)

  4. [4]

    E. H. Shortliffe, Mycin: A knowledge-based computer program applied to infectious diseases. Proc. Annu. Symp. Comput. Appl. Med. Care , 66–69 (1977)

  5. [5]

    E. B. Ing, M. Balas, G. Nassrallah, D. DeAngelis, N. Nijhawan, The Isabel differential diagnosis generator for orbital diagnosis. Ophthal. Plast. Reconstr. Surg. 39 , 461–464 (2023)

  6. [6]

    E. L. Burkett, B. R. Todd, A novel use of an electronic differential diagnosis generator in the emergency department setting. Cureus 15 , e34211 (2023)

  7. [7]

    H. Nori, N. Usuyama, N. King, S. McKinney, X. Fernandes, S. Zhang, E. Horvitz, From Medprompt to o1: Exploration of run-time strategies for medical challenge problems and beyond. arXiv (2024)

  8. [8]

    https://openai.com/index/openai-o1-system-card/

    OpenAI o1 System Card. https://openai.com/index/openai-o1-system-card/

Show all 29 references
  1. [9]

    P. Lee, S. Bubeck, J. Petro, Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 388 , 1233–1239 (2023)

  2. [10]

    Bubeck, V

    S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, Y. Zhang, Sparks of Artificial General Intelligence: Early experiments with GPT-4, arXiv [cs.CL] (2023). http://arxiv.org/abs/2303.12712

  3. [11]

    Johri, J

    S. Johri, J. Jeong, B. A. Tran, D. I. Schlessinger, S. Wongvibulsin, L. A. Barnes, H.-Y. Zhou, Z. R. Cai, E. M. Van Allen, D. Kim, R. Daneshjou, P. Rajpurkar, An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 31 , 77–86 (2025)

  4. [12]

    Kanjee, B

    Z. Kanjee, B. Crowe, A. Rodman, Accuracy of a Generative Artificial Intelligence Model in a Complex Diagnostic Challenge. JAMA 330 , 78–80 (2023)

  5. [13]

    Rodman, L

    A. Rodman, L. Zwaan, A. Olson, A. K. Manrai, When it comes to benchmarks, humans are the only way. NEJM AI 2 , AIe2500143 (2025)

  6. [14]

    R.-E. E. Abdulnour, A. S. Parsons, D. Muller, J. Drazen, E. J. Rubin, J. Rencic, Deliberate practice at the virtual bedside to improve clinical reasoning. N. Engl. J. Med. 386 , 1946–1947 (2022)

  7. [15]

    Cabral, D

    S. Cabral, D. Restrepo, Z. Kanjee, P. Wilson, B. Crowe, R.-E. Abdulnour, A. Rodman, Clinical Reasoning of a Generative Artificial Intelligence Model Compared With Physicians. JAMA Intern. Med. 184 , 581–583 (2024)

  8. [16]

    Schaye, L

    V. Schaye, L. Miller, D. Kudlowitz, J. Chun, J. Burk-Rafel, P. Cocks, B. Guzman, Y. Aphinyanaphongs, M. Marin, Development of a Clinical Reasoning Documentation Assessment Tool for Resident and Fellow Admission Notes: a Shared Mental Model for Feedback. J. Gen. Intern. Med. 37...

  9. [17]

    E. Goh, R. Gallo, E. Strong, Y. Weng, H. Kerman, J. Freed, J. A. Cool, Z. Kanjee, K. P. Lane, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, J. Hom, J. H. Chen, A. Rodman, Large language model influence on management reasoning: A randomized controll...

  10. [18]

    E. Goh, R. Gallo, J. Hom, E. Strong, Y. Weng, H. Kerman, J. A. Cool, Z. Kanjee, A. S. Parsons, N. Ahuja, E. Horvitz, D. Yang, A. Milstein, A. P. J. Olson, A. Rodman, J. H. Chen, Large language model influence on diagnostic reasoning: A randomized clinical trial: A randomized c...

  11. [19]

    E. S. Berner, G. D. Webster, A. A. Shugerman, J. R. Jackson, J. Algina, A. L. Baker, E. V. Ball, C. G. Cobbs, V. W. Dennis, E. P. Frenkel, Performance of four computer-based diagnostic systems. N. Engl. J. Med. 330 , 1792–1796 (1994)

  12. [20]

    D. J. Morgan, L. Pineles, J. Owczarzak, L. Magder, L. Scherer, J. P. Brown, C. Pfeiffer, C. Terndrup, L. Leykum, D. Feldstein, A. Foy, D. Stevens, C. Koch, M. Masnick, S. Weisenberg, D. Korenstein, Accuracy of practitioner estimates of probability of diagnosis before and after...

  13. [21]

    R. M. Ratwani, D. W. Bates, D. C. Classen, Patient safety and artificial intelligence in clinical care. JAMA Health Forum 5 , e235514 (2024)

  14. [22]

    Q. Jin, F. Chen, Y. Zhou, Z. Xu, J. M. Cheung, R. Chen, R. M. Summers, J. F. Rousseau, P. Ni, M. J. Landsman, S. L. Baxter, S. J. Al’Aref, Y. Li, A. Chen, J. A. Brejt, M. F. Chiang, Y. Peng, Z. Lu, Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicin...

  15. [23]

    D. E. Newman-Toker, S. M. Peterson, S. Badihian, A. Hassoon, N. Nassery, D. Parizadeh, L. M. Wilson, Y. Jia, R. Omron, S. Tharmarajah, Others, Diagnostic errors in the emergency department: a systematic review. (2023)

  16. [24]

    A. D. Auerbach, T. M. Lee, C. C. Hubbard, S. R. Ranji, K. Raffel, G. Valdes, J. Boscardin, A. K. Dalal, A. Harris, E. Flynn, J. L. Schnipper, UPSIDE Research Group, Diagnostic errors in hospitalized adults who died or were transferred to intensive care. JAMA Intern. Med. 184 ,...

  17. [25]

    Goldszmidt, J

    M. Goldszmidt, J. P. Minda, G. Bordage, Developing a unified list of physicians’ reasoning tasks during clinical encounters. Acad. Med. 88 , 390–394 (2013)

  18. [26]

    W. F. Bond, L. M. Schwartz, K. R. Weaver, D. Levick, M. Giuliano, M. L. Graber, Differential diagnosis generators: an evaluation of currently available computer programs. J. Gen. Intern. Med. 27 , 213–219 (2012)

  19. [27]

    McDuff, M

    D. McDuff, M. Schaekermann, T. Tu, A. Palepu, A. Wang, J. Garrison, K. Singhal, Y. Sharma, S. Azizi, K. Kulkarni, L. Hou, Y. Cheng, Y. Liu, S. Sara Mahdavi, S. Prakash, A. Pathak, C. Semturs, S. Patel, D. R. Webster, E. Dominowska, J. Gottweis, J. Barral, K. Chou, G. S. Corrad...

  20. [28]

    Fritz, A

    P. Fritz, A. Kleinhans, R. Raoufi, A. Sediqi, N. Schmid, S. Schricker, M. Schanz, C. Fritz-Kuisle, P. Dalquen, H. Firooz, G. Stauch, M. D. Alscher, Evaluation of medical decision support systems (DDX generators) using real medical cases of varying complexity and origin. BMC Me...

  21. [29]

    Rodman, T

    A. Rodman, T. A. Buckley, A. K. Manrai, D. J. Morgan, Artificial intelligence vs clinician performance in estimating probabilities of diagnoses before and after testing. JAMA Netw. Open 6 , e2347075 (2023)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.