{"id":"53d1b195-52d9-4195-b452-bf334a4bf470","arxiv_id":"2412.10849","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"One advanced language model matched or beat physician baselines on many diagnostic reasoning benchmarks, but the paper's 'superhuman in every experiment' claim is not fully supported by its own statistics.","lead":"The authors compared OpenAI's o1 models against hundreds of physicians on six medical reasoning tasks, including real emergency room cases. The models performed at or above physician level on most tasks, but the claim of superhuman performance in every experiment is stronger than the evidence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'superhuman' claim rests on an emergency-room baseline of only two physicians, and no confidence intervals or inferential statistics are reported for that experiment, so the real-world claim is not statistically supported.","rationale":"Read in good faith, the paper is a serious evaluation: it uses validated scales (Bond, R-IDEA, etc.), blinded adjudication, and a real-world dataset. The strongest evidence for the abstract's 'superhuman' claim is the emergency-room comparison, because it is the only non-vignette setting. That experiment, however, has only two human physicians as the comparator, lacks any inferential statistics, and is described by the authors as a proof-of-concept in the Discussion. The central 'superhuman in every experiment' sentence therefore goes beyond what the ER data can support; the observed margin at triage could be noise. The reader's weakest_assumption captures this (representativeness of the two physicians) along with the historical-control comparability problem. I focus on the ER baseline because the real-world claim is the most novel and most load-bearing. A conditional verdict is appropriate: the authors should reanalyze the ER data with proper statistical inference and either add physicians or calibrate against published performance norms. If the ER comparison fails to reach significance, the abstract's universal 'superhuman' statement should be narrowed.","tokens_in":24171,"tokens_out":3783,"duration_ms":36394,"concrete_test":"Reanalyze the 79-case emergency-room data with a mixed-effects logistic regression (or paired bootstrap) with physician as a random effect, comparing the probability of Bond score ≥4 for o1 versus the two physicians at each touchpoint, and report the difference with a 95% confidence interval. If the confidence interval for the triage touchpoint includes zero, the claim that o1 is superhuman at the point where it 'surpasses' physicians is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim—'superhuman diagnostic and reasoning abilities' in all experiments—leans most heavily on the emergency-room study (Figure 7), the only real-world component. But that experiment compares o1 and GPT-4o against exactly two attending physicians, and the paper reports no confidence intervals, p-values, or effect sizes for these comparisons. The headline numbers (e.g., 65.8% vs 54.4% and 48.1% at triage) could easily be within sampling variability with 79 cases, especially for a binary high/low Bond-score outcome. The Discussion itself concedes this experiment 'is best thought of as a proof-of-concept,' which is in tension with the abstract's unqualified 'superhuman' statement. A two-physician baseline is also not established as representative: the physicians are internal medicine attendings, not emergency physicians, and they generate differentials from transcribed EHR touchpoints, not in the live clinical workflow, and their lists are capped at five diagnoses. Without a statistical model that treats physician as a random effect, or evidence that these two physicians' accuracy is typical of practicing clinicians, the real-world 'superhuman' claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript evaluates OpenAI o1-preview and o1 against physician baselines in six settings: NEJM clinicopathological conference (CPC) differential diagnosis and test selection, NEJM Healer diagnostic reasoning documentation, Grey Matters management cases, Landmark diagnostic cases, probabilistic reasoning, and a real-world emergency department (ED) second-opinion study. Performance is scored by physician raters using the Bond score, the R-IDEA scale, and rubrics from prior studies. The authors claim that the LLM displayed 'superhuman diagnostic and reasoning abilities' in all experiments, based largely on comparisons with historical controls from the same author group and on a small ED study against two attending physicians.","tokens_in":24346,"tokens_out":4516,"duration_ms":40606,"significance":"The study is valuable for its scale and design features: 143 NEJM CPCs adjudicated by physicians, validated psychometric instruments (Bond score, R-IDEA), a memorization sensitivity analysis around the model's pretraining cutoff, and a blinded real-world second-opinion protocol. If the central claim were fully supported, the result would be an important benchmark for LLM diagnostic reasoning. The strongest findings—78.3% inclusion of the correct CPC diagnosis, 87.5% correct next-test selection, and 78/80 perfect R-IDEA scores—are genuinely notable. However, the paper's headline claim is not yet supported by the reported analyses: several experiments show parity or non-significant differences, and the real-world component lacks any inferential statistics.","major_comments":[{"comment":"The abstract and Discussion claim 'superhuman performance in every experiment,' but the Landmark Diagnostic Cases analysis shows o1-preview performed comparably to GPT-4 (4.4% higher, 95% CI, -19.0% to 27.7%; p=0.7) and its advantage over physicians with GPT-4 (p=0.076) and physicians with conventional resources (p=0.055) did not reach statistical significance. Figure 5B is labeled 'Ns: not statistically significant.' This directly contradicts the unqualified 'superhuman' claim. The authors should replace the global claim with a per-experiment quantitative summary of effect sizes, confidence intervals, and significance, and qualify the conclusion accordingly.","section":"Landmark Diagnostic Cases / Figure 5B"},{"comment":"The real-world ED second-opinion experiment is the cornerstone of the 'superhuman' claim, but it compares o1 and GPT-4o with exactly two board-certified internal medicine physicians, not emergency physicians, and reports only raw proportions (e.g., 65.8% vs 54.4% and 48.1% at triage) with no confidence intervals, p-values, effect sizes, or a mixed-effects model treating physician as a random effect. With 79 cases and a binary Bond 4/5 outcome, these differences may be within sampling variability. The Discussion concedes the experiment 'is best thought of as a proof-of-concept,' which is in tension with the abstract's unqualified real-world claim. The authors should add appropriate inferential statistics, report CIs for every comparison, and temper the real-world conclusion.","section":"Emergency Room Cases / Figure 7"},{"comment":"Several key comparisons use historical controls from prior studies by the same author group, with different prompts, case sets, scoring conditions, and no concurrently collected human data. Figure 1's caption explicitly states that 'the set of CPCs each model was evaluated on are not the same,' yet the barplot juxtaposes those models' accuracies without adjustment. The same comparability problem applies to the GPT-4, attending, and resident baselines in Figures 2, 4, and 5. This is load-bearing for the physician-baseline claims. The authors should either include a concurrent human baseline collected under identical conditions, or label all human comparisons as historical and provide a formal sensitivity analysis assessing cross-study comparability.","section":"Statistical Analysis and Figure 1"},{"comment":"The handling of missing responses is not justified: two responses left unanswered by Physician 1 and one by o1 are assigned a Bond score of 0. This imputation could differentially penalize the human physicians if unanswered responses reflect process constraints rather than diagnostic failure. The text should report the number of missing responses per arm and analyze the data with and without this imputation, or use a more principled missing-data approach.","section":"Methods, Blinded Physician Evaluation of Emergency Room Cases"}],"minor_comments":[{"comment":"The abstract says 'We conduct five experiments' and then refers to 'all experiments—both vignettes and emergency room second opinions,' which implies at least six distinct evaluations; the enumeration should be made consistent.","section":"Abstract"},{"comment":"Table 3 carries a footnote '*: p <= 0.05' but reports no p-values for the probabilistic reasoning comparisons; the footnote should be removed or the corresponding significance tests added.","section":"Table 3"},{"comment":"The inter-rater agreement for the test-selection outcome is κ=0.28, which is low; although the authors attribute this to class imbalance, the low reliability of a primary outcome in that experiment should be discussed explicitly as a limitation.","section":"NEJM CPC test selection"},{"comment":"The phrase 'has been seen an aspirational goal post' is ungrammatical and should read 'has been seen as an aspirational goal post' or similar; the same section also contains the typo 'computed based diagnostic systems' in the Methods for Landmark cases.","section":"Introduction"},{"comment":"The two human comparators are described as 'board-certified physicians' and 'expert attending physicians,' but they were internal medicine attendings reviewing transcribed EHR touchpoints outside the live clinical workflow; the description should state their specialty and task constraints explicitly.","section":"Emergency Room Cases / Methods"},{"comment":"The figure legend reports a total sample size of 70 with 18 responses from attending physicians, GPT-4, and o1-preview, and 16 responses from residents; the text does not explain why the resident count is 16 rather than 18 after excluding the two cases without cannot-miss diagnoses, so the discrepancy should be clarified.","section":"Figure 4B"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a strong, well-executed benchmark evaluation, but the title, abstract, and Discussion overclaim by asserting 'superhuman performance in every experiment' when the authors' own analyses show non-significant or parity results in the Landmark, probabilistic reasoning, and cannot-miss analyses, and no inferential statistics for the ED study. A revision that recalibrates the claims to the evidence and adds the missing statistical analyses would make the contribution both accurate and valuable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jeremy,\n\nQuick take: this is the most comprehensive o1 clinical reasoning evaluation I've seen, and the o1-vs-physician results are genuinely worth engaging with. But the 'superhuman in all experiments' claim in the abstract is not what the data show.\n\nWhat's new: the full 143-case NEJM CPC evaluation, the 80-response R-IDEA comparison, the five Grey Matters management cases, and especially the blinded ER second-opinion study using real EHR touchpoints. Those are real additions. The blinding check in the ER study is well done, and the pre/post cutoff contamination analysis for the CPCs is thoughtful. The paper also does something a lot of LLM papers skip: it uses physician adjudication with interrater kappas, and it is explicit when its own results are proof-of-concept.\n\nWhere it slips: several comparisons do not support the word 'superhuman.' In the Landmark cases, o1 is comparable to GPT-4 (p=0.7). In probabilistic reasoning, the paper itself says o1 performs similarly to GPT-4, and the 'cannot-miss' diagnosis result is not statistically significant. The ER study is the load-bearing real-world evidence, but it compares o1 and GPT-4o to exactly two internal medicine attendings, not emergency physicians, and reports no confidence intervals, p-values, or effect sizes for the headline proportions. With 79 cases, a 65.8% versus 54.4% gap is within sampling variability. The Discussion admits this experiment is best thought of as a proof-of-concept, which is in tension with the abstract's unqualified 'superhuman.'\n\nAlso, many human baselines are historical controls from the authors' earlier studies, with different prompts and case sets. That is not fatal, and the same-case GPT-4 comparison on 70 CPCs is fine, but the 'hundreds of physicians' line overstates the comparability.\n\nBottom line: this paper deserves a serious referee. It is important, the data are real, and the central message that o1 is strong and improving on clinical reasoning is probably right. But the abstract needs to be rewritten to match the evidence, and the ER analysis needs at least confidence intervals and ideally a random-effects model with physicians as a random factor. I would not desk reject it; I would send it out with a request for revision.","headline":"Strongest o1 clinical reasoning evaluation so far, but 'superhuman' overstates what the data show, especially in the two-physician ER study.","tokens_in":25003,"tokens_out":2784,"would_cite":true,"duration_ms":25342,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The o1 language model outperformed board-certified physicians on diagnostic and management reasoning across five benchmarks and real emergency-room second opinions.","keywords":["large language models","clinical reasoning","differential diagnosis","management reasoning","emergency department","physician comparison","diagnostic benchmarks","o1 model"],"falsifier":"A prospective, pre-registered trial that presents the same cases, prompts, and scoring rubrics to o1 and to a concurrently enrolled panel of board-certified physicians; if physicians match or exceed the model's Bond and R-IDEA scores, the 'superhuman' claim would be refuted.","tokens_in":23956,"feed_emoji":"🩺","tokens_out":7027,"duration_ms":62358,"temperature":0.7,"pith_summary":"The paper sets out to test whether a current large language model, the o1 series, has reached or passed the diagnostic and management reasoning level of board-certified physicians. It compares the model with hundreds of physicians across five vignette experiments covering differential diagnosis, documented reasoning, triage differentials, probability estimates, and management choices, then adds a blinded comparison on real emergency department cases at three decision points. The authors claim the model outperformed physicians in every experiment, with the largest margin at the earliest triage step where information is most limited. If true, this would mean the long-standing goal of computer systems that can reason through complex clinical cases has been met, and the next step is prospective trials of AI-assisted care.","feed_headline":"o1 model beats physicians on five reasoning tests","feed_subtitle":"Widest edge comes at triage, where clinicians act on the least information.","key_machinery":"The evaluation rests on three instruments: the Bond score, a 0-5 scale rating whether a differential list contains the exact or nearly exact diagnosis; the R-IDEA scale, a validated 10-point measure of documented clinical reasoning; and a battery of case sets with historical physician and GPT-4 controls from earlier studies. A blinded adjudication protocol, in which two attending physicians scored AI and human differentials without knowing the source, carries the emergency room comparison, with a locally hosted language model used only to standardize formatting. Statistical comparisons use mixed-effects models with random intercepts for case and participant.","core_discovery":"The central claim, stated on the paper's own terms, is that o1-preview and o1 display superhuman diagnostic and management reasoning. On 143 clinicopathological conference cases, the model included the correct diagnosis in 78.3% of differentials and selected a correct or helpful next test in 98.5% of scored cases. On structured reasoning documentation, it achieved a perfect R-IDEA score in 78 of 80 responses, above both attending and resident physicians from a prior study. In the emergency department study, its proportion of differentials with the exact or very close diagnosis was 65.8% at triage, 69.6% at physician evaluation, and 79.7% at admission, surpassing two attending physicians at each stage. The authors conclude that the model is ready for prospective trials.","pith_inferences":["The paper's own Figure 1 caption notes that prior models and o1 were not evaluated on the same case sets, and several comparisons use historical controls; a reader should treat the 'superhuman' label as conditional on baseline comparability until a same-case, same-prompt head-to-head trial is run.","If the model's edge is genuinely largest when information is scarce, the most promising near-term use is decision support at nurse triage, not replacement of specialist consultation, where the model's absolute advantage narrows.","Benchmark saturation may soon force the field to move from case-based accuracy to process measures such as calibration, uncertainty expression, and the value of a second opinion in changing management.","A testable extension would be to run the same physician adjudicators on both human and AI differentials for the same patients with concurrent enrollment, and to measure whether the AI's advantage changes with case rarity or patient complexity."],"forward_implications":["If the model's diagnostic edge holds, prospective randomized trials in emergency triage become the immediate next step, because that is where the largest measured gap appears.","AI-generated second opinions could be tested as a safety layer at admission and ICU-transfer decisions, where the model scored highest in absolute terms.","Clinical reasoning benchmarks will need concurrent physician baselines rather than historical controls, since prior model generations and doctors were evaluated on different case sets.","Documentation quality, as measured by R-IDEA, may change what residency training and note assessment expect from augmented physicians.","Deployment attention should shift from standalone diagnosis to human-computer interaction design, since the study measured model-only performance, not physician-plus-model teams."],"supporting_citations":[{"why":"Establishes the complex clinical case as the evaluation standard the paper uses to anchor its claims.","marker":"[1]"},{"why":"Supplies the 70 clinicopathological conference cases and prior GPT-4 differential results that o1-preview is compared against.","marker":"[12]"},{"why":"Provides the prior R-IDEA scores for GPT-4 and physicians on clinical reasoning cases used as the historical control.","marker":"[15]"},{"why":"Contributes the management-reasoning cases, rubrics, and physician/GPT-4 comparison data for the management experiment.","marker":"[17]"},{"why":"Contributes the diagnostic-reasoning cases and physician/GPT-4 comparison data for the landmark-case experiment.","marker":"[18]"},{"why":"Provides the six landmark diagnostic vignettes that were withheld from public release to protect evaluation validity.","marker":"[19]"},{"why":"Supplies the reference probability ranges and 553-practitioner human estimates for the probabilistic reasoning experiment.","marker":"[20]"},{"why":"Defines the Bond score used to rate differential diagnosis quality across the study.","marker":"[26]"},{"why":"Supplies the probabilistic reasoning task format and prior GPT-4 outputs for the probability-estimation experiment.","marker":"[29]"}],"fun_headline_variants":["o1 beats physicians on five clinical reasoning tests","LLM outperforms doctors in diagnosis and triage","o1 achieves higher diagnostic accuracy than physicians","AI model bests doctors on reasoning in emergency room"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The superhuman claim rests on the comparability of historical human and GPT-4 baselines that were collected in earlier studies with different prompts, case sets, and scoring conditions, with no concurrently enrolled physicians in most experiments.","fun_headline_variants_meta":{"raw":{"variants":["o1 beats physicians on five clinical reasoning tests","LLM outperforms doctors in diagnosis and triage","o1 achieves higher diagnostic accuracy than physicians","AI model bests doctors on reasoning in emergency room"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1548,"prompt_tokens":938,"completion_tokens":610,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":550}},"tokens_in":554,"tokens_out":610,"duration_ms":5876,"temperature":1.0,"reasoning_tokens":550,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:32:59.550567+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective, pre-registered trial that presents the same cases, prompts, and scoring rubrics to o1 and to a concurrently enrolled panel of board-certified physicians; if physicians match or exceed the model's Bond and R-IDEA scores, the 'superhuman' claim would be refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the complex clinical case as the evaluation standard the paper uses to anchor its claims."},{"cited_title":"Kanjee, B","cited_arxiv_id":null,"evidence_quote":"Supplies the 70 clinicopathological conference cases and prior GPT-4 differential results that o1-preview is compared against."},{"cited_title":"Cabral, D","cited_arxiv_id":null,"evidence_quote":"Provides the prior R-IDEA scores for GPT-4 and physicians on clinical reasoning cases used as the historical control."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the management-reasoning cases, rubrics, and physician/GPT-4 comparison data for the management experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the diagnostic-reasoning cases and physician/GPT-4 comparison data for the landmark-case experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the six landmark diagnostic vignettes that were withheld from public release to protect evaluation validity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reference probability ranges and 553-practitioner human estimates for the probabilistic reasoning experiment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the Bond score used to rate differential diagnosis quality across the study."},{"cited_title":"Rodman, T","cited_arxiv_id":null,"evidence_quote":"Supplies the probabilistic reasoning task format and prior GPT-4 outputs for the probability-estimation experiment."}],"review_version":1}