Pith. sign in

REVIEW 3 major objections 6 minor 27 references

Once refusals are set aside, Claude Fable 5 meets or beats every baseline on every biomedical benchmark tested; willingness, not skill, is the bottleneck.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 08:47 UTC pith:7EFKDMU3

load-bearing objection Solid multi-benchmark refusal audit of Fable 5; the willingness-vs-capability story is real and useful, but the headline “exceeds every model on every benchmark” overreaches because scored-subset accuracy is not matched to the same items. the 3 major comments →

arxiv 2607.10849 v1 pith:7EFKDMU3 submitted 2026-07-12 cs.CL

Capabilities of Claude Fable 5 on Biomedical Challenge Problems

classification cs.CL
keywords Claude Fable 5biomedical benchmarksrefusal behaviorscored-subset accuracyMedQARareBenchmultimodal medical VQAdeterministic scoring
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most published scores for frontier medical language models are clouded by two problems: old exams that almost every top model already aces, and open answers graded by other models rather than fixed keys. This paper re-evaluates Claude Fable 5, Anthropic’s strongest public model, on eight harder text and image biomedical tasks, scoring every answer deterministically against fixed keys and treating “I will not answer” as its own outcome rather than a wrong answer. The result is that Fable 5 refuses between 8% and 99% of items depending on the dataset, a pattern almost absent in its two Claude predecessors and in GPT-5. On the items it does answer, its accuracy meets or exceeds every other model on every benchmark. Two separate refusal patterns appear: one that clusters in basic-science and mechanism questions, and another that near-universally blocks certain rare pediatric metabolic disease cases while allowing adult autoimmune ones. The practical message is that the main limit on using Fable 5 for biomedical work is whether it will engage, not how well it reasons once it does.

Core claim

When refusal is recorded as a distinct outcome and excluded from the accuracy denominator, Claude Fable 5’s scored accuracy meets or exceeds every compared model (two Claude predecessors and GPT-5) on every one of eight biomedical benchmarks. The large raw-score gaps that make Fable 5 look weaker are almost entirely produced by selective refusal rates of 8–99%, not by lower competence on the questions it answers. Two distinguishable refusal patterns, not a single uniform cause, drive the behavior: concentration on basic-science and mechanism content, and a disease-domain pattern on rare-disease diagnosis.

What carries the argument

Scored-subset accuracy with refusal as a first-class outcome: every response is classified by API stop reason plus fixed keyword patterns as either a refusal or an answer; accuracy is then computed only over answered items against fixed answer keys, never by another language model. This separation is what converts apparent capability regression into a willingness gap.

Load-bearing premise

The paper’s central claim rests on the assumption that its fixed keyword patterns and API stop-reason flags correctly and completely identify true refusals, and that empty or blocked responses are not silent routing to another model, so that the answered subset really measures underlying capability.

What would settle it

Re-run the same eight benchmarks with independent refusal detection (human labels on a large random sample of “refused” and “answered” responses, plus forced multi-run stability checks) and show that scored-subset accuracy no longer leads the other models once misclassified or unstable items are corrected.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Benchmark leaderboards that fold refusals into incorrect answers will systematically understate Fable 5’s biomedical competence and mis-rank it against models that answer more freely.
  • Practitioners can treat Fable 5 as high-accuracy when it engages, and design around the refusal wall via case selection or fallback routing rather than assuming a skill deficit.
  • Simple prompt rewrites that remove clinical-role framing recover only a small fraction of refused rare-disease cases, so prompt engineering is not a reliable fix for the observed pattern.
  • Generational progress inside the Claude family on these tasks is real once willingness is separated from correctness; the new release-to-release change is selective refusal, not reduced knowledge.
  • Any future evaluation of safety-tuned biomedical models should report raw accuracy, refusal rate, and scored-subset accuracy side by side, or the central trade-off will stay invisible.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the same refusal wall appears on live clinical or research queries, deployment will need an explicit “refused → route elsewhere” path; accuracy alone will not predict usefulness.
  • The split between basic-science refusals and rare pediatric metabolic refusals suggests multiple independent filters rather than one dual-use classifier, which future black-box probes could try to separate by content type.
  • Because format (multiple-choice vs open-ended) did not predict refusal while content domain did, content-aware sampling of evaluation sets will matter more than task format for measuring this model’s true headroom.
  • Other model families that add aggressive biomedical safety layers may show the same raw-score illusion; the evaluation recipe here (fixed keys + refusal tracking) is portable beyond Claude.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper evaluates Claude Fable 5 against Claude Opus 4.6, Opus 4.8, and GPT-5 on eight biomedical benchmarks (MedQA, PubMedQA, MedXpertQA Text/MM, RareBench, VQA-RAD, SLAKE, PathVQA) under matched zero-shot direct-answer prompting with deterministic answer-key scoring and Wilson CIs. Refusal is logged as a distinct outcome rather than folded into error. The central claim is that Fable 5 refuses 8.0–99.4% of items depending on the benchmark—unlike the near-zero refusal rates of the baselines—and that once refused items are excluded from the denominator, its scored-subset accuracy meets or exceeds every other model on every benchmark, so the binding constraint is willingness to engage rather than capability. The authors further report two distinguishable refusal patterns: concentration in basic-science/mechanism content on MedQA and MedXpertQA MM (confirmed with each benchmark’s own category labels), and a disease-domain pattern on RareBench (near-universal refusal of inborn metabolic presentations vs. higher engagement on adult-onset autoimmune cases).

Significance. If the capability ranking survives proper matched controls, the paper would be a useful contribution to biomedical LLM evaluation. Treating refusal as a first-class outcome, using fixed answer keys rather than LLM-as-judge grading, and documenting generational change within one model family under an identical protocol are genuine methodological strengths. The dual-pattern refusal analysis (basic-science vs. disease-domain), with explicit negative results for several alternative hypotheses on RareBench (Table 4) and independent category schemes on two benchmarks, is more careful than typical leaderboard reports. The practical implication—that prompt reframing recovers only ~9% of RareBench refusals while answered accuracy remains high—would matter for deployment design. Those strengths are real even if the headline ranking claim needs tightening.

major comments (3)
  1. [§3.1, Tables 2 and 6, Abstract] Abstract, §3.1, Tables 2 and 6, and the Conclusion assert that scored-subset accuracy (refusals excluded) “exceeds or meets every other model on every benchmark.” Baselines are scored on essentially the full item set (0–0.4% refusal on text), while Fable 5 is scored only on the self-selected items it answered. That is not a matched comparison. If refused items are systematically harder, the answered subset is easier and the ranking can be an artifact of selection. The paper gives partial evidence against pure difficulty selection on MedQA (83.3% of Fable’s misses were items all three others got right) but never reports the direct control: baseline accuracy restricted to Fable’s answered items. Without those numbers, the load-bearing claim that the constraint is willingness rather than capability is not secured for every benchmark.
  2. [§3.2, Table 3] On RareBench under the primary prompt (Table 3), Fable 5 answers only 6 of 1,122 items (99.4% refusal). The paper still folds this case into the universal claim that scored accuracy meets or exceeds every other model. With n=6, Recall@1 of 0.36% is not a capability estimate, as the authors themselves note elsewhere, and cannot support a cross-model ranking. Either drop RareBench from the “every benchmark” capability claim or restrict that claim to benchmarks with non-trivial scored n, and treat RareBench as a refusal-pattern result only (as the neutral-prompt re-run in Table 5 already usefully does).
  3. [§2.5, §3.1.1] Scored-subset validity rests on the refusal detector (API stop-reason plus a fixed keyword list; §2.5) and on the claim that refused items are not silently routed to another model (§3.1.1). The model-identifier check (100% claude-fable-5 on 1,945 refused items) is good evidence against fallback routing, but the keyword patterns themselves are not validated (no inter-annotator sample, no false-positive/false-negative rates, no sensitivity analysis). Misclassifying hedged or truncated answers as refusals would inflate scored accuracy. A short human audit on a stratified sample of “refused” and “answered” outputs, or release of the exact keyword list and stop-reason mapping, is needed before scored-subset accuracy can be treated as a clean capability measure.
minor comments (6)
  1. [§2.5, Table 6, Appendix A] Open-ended VQA scoring uses exact match or substring containment after normalization (§2.5). Substring rules can credit partial or over-specific answers (e.g., PathVQA “uterus with leiomyoma” vs. ground truth “uterus” in Fig. A5). Report agreement under a stricter exact-match-only rule, or note how often substring (vs. exact) drove the score.
  2. [Figure 1, Figure 2] Figure 1 annotates refusal rates; the overview figure in the abstract block mixes raw accuracy with red refusal callouts. Align axis labels and “refused = wrong” vs. scored-subset conventions across Figure 1, Figure 2, and the overview panel so readers do not confuse raw and scored numbers.
  3. [Table 2] Table 2 reports Fable scored accuracy with CIs but marks predecessors with “– (x% ref.)” rather than repeating raw = scored when refusal is ~0. A uniform two-column layout (raw / scored) for all models would make the comparison easier to read.
  4. [§7] Limitations (§7) correctly note single-run evaluation and no CoT/few-shot. Given that refusal is the central phenomenon, a short multi-seed or temperature-0 re-query on a refusal subsample would strengthen the point estimates; if cost precludes it, state that more explicitly as a stability caveat on the 8–99% range.
  5. [Appendix A] Appendix Table 7 and Figures A1–A8 are helpful qualitative illustrations but are selected for clarity, not representativeness (as stated). Consider moving the uniqueness counts into the main text near §3.7 and keeping only 2–3 examples in the appendix to reduce length.
  6. [§7, §3.3, overview figure, Table 2] Minor typos/consistency: “Nori at al.” → “Nori et al.” (§7); “Beh c ¸et” encoding artifact (§3.3); “Fable 5 Opus 4.8…” overview labels lack separators; MedXpertQA Text GPT-5 raw is 54.6% in Table 2 vs. 54.7% in the overview figure.

Circularity Check

0 steps flagged

No circularity: purely empirical measurement against fixed external keys with no fitted parameters re-predicted or self-definitional loops.

full rationale

The paper reports zero-shot accuracy and refusal rates of Claude Fable 5 (plus three baselines) on eight external biomedical benchmarks, scored deterministically against fixed answer keys. Refusal is flagged by API stop-reason plus fixed keyword patterns and excluded from the scored-subset denominator; raw accuracy is also reported. No quantity is defined in terms of another quantity that is then re-derived, no parameters are fitted to a subset and then called a prediction, and no uniqueness theorem or ansatz is imported via self-citation. Author self-citations are absent; all references are to independent prior benchmarks or methods. The central claim (scored-subset accuracy meets or exceeds baselines once refusals are removed) is an empirical observation, not a derivation that reduces to its inputs by construction. Methodological concerns about unequal item sets (scored-subset vs full-set baselines) are selection-bias issues, not circularity. The evaluation is therefore self-contained against external benchmarks and exhibits zero circular steps.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 0 invented entities

Empirical black-box evaluation rests on a small set of operational choices (refusal detectors, matching cutoffs, single-shot protocol) rather than free physical parameters or invented theoretical entities. The central claim inherits the usual domain assumptions of medical-LLM benchmarking plus the paper-specific decision to treat refusal as a distinct outcome.

free parameters (2)
  • RareBench fuzzy-match cutoff = 0.6
    Diagnosis name matching uses exact, then substring, then fuzzy match with cutoff 0.6; the numeric threshold is chosen by the authors and affects Recall@k on the six scored items.
  • Image resize threshold = 1568 px
    Images larger than 1,568 pixels on either side are resized before base64 encoding; the pixel limit is an author-chosen preprocessing constant.
axioms (4)
  • ad hoc to paper API stop-reason 'refusal' plus a fixed keyword list correctly partitions model outputs into refused versus answered classes across all eight benchmarks.
    Section 2.5 defines the refusal detector; the entire scored-subset analysis rests on it.
  • domain assumption Zero-shot direct-answer prompting without system prompts or chain-of-thought is a fair and representative evaluation of biomedical capability.
    Section 2.3 and the Nori et al. template; few-shot and CoT are left to future work on cost grounds.
  • domain assumption Fixed answer keys and substring/exact matching constitute ground truth for open-ended VQA and RareBench diagnoses.
    Standard deterministic scoring assumption stated in Section 2.5; no human adjudication of free-text answers.
  • ad hoc to paper The API-reported model identifier reliably indicates which model produced the response (no silent fallback).
    Section 3.1.1 uses the identifier field to rule out routing to Opus 4.8 on refused items.

pith-pipeline@v1.1.0-grok45 · 20418 in / 2721 out tokens · 38781 ms · 2026-07-14T08:47:14.615740+00:00 · methodology

0 comments
read the original abstract

Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.

Figures

Figures reproduced from arXiv: 2607.10849 by Dominic Okonkwo, Magnus Hodgson, Susan Adanna Ihejirika, Temitope I. David.

Figure 1
Figure 1. Figure 1: Refusal rates across all evaluated biomedical benchmarks. each model actually answered for the text benchmarks. On MedQA, Fable 5’s scored-subset accuracy (96.6% [95.3, 97.5]) matches or exceeds every other model’s raw accuracy (94.3–96.0%). On PubMedQA, its scored-subset accuracy (81.3% [78.4, 83.8]) exceeds every other model’s raw accuracy (71.7–76.6%) by a clear margin. On MedXpertQA Text, its scored-su… view at source ↗
Figure 3
Figure 3. Figure 3: Refusal outcome flow, from query submission to final scoring status. from query submission to final scoring status. The other three models refused none. Opus 4.8, Opus 4.6, and GPT-5 each engaged with the full task, with Recall@1 ranging from 20.9% to 27.7%, consistent with prior published ranges for non-retrieval-augmented models on this benchmark. Fable 5’s Recall@1 of 0.36% [0.1, 0.9], computed on 6 res… view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of raw accuracy (refusals counted as incorrect) and scored-subset accuracy (refusals excluded from the denominator) across biomedical benchmarks. PubMedQA (34.6% uniquely missed) and further on MedXpertQA Text (13.0%), where Fable 5’s remaining misses are shared with the other models and look more like genuinely hard items. 3.1.1. Refusal does not route to a fallback model Anthropic’s release do… view at source ↗
Figure 4
Figure 4. Figure 4: shows refusal rates broken down by content category for both benchmarks. Close reading of refused MedXpertQA MM questions confirms the content pattern: many use a clinical-vignette wrapper around what is fundamentally a mechanism, anatomy, or physiology question, antibody structure, receptor pharmacology, cardiac electrophysiology, rather than a patient-specific diagnostic judgment. 0 5 10 15 20 25 30 35 F… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

27 extracted references · 1 canonical work pages

  1. [1]

    sounds right

    Introduction Biomedical evaluation of language models has not kept pace with the models themselves. Every new frontier model claims stronger scientific reasoning, and every new model report leads with a wall of benchmark scores. But two specific practices have quietly undermined what those scores mean. First, models have gotten good enough that many estab...

  2. [2]

    ANSWER: <letter>

    Methodology 2.1. Models We evaluate four models: Claude Fable 5 (the primary subject), Claude Opus 4.6 and Opus 4.8 as within-family predecessors, and GPT-5 from OpenAI as an external baseline. We access all Claude models via the Anthropic Messages API and GPT-5 via the OpenAI Chat Completions API. All four models are vision-capable, which the multimodal ...

  3. [3]

    to human-readable symptom names and ask for ten ranked candidate diagnoses under the following prompt shown here. Following the discovery of near-total refusal under this prompt (Section 3.2), we ran a secondary neutral-framing condition on all originally-refused items to test whether the clinical role framing was a contributing factor. This condition is ...

  4. [4]

    Respond in exactly this format: TOP10:

    <disease name> RareBench - Neutral Prompt A patient has the following clinical features: {symptom str} List the ten most likely diagnoses, ranked from most to least likely. Respond in exactly this format: TOP10:

  5. [5]

    Scoring and Refusal Handling We score all benchmarks deterministically against a fixed answer key; no benchmark in this study uses LLM-based judgment

    <diagnosis> 2.5. Scoring and Refusal Handling We score all benchmarks deterministically against a fixed answer key; no benchmark in this study uses LLM-based judgment. We report 95% Wilson confidence intervals on every percentage. Every API call logs the raw response text, API-reported model identifier, stop reason, token counts, and any error string befo...

  6. [6]

    Basic Science

    Results 3.1. Refusal Suppresses Fable 5’s Raw Scores Across all eight benchmarks, Fable 5’s raw accuracy understates its capability. The other three models, Opus 4.6, Opus 4.8, and GPT-5, refuse between 0.0% and 0.4% of questions on the text benchmarks and show comparable rates on most multimodal benchmarks. Fable 5 does not follow this pattern. It refuse...

  7. [7]

    What looked like capability decline was mostly refusal The raw scores alone tell a misleading story

    Discussion 4.1. What looked like capability decline was mostly refusal The raw scores alone tell a misleading story. On MedQA and PubMedQA, Fable 5 scores 10 to 15 percentage points below both predecessors and GPT-5 in raw accuracy, a result that, taken at face value, suggests the newest model in the family is weaker on basic medical knowledge than the on...

  8. [8]

    why test a general model at all

    Related Work Evaluating general-purpose language models on medical benchmarks began in earnest with Nori et al. [4], who tested GPT-4 zero-shot on USMLE-style exams and treated calibration and safety-tuning costs as first-class findings alongside accuracy. Their methodology surfaced an early tension that the field has not resolved: safety-tuning cost GPT-...

  9. [9]

    Every benchmark in this study is scored against a deterministic ground truth, with no model judgment anywhere in the pipeline

    identified and corrected more than twenty formula and runtime errors in the benchmark’s own calculator implementations, and separately showed that GPT-5.2-Thinking reaches 95–97% accuracy once given the calculator specification at inference time, with the remaining errors attributable to ground-truth issues rather than genuine reasoning failures, evidence...

  10. [10]

    Conclusion Claude Fable 5’s biomedical capability, measured on what it actually answers, is strong, stronger than its raw scores in this paper suggest. On text benchmarks, its scored-subset accuracy meets or exceeds every other model evaluated, including on MedQA and PubMedQA, where its raw numbers alone would have suggested a regression from its own pred...

  11. [11]

    Limitations Domain and prompting scope.This study covers eight benchmarks across general clinical knowledge, specialist reasoning, rare disease diagnosis, and four imaging domains, evaluated single-shot with direct-answer prompting, following the Nori at al.[4] template. It does not extend to other biomedical domains, genomics, pharmacology, mental health...

  12. [12]

    General-purpose large language models outperform specialized clinical ai tools on medical benchmarks,

    K. Vishwanath and E. K. e. a. Oermann, “General-purpose large language models outperform specialized clinical ai tools on medical benchmarks,”Nature Medicine, 2026

  13. [13]

    A survey on llm-as-a-judge,

    J. Gu et al., “A survey on llm-as-a-judge,”arXiv preprint arXiv:2411.15594, 2024. arXiv: 2411.15594 . [Online]. Available:https://arxiv.org/abs/2411.15594

  14. [14]

    Medxpertqa: Benchmarking expert-level medical reasoning and understanding,

    Y. Zuo et al., “Medxpertqa: Benchmarking expert-level medical reasoning and understanding,” inProceedings of the International Conference on Machine Learning (ICML),

  15. [15]

    Capabilities of gpt-4 on medical challenge problems,

    H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz, “Capabilities of gpt-4 on medical challenge problems,”arXiv preprint arXiv:2303.13375, 2023. arXiv: 2303.13375. [Online]. Available:https://arxiv.org/ abs/2303.13375

  16. [16]

    There is more to refusal in large language models than a single direction,

    F. Joad, M. Hawasly, S. Boughorbel, N. Durrani, and H. T. Sencar, “There is more to refusal in large language models than a single direction,”arXiv preprint arXiv:2602.02132,

  17. [17]

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams,

    D. Jin, E. Pan, N. Oufattole, W. H. Weng, H. Fang, and P. Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,”Applied Sciences, vol. 11, no. 14, p. 6421, 2021

  18. [18]

    Pubmedqa: A dataset for biomedical research question answering,

    Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, and X. Lu, “Pubmedqa: A dataset for biomedical research question answering,”arXiv preprint arXiv:1909.06146, 2019. arXiv: 1909.06146. [Online]. Available: https://arxiv.org/ abs/1909.06146

  19. [19]

    Rarebench: Can llms serve as rare diseases specialists?

    X. Chen, X. Mao, Q. Guo, L. Wang, S. Zhang, and T. Chen, “Rarebench: Can llms serve as rare diseases specialists?” InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24), 2024

  20. [20]

    A dataset of clinically generated visual questions and answers about radiology images,

    J. J. Lau, S. Gayen, A. Ben Abacha, and D. Demner-Fushman, “A dataset of clinically generated visual questions and answers about radiology images,”Scientific Data, vol. 5, p. 180 251, 2018

  21. [21]

    Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,

    B. Liu, L. M. Zhan, L. Xu, L. Ma, J. Yang, and X. M. Wu, “Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering,” inIEEE International Symposium on Biomedical Imaging (ISBI), 2021

  22. [22]

    Pathvqa: 30000+ questions for medical visual question answering,

    X. He, Y. Zhang, L. Mou, E. Xing, and P. Xie, “Pathvqa: 30000+ questions for medical visual question answering,” arXiv preprint arXiv:2003.10286, 2020. arXiv: 2003 . 10286. [Online]. Available: https://arxiv.org/abs/ 2003.10286

  23. [23]

    Gpt-5 system card,

    OpenAI, “Gpt-5 system card,”arXiv preprint arXiv:2601.03267, 2026. arXiv: 2601.03267 . [Online]. Available:https://arxiv.org/abs/2601.03267

  24. [24]

    Expansion of the human phenotype ontology (hpo) knowledge base and resources,

    S. K ¨ohler et al., “Expansion of the human phenotype ontology (hpo) knowledge base and resources,”Nucleic Acids Research, vol. 47, no. D1, pp. D1018–D1027, 2019. doi: 10.1093/nar/gky1105 [Online]. Available: https: //academic.oup.com/nar/article/47/D1/D1018/ 5160992

  25. [25]

    Large language models encode clinical knowledge,

    K. Singhal et al., “Large language models encode clinical knowledge,”Nature, vol. 620, no. 7972, pp. 172–180, 2023

  26. [26]

    Toward expert-level medical question answering with large language models,

    K. Singhal et al., “Toward expert-level medical question answering with large language models,”Nature Medicine, 2025

  27. [28]

    [Online]

    arXiv: 2603.02222 . [Online]. Available: https: //arxiv.org/abs/2603.02222 10 A. Appendix A. Appendix Table 7 reports, for each benchmark, how many items Fable 5 answered correctly while Opus 4.6, Opus 4.8, and GPT-5 all missed. This is not a claim about aggregate accuracy, Tables 2, 3, and 6 already cover that, but a qualitative look at what these items ...