Pith. sign in

REVIEW 3 major objections 5 minor 12 references

Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Four leading reasoning models as of May 2025 answered 95–99 of 100 UK general-practice exam questions correctly, far above the 73% average reported for the GPs and registrars who took the same set.

desk verdict A transparent, internally consistent benchmark of four reasoning models on 100 MRCGP-style questions, but the headline claim that the models 'substantially exceeded' GP performance rests on an unverifiable 73% peer average. read the letter →

arxiv 2506.02987 v1 pith:AVRVLSU7 submitted 2025-06-03 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords largelanguagemodelsartificialintelligencegenerativeAImedicaleducationprimarycarereasoningMRCGPclinicaldecisionsupport
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Four leading reasoning-model large language models as of May 2025—o3, Claude Opus 4, Grok-3, and Gemini 2.5 Pro—were given 100 randomly selected multiple-choice questions from the UK general-practice MRCGP-style question bank GP SelfTest, with full question text, images, and laboratory results, each question answered once. The paper reports total scores of 99.0%, 95.0%, 95.0%, and 95.0%, respectively, against a reported average peer score of 73.0% for GPs and GP registrars who had taken the same questions. The paper's central claim is that all four models performed very well and substantially exceeded the average human comparator, with o3 near-perfect and the other three comparable on this set. The finding matters because it extends evidence that the newest reasoning models, which produce step-by-step reasoning and have live internet access, may be able to support primary care decision-making, while still not being ready to replace clinicians. The authors note the peer average is as reported by the question bank, without the number of peer test-takers.

What carries the argument

The central object is the 'reasoning model' class of large language model—models fine-tuned to think through intermediate steps before giving a final answer—together with the question set: 100 randomly generated GP SelfTest multiple-choice questions scored one point each. The controlled procedure is the mechanism each model received the same instruction to answer as a UK GP, the full text and any images or tables, no further prompting, and each question was attempted exactly once. The resulting percentage scores are the observable that carries the comparison to the reported GP and registrar peer average.

What would settle it

Re-administer the same 100 questions to a representative sample of practising GPs and GP registrars under timed, closed-book conditions with no internet access and the same full question information; if their mean score approaches or exceeds the models' 95–99%, the claim that models substantially exceeded human performance would be refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that, as of 25 May 2025, the four leading publicly available reasoning models all perform at or near the top of MRCGP-style examination questions. o3 answered 99 of 100 questions correctly; Claude Opus 4, Grok-3, and Gemini 2.5 Pro each answered 95 correctly. The average peer score for the same questions was 73.0% as reported by GP SelfTest, so every model substantially exceeded the average human performance. The paper also concludes that the models' incorrect answers arose mainly from factual errors built into otherwise correct reasoning, and that models expressed equal confidence in correct and incorrect answers, which the author treats as a caution against full delegation to LLMs.

Load-bearing premise

The study's comparison stands on the 73.0% average peer score reported by GP SelfTest being a fair measure of how well GPs and GP registrars answer these questions, but the number and conditions of those peers are not disclosed.

Editorial extensions

If this is right

  • o3 was the only model to exceed 95%, and its single error out of 100 makes its performance near-perfect on this question set.
  • Claude Opus 4, Grok-3, and Gemini 2.5 Pro were comparable to each other at 95% in this sample, suggesting that among leading reasoning models the top-tier differences are small.
  • All four models scored well above the 73.0% reported average of GPs and GP registrars, which the paper reads as evidence that reasoning models could strengthen primary care knowledge support.
  • The two questions missed by multiple models—antibiotic duration for acute prostatitis and flying restriction guidelines—were answered correctly only by o3, pointing to specific guideline-recall gaps.
  • The paper argues that equal confidence in right and wrong answers is a patient-safety concern and supports keeping LLMs in an augmenting, not replacing, role.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the 73.0% peer figure is not representative of GP performance—the denominator is unreported and peers were self-selected, untimed, and possibly without the same 'answer as a UK GP' framing—then the headline conclusion that models 'substantially exceeded' GPs is an inference, not yet a controlled result.
  • A controlled head-to-head replication that asks GPs and registrars to answer the same 100 questions under identical information conditions and time limits would settle whether the gap is real; the author did not report such data.
  • Because the question bank sits behind a paywall, the protection against training-data contamination is stronger than for public exams; this makes the result informative but also means public archives of these questions could later contaminate future model training.
  • The near-ceiling scores on a 100-question sample mean small sample noise is non-negligible: a one-question miss changes o3's score by 1%, and the ordering of the three 95% models could plausibly shuffle with a different random draw.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a cross-sectional study in which four large language models (OpenAI o3, Claude Opus 4, Grok-3, and Gemini 2.5 Pro) answered 100 multiple-choice questions from the Royal College of General Practitioners' GP SelfTest on 25 May 2025. Each model was prompted to answer as a UK GP and was provided with the full question text and any accompanying images. The reported total scores were 99.0% for o3 and 95.0% for each of the other three models. The paper also reports that the average peer score for the same questions, as shown by GP SelfTest, was 73.0%. The author concludes that all models performed remarkably well and substantially exceeded the average performance of GPs and GP registrars who answered the same questions, and uses this to argue that reasoning models could support primary care delivery.

Significance. If the reported result were robust, the near-perfect o3 score on a standard UK primary-care examination would be a notable data point for the capabilities of reasoning LLMs, and the comparison with GP peers would be of interest to medical educators and AI researchers. The paper offers a timely snapshot of four leading models as of May 2025, and its item-level results table is a useful, checkable artifact. The protocol is clearly described in most respects: the prompt is stated, the question source is named, and the scoring rule is explicit. However, the significance of the headline finding depends entirely on the validity of the 73.0% peer average, which the paper itself acknowledges is an unverified aggregate with an unknown number of respondents and unstated conditions. Absent that comparator, the paper demonstrates high absolute performance on 100 questions but does not establish the claimed superiority over GPs and registrars.

major comments (3)
  1. [Results and Discussion] The central comparative claim rests on an unverified peer average. The Results state that the 73.0% figure is 'as reported by GP SelfTest' and that 'the total number of peers who had taken the questions was not made available by GP SelfTest.' No denominator, no response conditions (time limits, internet access, prior exposure, career stage, or incentives) are provided. The models, in contrast, received the full question text and images, no time pressure, and live internet access. This asymmetry makes the peer score an uncontrolled comparator rather than an established measure of 'the average performance of GPs and GP registrars.' The Discussion's conclusion that models 'substantially exceeded' GP performance is therefore unsupported by the evidence presented in this manuscript.
  2. [Results] No confidence intervals or significance tests accompany the comparison. Because each question was attempted once and the peer average is an aggregate over an unknown number of respondents, there is no way to assess whether the observed gap between 95-99% and 73% reflects a genuine difference or sampling/selection artifacts. The claim that the three 95% scores are 'comparable' is also made without any measure of variance; with a single attempt per question, a difference of one or two answers would change the ranking. The paper should present at least binomial confidence intervals for each model's score and a clearly qualified statement about the uncertainty of the peer mean.
  3. [Discussion (limitations)] The paper's own limitation paragraph mentions only the number of questions and the proportion of image-based items. It does not acknowledge that the provenance of the peer comparator is unknown, even though that comparator is the sole basis for the principal comparative claim. This omission is material. At minimum, the author should either obtain the peer-level data from GP SelfTest or restrict the claims to absolute model performance and describe the 73.0% figure merely as a platform-reported aggregate with unknown comparability.
minor comments (5)
  1. [Abstract vs. full text] The abstract uses 'Grok3' while the full text uses 'Grok-3'; please standardize the model name throughout.
  2. [Methods] The prompt is summarized but not quoted in full. For reproducibility, please include the exact prompt text and any model settings (for example, temperature or reasoning effort) used in each API call.
  3. [Methods] The scoring rule says 'the response had to be completely correct to receive a mark.' It would be helpful to state whether any model responses required adjudication for partially correct or ambiguously formatted answers, and if so, how that was handled.
  4. [Introduction] The paper repeatedly uses the term 'MRCGP-style examination questions' without clarifying whether GP SelfTest questions are actual past MRCGP items or original items written to mimic the exam. A sentence on provenance would help readers interpret generalizability.
  5. [References] The reference list includes several self-citations (for example references 18, 20, 28, and 31). This is not improper, but a broader set of independent sources on LLM performance in medical examinations would strengthen the related-work discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the study is an externally scored empirical benchmark with no fitted parameters and no derivation chain that reduces to its inputs.

full rationale

The paper's central numerical results are direct counts: each model answered 100 externally supplied MRCGP-style questions and was scored against correct answers provided by GP SelfTest. There is no fitted parameter, no model constructed from the data, and no equation in which an output is defined in terms of the input. The model scores (99.0%, 95.0%, 95.0%, 95.0%) are arithmetic transformations of the answer counts, and the comparator (73.0%) is a single externally reported peer average. The conclusion that the models 'substantially exceeded' GP performance depends on the validity and comparability of that peer average, which is a legitimate methodological concern about an uncontrolled, self-selected comparator with unknown denominator and administration conditions, but it is not circularity: the claim is not guaranteed by construction, and the result would not be identical to the input even if the 73.0% figure were treated as a benchmark. The paper's self-citations (references 7, 18, 20, 28, 31) provide background and prior related studies, but none supplies a load-bearing premise from which the current scores are derived, and none is invoked as a uniqueness theorem or as a framework that forces the reported outcomes. The training-data contamination risk is discussed as a limitation and is a validity threat, not a circular step. No circular step can be exhibited because there is no derivation chain to walk: the study is a direct, externally scored empirical measurement.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No fitted parameters or invented entities. The study rests on four unverified domain assumptions about the validity of the question bank, the representativeness of the random draw, the comparability of the peer average, and the absence of training-data contamination.

assumptions (4)
  • domain assumption GP SelfTest questions and RCGP-provided answer keys are a valid operationalization of MRCGP-style examination competence.
    The study scores models against this paywalled question bank as ground truth; no independent validation of question quality or answer key is provided.
  • domain assumption The Lucky Dip random selection produced a representative 100-question sample of the GP SelfTest bank.
    No seed, stratification, or comparison to the full bank is given; sampling variability is not addressed.
  • domain assumption The 73.0% peer average reported by GP SelfTest is a meaningful comparator for GP performance.
    The number of peers, their career stage, and their answering conditions are unknown; the average may reflect self-selected users rather than typical GPs.
  • domain assumption Models had not previously encountered the questions because the bank is paywalled.
    The paper argues this is unlikely but cannot verify training data or live retrieval; if false, scores would be inflated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis." pith.science (2026). https://pith.science/paper/AVRVLSU7

@misc{pith2026250602987,
  author       = {Pith},
  title        = {Pith review of: Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AVRVLSU7}},
  note         = {Machine review of arXiv:2506.02987}
}
read the original abstract

Background: Large language models (LLMs) have demonstrated substantial potential to support clinical practice. Other than Chat GPT4 and its predecessors, few LLMs, especially those of the leading and more powerful reasoning model class, have been subjected to medical specialty examination questions, including in the domain of primary care. This paper aimed to test the capabilities of leading LLMs as of May 2025 (o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro) in primary care education, specifically in answering Member of the Royal College of General Practitioners (MRCGP) style examination questions. Methods: o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro were tasked to answer 100 randomly chosen multiple choice questions from the Royal College of General Practitioners GP SelfTest on 25 May 2025. Questions included textual information, laboratory results, and clinical images. Each model was prompted to answer as a GP in the UK and was provided with full question information. Each question was attempted once by each model. Responses were scored against correct answers provided by GP SelfTest. Results: The total score of o3, Claude Opus 4, Grok3, and Gemini 2.5 Pro was 99.0%, 95.0%, 95.0%, and 95.0%, respectively. The average peer score for the same questions was 73.0%. Discussion: All models performed remarkably well, and all substantially exceeded the average performance of GPs and GP registrars who had answered the same questions. o3 demonstrated the best performance, while the performances of the other leading models were comparable with each other and were not substantially lower than that of o3. These findings strengthen the case for LLMs, particularly reasoning models, to support the delivery of primary care, especially those that have been specifically trained on primary care clinical data.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 9 canonical work pages

  1. [1]

    Conflict of interest RA declares no competing interest

    Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis Corresponding author Dr Richard C Armitage Academic Unit of Population and Lifespan Sciences, School of Medicine, University of Nottingham, Clinical Sciences Building, Nottingham City ...

  2. [3]

    and the comparison of their performance. Another strength is the unlikeliness that the models would previously have ‘seen’ the examination questions and their answers, either in their training data or via direct live internet access, because the questions lie behind a paywall. The study has limitations. One way in which it could have been strengthened is ...

  3. [7]

    Introducing Claude

    https://openai.com/index/introducing-o3-and-o4-mini/ [accessed 25 May 2025] 23 Anthropic. Introducing Claude

  4. [8]

    Grok 3 Beta — The Age of Reasoning Agents

    https://www.anthropic.com/news/claude-4 [accessed 25 May 2025] 24 xAI. Grok 3 Beta — The Age of Reasoning Agents. 19 February

  5. [9]

    Introducing Gemini 2.5 Flash, Veo 2, and updates to the Live API

    https://x.ai/news/grok-3 [accessed 25 May 2025] 25 Google AI for Developers. Introducing Gemini 2.5 Flash, Veo 2, and updates to the Live API. 25 March

  6. [10]

    Chatbot Arena LLM Leaderboard

    https://ai.google.dev/gemini-api/docs/changelog [accessed 25 May 2025] 26 LMArena. Chatbot Arena LLM Leaderboard. https://lmarena.ai/?leaderboard [accessed 25 May 2025] 27 RCGP Learning. GP SelfTest. https://elearning.rcgp.org.uk/course/index.php?categoryid=56 [accessed 25 May 2025] 28 R Armitage. Large language models must serve clinicians, not the rever...

  7. [12]

    DOI: 10.1111/jep.14250 Table 1: Results of each model’s performance Question o3 Claude Opus 4 Gemini 2.5 Pro Grok-3 1 Correct Correct Correct Correct 2 Correct Correct Correct Correct 3 Correct Correct Correct Correct 4 Correct Correct Correct Correct 5 Correct Incorrect Correct Correct 6 Correct Correct Correct Correct 7 Correct Correct Correct Correct 8...

  8. [781]

    ChatGPT for low- and middle-income countries: a Greek gift? The Lancet Regional Health – Western Pacific December 2023; 41: 100906

    DOI: 10.1016/S1473-3099(23)00290-6 10 K Lam. ChatGPT for low- and middle-income countries: a Greek gift? The Lancet Regional Health – Western Pacific December 2023; 41: 100906. DOI: 10.1016/j.lanwpc.2023.100906 11 X Wang, HM Sanders, Y Liu, et al. ChatGPT: promise and challenges for deployment in low- and middle-income countries. The Lancet Regional Healt...

Show all 12 references
  1. [2022]

    Introducing OpenAI o3 and o4-mini

    DOI: 10.48550/arXiv.2201.11903 22 OpenAI. Introducing OpenAI o3 and o4-mini. 16 April

  2. [2023]

    ChatGPT: the threats to medical education

    DOI: 10.1007/s00330-023-10213-1 7 R Armitage R. ChatGPT: the threats to medical education. Postgraduate Medical Journal 21 September 2023; 99(1176): 1130-1131. DOI: 10.1093/postmj/qgad046 8 M Liebrenz, R Schleifer, A Buadze, et al. Generating scholarly content with ChatGPT: et...

  3. [2024]

    Implications of large language models for clinical practice: ethical analysis through the Principlism framework

    DOI: 10.1016/S2589-7500(24)00216-4 31 R Armitage. Implications of large language models for clinical practice: ethical analysis through the Principlism framework. Journal of Evaluation in Clinical Practice. 01 December

  4. [2025]

    Each model was prompted to answer as a GP in the UK and was provided with full question information

    Questions included textual information, laboratory results, and clinical images. Each model was prompted to answer as a GP in the UK and was provided with full question information. Each question was attempted once by each model. Responses were scored against correct answers p...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.