Frontier reasoning LLMs scored 95 to 99 percent on 100 MRCGP-style questions, outperforming a reported GP peer average of 73 percent.
Grok 3 Beta — The Age of Reasoning Agents
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis
Frontier reasoning LLMs scored 95 to 99 percent on 100 MRCGP-style questions, outperforming a reported GP peer average of 73 percent.