A new three-level medical benchmark shows LLM accuracy falls sharply from factual recall (up to 78%) to full clinical diagnosis (max 19%), with larger models and inference-time scaling helping most at intermediate levels.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
A new three-level medical benchmark shows LLM accuracy falls sharply from factual recall (up to 78%) to full clinical diagnosis (max 19%), with larger models and inference-time scaling helping most at intermediate levels.