A benchmark with expert-annotated reasoning steps and an LLM judge that scores reasoning by step coverage, showing strong correlation with expert evaluation.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
A benchmark with expert-annotated reasoning steps and an LLM judge that scores reasoning by step coverage, showing strong correlation with expert evaluation.