A new 750-case EHR benchmark that jointly scores verdicts, evidence grounding, policy use, process adherence, and abstention finds that all 16 tested agents under-abstain and that defect-free accuracy reorders the leaderboard.
Fleming, Alejandro Lozano, William J
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
A new 750-case EHR benchmark that jointly scores verdicts, evidence grounding, policy use, process adherence, and abstention finds that all 16 tested agents under-abstain and that defect-free accuracy reorders the leaderboard.