A blinded expert evaluation of 620 real physician queries finds the specialized OpenEvidence tool outperforming frontier general LLMs by 25-39 percentage points on accuracy, clinical utility, source quality, verifiability, and completeness.
Rank analysis of incomplete block designs: I
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
A blinded expert evaluation of 620 real physician queries finds the specialized OpenEvidence tool outperforming frontier general LLMs by 25-39 percentage points on accuracy, clinical utility, source quality, verifiability, and completeness.