One advanced language model matched or beat physician baselines on many diagnostic reasoning benchmarks, but the paper's 'superhuman in every experiment' claim is not fully supported by its own statistics.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Superhuman performance of a large language model on the reasoning tasks of a physician
One advanced language model matched or beat physician baselines on many diagnostic reasoning benchmarks, but the paper's 'superhuman in every experiment' claim is not fully supported by its own statistics.