A new benchmark covering all ICD-10 HPB disease categories shows that LLMs, including specialized medical models, perform far worse on real clinical cases than on exam-style questions.
Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework.NPJ digital medicine, 7(1):102, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
A new benchmark covering all ICD-10 HPB disease categories shows that LLMs, including specialized medical models, perform far worse on real clinical cases than on exam-style questions.