Across 1,500 MIMIC-IV discharge summaries and 11 LLMs, no model exceeded 57% F1 on ICD-10 coding, with reasoning-labeled models slightly ahead of others, but the comparison is confounded by model differences.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
Across 1,500 MIMIC-IV discharge summaries and 11 LLMs, no model exceeded 57% F1 on ICD-10 coding, with reasoning-labeled models slightly ahead of others, but the comparison is confounded by model differences.