A benchmark of 1,200 clinical documents shows three LLMs preserve diagnostic uncertainty expressions less than half the time and struggle with adjacent levels.
arXiv preprint arXiv:2402.16040 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
A 22M-parameter hyperbolic model answers structured EHR questions with accuracy close to LLM-based systems (EHRXQA 89.5%, MIMIC-Instr 76.0%).
citing papers explorer
-
Possible or Definite? A Benchmark for Evaluating Diagnostic Uncertainty Preservation in Clinical Text
A benchmark of 1,200 clinical documents shows three LLMs preserve diagnostic uncertainty expressions less than half the time and struggle with adjacent levels.
-
HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering
A 22M-parameter hyperbolic model answers structured EHR questions with accuracy close to LLM-based systems (EHRXQA 89.5%, MIMIC-Instr 76.0%).