REVIEW 4 cited by
EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Discharge summaries in Electronic Health Records (EHRs) are crucial for clinical decision-making, but their length and complexity make information extraction challenging, especially when dealing with accumulated summaries across multiple patient admissions. Large Language Models (LLMs) show promise in addressing this challenge by efficiently analyzing vast and complex data. Existing benchmarks, however, fall short in properly evaluating LLMs' capabilities in this context, as they typically focus on single-note information or limited topics, failing to reflect the real-world inquiries required by clinicians. To bridge this gap, we introduce EHRNoteQA, a novel benchmark built on the MIMIC-IV EHR, comprising 962 different QA pairs each linked to distinct patients' discharge summaries. Every QA pair is initially generated using GPT-4 and then manually reviewed and refined by three clinicians to ensure clinical relevance. EHRNoteQA includes questions that require information across multiple discharge summaries and covers eight diverse topics, mirroring the complexity and diversity of real clinical inquiries. We offer EHRNoteQA in two formats: open-ended and multi-choice question answering, and propose a reliable evaluation method for each. We evaluate 27 LLMs using EHRNoteQA and examine various factors affecting the model performance (e.g., the length and number of discharge summaries). Furthermore, to validate EHRNoteQA as a reliable proxy for expert evaluations in clinical practice, we measure the correlation between the LLM performance on EHRNoteQA, and the LLM performance manually evaluated by clinicians. Results show that LLM performance on EHRNoteQA have higher correlation with clinician-evaluated performance (Spearman: 0.78, Kendall: 0.62) compared to other benchmarks, demonstrating its practical relevance in evaluating LLMs in clinical settings.
Forward citations
Cited by 4 Pith papers
-
MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models
A new EHR-grounded benchmark of 100 long multi-session patient dialogues shows LLMs are consistently worse at cross-admission reasoning than at single-admission QA.
-
HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering
HypEHR is a hyperbolic embedding model for EHR data that uses Lorentzian geometry and hierarchy-aware pretraining to answer clinical questions nearly as well as large language models but with much smaller size.
-
MeDiSumQA: Patient-Oriented Question-Answer Generation from Discharge Letters
MeDiSumQA is a physician-curated benchmark of 416 patient-oriented QA pairs from MIMIC-IV discharge letters, and evaluation shows general-purpose LLMs often outperform biomedical-adapted models.
-
Evaluating LLM Reasoning in the Operations Research Domain with ORQA
ORQA is a new 1,513-question multiple-choice benchmark showing that open-source LLMs score up to 77% on operations research modeling questions, well below a 93% expert baseline.
Discussion (0). Continue with ORCID to comment.