Pith. sign in

REVIEW 2 cited by

A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.03948 v1 pith:PX5XEYRS submitted 2023-03-07 cs.CL

classification cs.CL
keywords metricsclinicalfaithfulnesssummarieshospitallong-formsummarizationsummary
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long-form clinical summarization of hospital admissions has real-world significance because of its potential to help both clinicians and patients. The faithfulness of summaries is critical to their safe usage in clinical settings. To better understand the limitations of abstractive systems, as well as the suitability of existing evaluation metrics, we benchmark faithfulness metrics against fine-grained human annotations for model-generated summaries of a patient's Brief Hospital Course. We create a corpus of patient hospital admissions and summaries for a cohort of HIV patients, each with complex medical histories. Annotators are presented with summaries and source notes, and asked to categorize manually highlighted summary elements (clinical entities like conditions and medications as well as actions like "following up") into one of three categories: ``Incorrect,'' ``Missing,'' and ``Not in Notes.'' We meta-evaluate a broad set of proposed faithfulness metrics and, across metrics, explore the importance of domain adaptation (e.g. the impact of in-domain pre-training and metric fine-tuning), the use of source-summary alignments, and the effects of distilling a single metric from an ensemble of pre-existing metrics. Off-the-shelf metrics with no exposure to clinical text correlate well yet overly rely on summary extractiveness. As a practical guide to long-form clinical narrative summarization, we find that most metrics correlate best to human judgments when provided with one summary sentence at a time and a minimal set of relevant source context.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Development and Validation of the Provider Documentation Summarization Quality Instrument for Large Language Models

    cs.AI 2025-01 conditional novelty 6.0 of 10

    PDSQI-9 is a 9-item instrument for rating LLM-generated clinical summaries; in a validation study with 7 physician raters and 779 summary evaluations, it showed Cronbach's alpha 0.879 and ICC 0.867, but Krippendorff's...

  2. Ontology-Constrained Generation of Domain-Specific Clinical Summaries

    cs.CL 2024-11 conditional novelty 6.0 of 10

    An ontology-guided constrained decoding method produces specialty-specific clinical summaries and lowers hallucination scores on MIMIC-III relative to greedy and beam search baselines.

Pith tools