Pith. sign in

REVIEW 3 cited by

Hallucinations and Key Information Extraction in Medical Texts: A Comprehensive Assessment of Open-Source Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.19061 v3 pith:TPBFWXA2 submitted 2025-04-27 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords modelsinformationllmsclinicalcomprehensiveeventshallucinationslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Clinical summarization is crucial in healthcare as it distills complex medical data into digestible information, enhancing patient understanding and care management. Large language models (LLMs) have shown significant potential in automating and improving the accuracy of such summarizations due to their advanced natural language understanding capabilities. These models are particularly applicable in the context of summarizing medical/clinical texts, where precise and concise information transfer is essential. In this paper, we investigate the effectiveness of open-source LLMs in extracting key events from discharge reports, including admission reasons, major in-hospital events, and critical follow-up actions. In addition, we also assess the prevalence of various types of hallucinations in the summaries produced by these models. Detecting hallucinations is vital as it directly influences the reliability of the information, potentially affecting patient care and treatment outcomes. We conduct comprehensive simulations to rigorously evaluate the performance of these models, further probing the accuracy and fidelity of the extracted content in clinical summarization. Our results reveal that while the LLMs (e.g., Qwen2.5 and DeepSeek-v2) perform quite well in capturing admission reasons and hospitalization events, they are generally less consistent when it comes to identifying follow-up recommendations, highlighting broader challenges in leveraging LLMs for comprehensive summarization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders

    cs.CV 2025-08 reject novelty 5.0 of 10

    The abstract proposes TerraMAE, an adaptive channel-grouping masked autoencoder for hyperspectral Earth observation, but the manuscript body is a different paper, leaving the proposal without any supporting method or ...

  2. Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

    eess.IV 2025-08 conditional novelty 4.0 of 10

    AI models hallucinate when reading medical images and when generating them from text, producing false findings and anatomically impossible pictures.

  3. Can Large Language Models Challenge CNNs in Medical Image Analysis?

    eess.IV 2025-05 conditional novelty 3.0 of 10

    CNNs outperform GPT-4o and Llama3.2-vision on X-ray, MRI, and CT classification, but an added filtering step substantially improves GPT-4o's chest X-ray accuracy.

Pith tools