Pith. sign in

REVIEW 1 cited by

CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.17265 v1 pith:S4MVFVYM submitted 2025-05-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords caseinformationpromptingreportsextractionclinicaldatasetdense
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Rare diseases, including Inborn Errors of Metabolism (IEM), pose significant diagnostic challenges. Case reports serve as key but computationally underutilized resources to inform diagnosis. Clinical dense information extraction refers to organizing medical information into structured predefined categories. Large Language Models (LLMs) may enable scalable information extraction from case reports but are rarely evaluated for this task. We introduce CaseReportBench, an expert-annotated dataset for dense information extraction of case reports, focusing on IEMs. Using this dataset, we assess various models and prompting strategies, introducing novel approaches such as category-specific prompting and subheading-filtered data integration. Zero-shot chain-of-thought prompting offers little advantage over standard zero-shot prompting. Category-specific prompting improves alignment with the benchmark. The open-source model Qwen2.5-7B outperforms GPT-4o for this task. Our clinician evaluations show that LLMs can extract clinically relevant details from case reports, supporting rare disease diagnosis and management. We also highlight areas for improvement, such as LLMs' limitations in recognizing negative findings important for differential diagnosis. This work advances LLM-driven clinical natural language processing and paves the way for scalable medical AI applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Infherno: End-to-end Agent-based FHIR Resource Synthesis from Free-form Clinical Notes

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Infherno deploys LLM agents with code execution and terminology tools to synthesize FHIR resources from unstructured clinical notes, matching human baseline performance on synthetic and real datasets.

Pith tools