Pith. sign in

REVIEW 3 cited by

LongHealth: A Question Answering Benchmark with Long Clinical Documents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.14490 v1 pith:UOBOWZJY submitted 2024-01-25 cs.CL

classification cs.CL
keywords clinicalllmsbenchmarkinformationdocumentslonghealthpatientaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Background: Recent advancements in large language models (LLMs) offer potential benefits in healthcare, particularly in processing extensive patient records. However, existing benchmarks do not fully assess LLMs' capability in handling real-world, lengthy clinical data. Methods: We present the LongHealth benchmark, comprising 20 detailed fictional patient cases across various diseases, with each case containing 5,090 to 6,754 words. The benchmark challenges LLMs with 400 multiple-choice questions in three categories: information extraction, negation, and sorting, challenging LLMs to extract and interpret information from large clinical documents. Results: We evaluated nine open-source LLMs with a minimum of 16,000 tokens and also included OpenAI's proprietary and cost-efficient GPT-3.5 Turbo for comparison. The highest accuracy was observed for Mixtral-8x7B-Instruct-v0.1, particularly in tasks focused on information retrieval from single and multiple patient documents. However, all models struggled significantly in tasks requiring the identification of missing information, highlighting a critical area for improvement in clinical data interpretation. Conclusion: While LLMs show considerable potential for processing long clinical documents, their current accuracy levels are insufficient for reliable clinical use, especially in scenarios requiring the identification of missing information. The LongHealth benchmark provides a more realistic assessment of LLMs in a healthcare setting and highlights the need for further model refinement for safe and effective clinical application. We make the benchmark and evaluation code publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Back from the Future: Key-Value Cache Management by Counter-Causal Surprise

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Past tokens that the model can predict from their future context are evicted from the KV cache, judged by a counter-causal attention pass that reuses cached keys and values.

  2. Cartridges: Lightweight and general-purpose long context representations via self-study

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A per-corpus trained KV cache, called a Cartridge, matches full-context in-context learning quality on long-document benchmarks while using up to 38.6x less serving memory.

  3. Benchmarking Foundation Models with Multimodal Public Electronic Health Records

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A standardized MIMIC-IV benchmark comparing eight unimodal and multimodal foundation models shows multimodal inputs improve predictive performance without adding bias, while medical LVLMs underperform on length-of-sta...

Pith tools