Pith. sign in

REVIEW 5 cited by

VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.16672 v1 pith:UG7I4OOP submitted 2025-01-28 cs.AI cs.CLcs.IRcs.LO

classification cs.AIcs.CLcs.IRcs.LO
keywords verifacttextclinicalpatientagreementaverageclinicianelectronic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines retrieval-augmented generation and LLM-as-a-Judge to verify whether LLM-generated text is factually supported by a patient's medical history based on their electronic health record (EHR). To evaluate this system, we introduce VeriFact-BHC, a new dataset that decomposes Brief Hospital Course narratives from discharge summaries into a set of simple statements with clinician annotations for whether each statement is supported by the patient's EHR clinical notes. Whereas highest agreement between clinicians was 88.5%, VeriFact achieves up to 92.7% agreement when compared to a denoised and adjudicated average human clinican ground truth, suggesting that VeriFact exceeds the average clinician's ability to fact-check text against a patient's medical record. VeriFact may accelerate the development of LLM-based EHR applications by removing current evaluation bottlenecks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment

    cs.CY 2026-05 unverdicted novelty 6.0 of 10

    Scoping review of 134 studies on LLM-as-a-Judge in healthcare finds concentration in clinical decision support and NLP, frequent use of OpenAI models with prompt engineering, and moderate-to-strong human alignment whe...

  2. Verification Mirage: Mapping the Reliability Boundary of Self-Verification in Medical VQA

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Self-verification in medical VQA creates a verification mirage where verifiers exhibit high error and agreement bias on wrong answers, with reliability strongly conditioned on task type.

  3. MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts

    cs.CL 2025-09 conditional novelty 6.0 of 10

    MedFact, a new Chinese medical fact-checking benchmark, shows LLMs often detect errors but localize them poorly, and more reasoning time triggers over-criticism.

  4. Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    A pre-response classifier predicts user rejection risk for clinical LLM outputs with AUROC 0.719 over 4.5 months of deployment data by incorporating deployment-specific context.

  5. MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning

    cs.CL 2025-07 conditional novelty 4.0 of 10

    MedReadCtrl instruction-tunes LLaMA3 to control readability at 12 grade levels, reporting lower readability errors than GPT-4 and higher content scores on unseen clinical simplification.

Pith tools