Pith. sign in

REVIEW 7 cited by

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18947 v4 pith:YI7DUMWT submitted 2024-12-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords medicalhallucinationsmllmsmodelsapplicationsbenchmarkframeworkhallucination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical Large Language Models (MLLMs) have demonstrated potential in healthcare applications, yet their propensity for hallucinations -- generating medically implausible or inaccurate information -- presents substantial risks to patient care. This paper introduces MedHallBench, a comprehensive benchmark framework for evaluating and mitigating hallucinations in MLLMs. Our methodology integrates expert-validated medical case scenarios with established medical databases to create a robust evaluation dataset. The framework employs a sophisticated measurement system that combines automated ACHMI (Automatic Caption Hallucination Measurement in Medical Imaging) scoring with rigorous clinical expert evaluations and utilizes reinforcement learning methods to achieve automatic annotation. Through an optimized reinforcement learning from human feedback (RLHF) training pipeline specifically designed for medical applications, MedHallBench enables thorough evaluation of MLLMs across diverse clinical contexts while maintaining stringent accuracy standards. We conducted comparative experiments involving various models, utilizing the benchmark to establish a baseline for widely adopted large language models (LLMs). Our findings indicate that ACHMI provides a more nuanced understanding of the effects of hallucinations compared to traditional metrics, thereby highlighting its advantages in hallucination assessment. This research establishes a foundational framework for enhancing MLLMs' reliability in healthcare settings and presents actionable strategies for addressing the critical challenge of AI hallucinations in medical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models

    cs.LG 2026-04 conditional novelty 6.5 of 10

    At matched accuracy, 85 of 210 open-weight LLM pairs have disjoint severity-tail slopes, so error rate alone cannot rank catastrophic-failure risk.

  2. KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

    cs.AI 2026-08 conditional novelty 6.0 of 10

    KnowHal is a new benchmark that jointly tests entity, attribute, relation, and knowledge hallucinations in multimodal language models using paired true/false questions on shared images.

  3. TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders

    cs.CV 2025-08 reject novelty 5.0 of 10

    The abstract proposes TerraMAE, an adaptive channel-grouping masked autoencoder for hyperspectral Earth observation, but the manuscript body is a different paper, leaving the proposal without any supporting method or ...

  4. A Multi-Task Evaluation of LLMs' Processing of Academic Text Input

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    The abstract reports Gemini underperforms on four academic text tasks, but the attached full text is an unrelated biomedical retrieval paper, leaving the claims unverifiable.

  5. Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

    eess.IV 2025-08 conditional novelty 4.0 of 10

    AI models hallucinate when reading medical images and when generating them from text, producing false findings and anatomically impossible pictures.

  6. Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models

    cs.CL 2025-06 reject novelty 3.0 of 10

    A survey of LLM hallucination research that formalizes hallucination types and argues, via incompleteness and undecidability arguments, that hallucinations cannot be fully eliminated.

  7. A comprehensive taxonomy of hallucinations in Large Language Models

    cs.CL 2025-08 conditional novelty 2.0 of 10

    A survey that organizes LLM hallucination types, causes, benchmarks, and mitigations, and restates the theorem that hallucination is inevitable for computable LLMs.

Pith tools