Pith. sign in

REVIEW 6 cited by

MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02730 v1 pith:DWHJCWDW submitted 2024-07-03 cs.CV cs.AI

classification cs.CVcs.AI
keywords medicallvlmshallucinationmodelstaskslargemedvhcontext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Vision Language Models (LVLMs) have recently achieved superior performance in various tasks on natural image and text data, which inspires a large amount of studies for LVLMs fine-tuning and training. Despite their advancements, there has been scant research on the robustness of these models against hallucination when fine-tuned on smaller datasets. In this study, we introduce a new benchmark dataset, the Medical Visual Hallucination Test (MedVH), to evaluate the hallucination of domain-specific LVLMs. MedVH comprises five tasks to evaluate hallucinations in LVLMs within the medical context, which includes tasks for comprehensive understanding of textual and visual input, as well as long textual response generation. Our extensive experiments with both general and medical LVLMs reveal that, although medical LVLMs demonstrate promising performance on standard medical tasks, they are particularly susceptible to hallucinations, often more so than the general models, raising significant concerns about the reliability of these domain-specific models. For medical LVLMs to be truly valuable in real-world applications, they must not only accurately integrate medical knowledge but also maintain robust reasoning abilities to prevent hallucination. Our work paves the way for future evaluations of these studies.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs

    cs.CV 2026-07 accept novelty 6.0 of 10

    Medical VLM attention and saliency heatmaps are not causally faithful: they miss radiologist-annotated regions and anti-correlate with patch-occlusion importance, unlike CXR classifier baselines.

  2. Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

    cs.LG 2025-07 reject novelty 6.0 of 10

    A dynamic red-teaming audit reports that 94% of MedQA-correct answers fail under adversarial mutation, with 86% privacy leak rates, 81% bias shift rates, and 66-74% hallucination rates across 15 medical LLMs.

  3. MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.

  4. Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A3Tune aligns the visual attention of medical LVLMs to prompt-relevant regions via SAM and BioMedCLIP weak labels plus a Mixture-of-Experts over LoRA, improving VQA and report generation accuracy.

  5. TerraMAE: Learning Spatial-Spectral Representations from Hyperspectral Earth Observation Data via Adaptive Masked Autoencoders

    cs.CV 2025-08 reject novelty 5.0 of 10

    The abstract proposes TerraMAE, an adaptive channel-grouping masked autoencoder for hyperspectral Earth observation, but the manuscript body is a different paper, leaving the proposal without any supporting method or ...

  6. Trustworthy Medical Imaging with Large Language Models: A Study of Hallucinations Across Modalities

    eess.IV 2025-08 conditional novelty 4.0 of 10

    AI models hallucinate when reading medical images and when generating them from text, producing false findings and anatomically impossible pictures.

Pith tools