Pith. sign in

REVIEW 1 cited by

Hallucination Benchmark in Medical Visual Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.05827 v2 pith:LGY7G566 submitted 2024-01-11 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords modelshallucinationansweringbenchmarkmedicalquestionvisionvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of large language and vision models (LLVMs) on vision question answering (VQA), particularly their applications in medicine (Med-VQA), has shown a great potential of realizing effective visual assistants for healthcare. However, these models are not extensively tested on the hallucination phenomenon in clinical settings. Here, we created a hallucination benchmark of medical images paired with question-answer sets and conducted a comprehensive evaluation of the state-of-the-art models. The study provides an in-depth analysis of current models' limitations and reveals the effectiveness of various prompting strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

    cs.CV 2025-07 reject novelty 5.0 of 10

    A video-based eye-gaze prompt improved report generation and diagnosis for one general-purpose vision-language model, LLaVA-OneVision, but hurt or barely helped two others, and the main comparison to medical models re...

Pith tools