Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.
Xraygpt: Chest radiographs summarization using medical vision-language models
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Presents Med-HallMark benchmark, MediHall Score metric, and MediHallDetector model for hallucination detection and evaluation in medical LVLMs.
DIYHealth Suite introduces a large home-care dataset, DIYHealthGPT model with Hybrid Hyper Low-Rank Adaptation, and DIYHealthBench, claiming SOTA results on 11 tasks over general and medical baselines.
LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.
CoEV is a plug-and-play bidirectional verification method that maps text statements to visual evidence regions, assigns them to a four-quadrant factuality-grounding map, and uses this to detect and correct hallucinations in medical VLMs without retraining.
HTSC-CIF applies hierarchical task decomposition and cross-modal causal intervention to generate medical reports from images while addressing domain knowledge, alignment, and bias challenges.
Synthetic clinical demonstrations at inference time improve safety of Med-VLMs against visual and textual jailbreaks while preserving general performance on medical tasks.
M4CXR is a multi-modal large language model that performs multiple tasks in chest X-ray analysis including report generation with claimed SOTA clinical accuracy using chain-of-thought prompting.
Comparative evaluation of Qwen-VL-Max, Gemini 2.5 Pro, and Llama 4 Maverick on 100 blood report images using Sentence-BERT similarity indicates general-purpose VLMs show promise for preliminary patient-facing analysis.
citing papers explorer
-
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs
Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.
-
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models
Presents Med-HallMark benchmark, MediHall Score metric, and MediHallDetector model for hallucination detection and evaluation in medical LVLMs.
-
DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home
DIYHealth Suite introduces a large home-care dataset, DIYHealthGPT model with Hybrid Hyper Low-Rank Adaptation, and DIYHealthBench, claiming SOTA results on 11 tasks over general and medical baselines.
-
Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks
LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.
-
Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification
CoEV is a plug-and-play bidirectional verification method that maps text statements to visual evidence regions, assigns them to a four-quadrant factuality-grounding map, and uses this to detect and correct hallucinations in medical VLMs without retraining.
-
Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework
HTSC-CIF applies hierarchical task decomposition and cross-modal causal intervention to generate medical reports from images while addressing domain knowledge, alignment, and bias challenges.
-
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
Synthetic clinical demonstrations at inference time improve safety of Med-VLMs against visual and textual jailbreaks while preserving general performance on medical tasks.
-
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
M4CXR is a multi-modal large language model that performs multiple tasks in chest X-ray analysis including report generation with claimed SOTA clinical accuracy using chain-of-thought prompting.
-
Analysis of Blood Report Images Using General Purpose Vision-Language Models
Comparative evaluation of Qwen-VL-Max, Gemini 2.5 Pro, and Llama 4 Maverick on 100 blood report images using Sentence-BERT similarity indicates general-purpose VLMs show promise for preliminary patient-facing analysis.