Pith. sign in

REVIEW 10 cited by

XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07971 v2 pith:VFZEBFNU submitted 2023-06-13 cs.CV

XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models

classification cs.CV
keywords medicalmodelsradiographschestmodelperformancevision-languagexraygpt
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The latest breakthroughs in large vision-language models, such as Bard and GPT-4, have showcased extraordinary abilities in performing a wide range of tasks. Such models are trained on massive datasets comprising billions of public image-text pairs with diverse tasks. However, their performance on task-specific domains, such as radiology, is still under-investigated and potentially limited due to a lack of sophistication in understanding biomedical images. On the other hand, conversational medical models have exhibited remarkable success but have mainly focused on text-based analysis. In this paper, we introduce XrayGPT, a novel conversational medical vision-language model that can analyze and answer open-ended questions about chest radiographs. Specifically, we align both medical visual encoder (MedClip) with a fine-tuned large language model (Vicuna), using a simple linear transformation. This alignment enables our model to possess exceptional visual conversation abilities, grounded in a deep understanding of radiographs and medical domain knowledge. To enhance the performance of LLMs in the medical context, we generate ~217k interactive and high-quality summaries from free-text radiology reports. These summaries serve to enhance the performance of LLMs through the fine-tuning process. Our approach opens up new avenues the research for advancing the automated analysis of chest radiographs. Our open-source demos, models, and instruction sets are available at: https://github.com/mbzuai-oryx/XrayGPT.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs

    cs.CV 2026-05 conditional novelty 7.0

    Medical VLMs frequently select negated options that contradict visible chest X-ray findings, achieving only ~30% accuracy on direct presence probes, but a post-hoc consistency verifier raises accuracy above 95%.

  2. Detecting and Evaluating Medical Hallucinations in Large Vision Language Models

    cs.CV 2024-06 unverdicted novelty 7.0

    Presents Med-HallMark benchmark, MediHall Score metric, and MediHallDetector model for hallucination detection and evaluation in medical LVLMs.

  3. DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home

    cs.CY 2026-05 unverdicted novelty 6.0

    DIYHealth Suite introduces a large home-care dataset, DIYHealthGPT model with Hybrid Hyper Low-Rank Adaptation, and DIYHealthBench, claiming SOTA results on 11 tasks over general and medical baselines.

  4. Visual Instruction-Finetuned Language Model for Versatile Brain MR Image Tasks

    cs.CV 2026-04 unverdicted novelty 6.0

    LLaBIT is a single instruction-finetuned LLM that performs report generation, VQA, segmentation, and translation on brain MRI images while outperforming task-specific models.

  5. Hallucination Detection and Correction in Medical VLMs via Counter-Evidence Verification

    cs.CV 2026-06 unverdicted novelty 5.0

    CoEV is a plug-and-play bidirectional verification method that maps text statements to visual evidence regions, assigns them to a four-quadrant factuality-grounding map, and uses this to detect and correct hallucinati...

  6. MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence

    cs.CV 2026-03 reject novelty 5.0

    MEDIC-AD adds anomaly-aware and difference tokens to a medical VLM, claiming SOTA lesion detection, temporal tracking, and visual grounding; the zero-shot claim is undermined by likely train/test overlap.

  7. Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework

    cs.CV 2025-11 unverdicted novelty 5.0

    HTSC-CIF applies hierarchical task decomposition and cross-modal causal intervention to generate medical reports from images while addressing domain knowledge, alignment, and bias challenges.

  8. Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations

    cs.CV 2025-06 unverdicted novelty 5.0

    Synthetic clinical demonstrations at inference time improve safety of Med-VLMs against visual and textual jailbreaks while preserving general performance on medical tasks.

  9. M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation

    cs.CV 2024-08 unverdicted novelty 5.0

    M4CXR is a multi-modal large language model that performs multiple tasks in chest X-ray analysis including report generation with claimed SOTA clinical accuracy using chain-of-thought prompting.

  10. Analysis of Blood Report Images Using General Purpose Vision-Language Models

    cs.CV 2025-09 unverdicted novelty 2.0

    Comparative evaluation of Qwen-VL-Max, Gemini 2.5 Pro, and Llama 4 Maverick on 100 blood report images using Sentence-BERT similarity indicates general-purpose VLMs show promise for preliminary patient-facing analysis.