Pith. sign in

REVIEW 1 cited by

Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12356 v1 pith:B5RFF6OX submitted 2025-01-21 cs.CV

classification cs.CV
keywords modelsreportsgenerationradiologyswinachievingchestgpt-2
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Radiology plays a pivotal role in modern medicine due to its non-invasive diagnostic capabilities. However, the manual generation of unstructured medical reports is time consuming and prone to errors. It creates a significant bottleneck in clinical workflows. Despite advancements in AI-generated radiology reports, challenges remain in achieving detailed and accurate report generation. In this study we have evaluated different combinations of multimodal models that integrate Computer Vision and Natural Language Processing to generate comprehensive radiology reports. We employed a pretrained Vision Transformer (ViT-B16) and a SWIN Transformer as the image encoders. The BART and GPT-2 models serve as the textual decoders. We used Chest X-ray images and reports from the IU-Xray dataset to evaluate the usability of the SWIN Transformer-BART, SWIN Transformer-GPT-2, ViT-B16-BART and ViT-B16-GPT-2 models for report generation. We aimed at finding the best combination among the models. The SWIN-BART model performs as the best-performing model among the four models achieving remarkable results in almost all the evaluation metrics like ROUGE, BLEU and BERTScore.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2

    eess.IV 2025-05 conditional novelty 3.0 of 10

    Krum aggregation produced the highest automatic text metrics for a federated ViT-GPT-2 chest X-ray report generator on IU-Xray, but margins over FedAvg and centralized training are tiny and no privacy mechanism backs ...

Pith tools