Pith. sign in

REVIEW 3 cited by

3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19330 v1 pith:YDSAFCMD submitted 2024-09-28 cs.CV cs.AI

classification cs.CVcs.AI
keywords medicald-ct-gptreportdatagenerationradiologyreportsaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical image analysis is crucial in modern radiological diagnostics, especially given the exponential growth in medical imaging data. The demand for automated report generation systems has become increasingly urgent. While prior research has mainly focused on using machine learning and multimodal language models for 2D medical images, the generation of reports for 3D medical images has been less explored due to data scarcity and computational complexities. This paper introduces 3D-CT-GPT, a Visual Question Answering (VQA)-based medical visual language model specifically designed for generating radiology reports from 3D CT scans, particularly chest CTs. Extensive experiments on both public and private datasets demonstrate that 3D-CT-GPT significantly outperforms existing methods in terms of report accuracy and quality. Although current methods are few, including the partially open-source CT2Rep and the open-source M3D, we ensured fair comparison through appropriate data conversion and evaluation methodologies. Experimental results indicate that 3D-CT-GPT enhances diagnostic accuracy and report coherence, establishing itself as a robust solution for clinical radiology report generation. Future work will focus on expanding the dataset and further optimizing the model to enhance its performance and applicability.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Supervision For Vision-Language Modeling in 3D Computed Tomography

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A volumetric vision-language model trained jointly on classification labels and segmentation masks from three CT datasets reaches 83% AUROC on CT-RATE and shows cross-dataset zero-shot behavior.

  2. MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports

    cs.CV 2025-06 conditional novelty 6.0 of 10

    The authors introduce a 3D CT-based visual question answering benchmark with six error types and three task levels, and show that current 3D medical MLLMs perform poorly on it.

  3. HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HSENet improves 3D CT vision-language understanding by combining global and local 3D encoders with a centroid-based spatial token compressor, posting state-of-the-art results on CT-RATE and RadGenome-ChestCT.

Pith tools