Pith. sign in

REVIEW 2 cited by

A Survey of Deep Learning-based Radiology Report Generation Using Multimodal Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12833 v2 pith:BUT7DNQ3 submitted 2024-05-21 cs.CV

classification cs.CV
keywords generationreportdatamedicalmethodsfieldinformationradiology
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic radiology report generation can alleviate the workload for physicians and minimize regional disparities in medical resources, therefore becoming an important topic in the medical image analysis field. It is a challenging task, as the computational model needs to mimic physicians to obtain information from multi-modal input data (i.e., medical images, clinical information, medical knowledge, etc.), and produce comprehensive and accurate reports. Recently, numerous works have emerged to address this issue using deep-learning-based methods, such as transformers, contrastive learning, and knowledge-base construction. This survey summarizes the key techniques developed in the most recent works and proposes a general workflow for deep-learning-based report generation with five main components, including multi-modality data acquisition, data preparation, feature learning, feature fusion and interaction, and report generation. The state-of-the-art methods for each of these components are highlighted. Additionally, we summarize the latest developments in large model-based methods and model explainability, along with public datasets, evaluation methods, current challenges, and future directions in this field. We have also conducted a quantitative comparison between different methods in the same experimental setting. This is the most up-to-date survey that focuses on multi-modality inputs and data fusion for radiology report generation. The aim is to provide comprehensive and rich information for researchers interested in automatic clinical report generation and medical image analysis, especially when using multimodal inputs, and to assist them in developing new algorithms to advance the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    BUSTR combines a descriptor-predicting vision encoder with a frozen language model to generate breast-ultrasound reports without paired image–report data, improving NLG and clinical-efficacy metrics over five baseline...

  2. Uterine Ultrasound Image Captioning Using Deep Learning Techniques

    cs.CV 2024-11 reject novelty 4.0 of 10

    A CNN-BiGRU model trained on 505 expert-annotated uterine ultrasound images reportedly outperforms unidirectional LSTM/GRU baselines on BLEU and ROUGE, though the described architecture lacks a decoding loop.

Pith tools