REVIEW 5 major objections 5 minor 48 references
CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A two-module model that detects anatomical regions to draft radiology reports, then aligns image and text features with contrastive learning, is claimed to match full-data baselines on report generation and surpass the ConVIRT classifier…
desk verdict A low-resource radiology pipeline with honest raw numbers but comparative claims that outrun the evidence; re-benchmarking required before it can be accepted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a hybrid of local object detection and contrastive image-text alignment. For report generation, the load-bearing component is the region-selection classifier: a three-layer fully connected network that, after Faster R-CNN detects 29 anatomical regions, decides which bounding boxes should receive diagnostic sentences, so that each report is assembled from sentences written only about clinically salient regions. For classification, the load-bearing component is the contrastive alignment of radiograph and report embeddings, with ResNet-50 image features and BioClinicalBERT text features projected into a shared space and trained with multi-view, instance, and triplet contrastive losses, which lets unlabelled image-report pairs teach the model features that transfer to pneumonia detection. A single linear layer then performs the downstream classification.
What would settle it
Rerun the S&T, ADAATT, and ConVIRT baselines on the exact test split and preprocessing used for RRG-opt and R-CLIP, or rerun R-CLIP on the exact same 1% split ConVIRT was evaluated on; if the reported margins shrink or reverse on a fixed protocol, the central comparison collapses.
Extended reading notes
Core claim
The core claim is that a two-module pipeline, CRRG-CLIP, can generate chest radiology reports and classify chest radiographs at state-of-the-art level despite being trained on only 3.7% of the available data for generation and a comparable small split for classification. The generation module uses Faster R-CNN to locate 29 anatomical regions, a binary classifier to select which regions carry diagnostic sentences, and a fine-tuned GPT-2 to write a sentence per region; the classification module uses a ResNet-50 and BioClinicalBERT contrastive pair trained with multi-view, instance, and triplet contrastive losses, followed by a single linear layer for binary pneumonia classification. The authors report that the optimized generation model scores within roughly 14% of the S&T and ADAATT baselines on BLEU, METEOR, and ROUGE-L while using a fraction of their data, and outperforms GPT-4o on BLEU-2, BLEU-3, BLEU-4, and ROUGE-L. The optimized classifier is reported to surpass ConVIRT's AUC (0.848 vs 0.831) and to add an accuracy of 0.780. They also report that using generated reports rather than radiologist-written reports as the text side of contrastive training yields nearly the same classification performance, which they take as evidence that the generated reports are clinically meaningful.
Load-bearing premise
The biggest load-bearing premise is that the baseline numbers quoted from other papers were computed on the same test set with the same preprocessing as the new model's numbers, so the performance gaps are real rather than artifacts of different evaluation setups.
Editorial extensions
If this is right
- If the claims hold, report generation no longer requires massive full-dataset training: a 3.7% sample with region-level supervision can approach full-data captioning baselines.
- Generated reports prove usable as text data for contrastive classification, so the classification module can be trained without radiologist-written reports for every image.
- Contrastive pretraining on unlabelled image-report pairs can beat a supervised baseline such as ConVIRT for pneumonia detection in AUC and accuracy, suggesting label-efficient pathways for medical image classification.
- The region-selection step makes the generation process locally interpretable, because each sentence is traceable to a specific anatomical bounding box, which supports auditing of the report.
- The two-phase training schedule, with generation trained first and classification second, lets a single model serve both report drafting and diagnostic screening.
Reading between the lines
- A fair comparison between RRG-opt and GPT-4o would need identical prompt templates, decoding settings, and report sections; the paper's use of BLEU and ROUGE may understate GPT-4o's clinical language quality, since those metrics reward surface overlap rather than medical correctness.
- Because the region-selection classifier is trained on which anatomical regions have annotated sentences in the ImaGenome dataset, its notion of 'valuable' is locked to that dataset's annotation style; on other report corpora the selected regions may not match what radiologists would prioritize.
- The same contrastive backbone could be applied to other imaging-report pairs such as MRI, pathology, or ultrasound with little architectural change, but the region-level object detection would need a new anatomical atlas per modality.
- A testable extension is to feed the region-selection probabilities back into the classification module as a prior, or to measure whether R-CLIP's accuracy varies with report quality by systematically corrupting generated reports.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CRRG-CLIP, an end-to-end model that performs chest radiology report generation (RRG module: Faster R-CNN object detection, a binary region-selection classifier, and fine-tuned GPT-2) and radiograph classification (R-CLIP module: a CLIP-style image/text encoder pair with contrastive losses and a downstream linear classifier). The authors train on a small subset of MIMIC-CXR/ImaGenome data and evaluate report generation against S&T, ADAATT, and GPT-4o, and classification against ConVIRT. They claim the generation module performs comparably to S&T/ADAATT, outperforms GPT-4o on several metrics, and that the classification module significantly surpasses ConVIRT.
Significance. If the central comparative claims held, the paper would demonstrate that a multimodal pipeline trained on a few percent of public chest X-ray data can draft radiology-style reports and produce transferable features for pneumonia classification. The authors provide a public code repository and list detailed hyperparameters (Tables 3–6), which is helpful for reproducibility. However, the evaluation is not designed to support the headline claims: the baselines are not measured under the same test conditions, and the reported numeric gaps are small or negative. The claimed significance therefore rests on invalid comparisons rather than on the intrinsic utility of the proposed architecture, which builds on existing components (Faster R-CNN, GPT-2, CLIP, contrastive losses from CXR-CLIP) without introducing a new learning principle.
major comments (5)
- [Table 1, Section 5.1] The comparison between RRG-opt and S&T/ADAATT is not valid because the baselines were trained on 100% of their original datasets and evaluated on their own test sets, while RRG-opt was trained on 3.70% of a different data assembly (MIMIC-CXR reports, MIMIC-CXR-JPG images, Chest ImaGenome). The paper does not state that S&T and ADAATT were re-run on the RRG test set, nor does it describe identical tokenization, reference preprocessing, or metric scripts. The reported RRG-opt scores are 8–24% lower than S&T/ADAATT on every metric (e.g., BLEU-1 0.241 vs. 0.299); calling this 'comparable' is an interpretive claim that the evidence does not support. No error bars or significance tests are provided.
- [Table 1, Section 5.1] The abstract's statement that RRG-opt 'outperformed the GPT-4o model on BLEU-2, BLEU-3, BLEU-4, and ROUGE-L metrics' is selective and potentially misleading: GPT-4o's METEOR is 0.25 vs. RRG-opt's 0.109 and its BLEU-1 is 0.273 vs. 0.241, so the proposed model is not generally better. The evaluation protocol for GPT-4o (prompt, decoding parameters, number of reports, whether the same test set was used) is not described, so even the reported comparisons cannot be assessed.
- [Table 2, Section 5.2] The claim that R-CLIP 'significantly surpasses' ConVIRT is unsupported. The evaluation setup is ambiguous: Section 4.1 says the downstream classifier is trained and evaluated on the RSNA Pneumonia Dataset, but Table 2 and its note refer to a '1% MIMIC-CXR database'. ConVIRT's results are taken from the original publication rather than recomputed on the same test split, and the claimed superiority rests on a 0.021 AUC gap (0.852 vs. 0.831) with no error bars, confidence intervals, or statistical test.
- [Section 4.2, Section 5.2] There is an internal inconsistency in the data description that undermines the classification evaluation: the text states that MIMIC-CXR DICOM images were not used and that the downstream classifier uses RSNA Pneumonia, but then Section 4.2 refers to resizing 'images from the MIMIC-CXR dataset' and Table 2 indicates a 1% MIMIC-CXR split. This ambiguity prevents the reader from knowing which dataset was actually used for the results in Table 2.
- [Section 5.2, Table 2] The argument that R-CLIP-opt's performance is similar to R-CLIP-base (0.848 vs. 0.852 AUC) is used to claim that generated reports are comparable to radiologist-written reports. This is an indirect self-comparison: the reports used for R-CLIP-opt were produced by the RRG module, which was trained on the same radiologist reports, so similarity is partly by construction. The paper itself acknowledges in the conclusion that human evaluation of generated reports is future work, which is necessary before making this claim.
minor comments (5)
- [Section 4.3] The ConVIRT entry in the baseline list contains the placeholder text 'using .....', which is incomplete and must be filled in.
- [Section 3.2] The text says the image encoder uses 'RestNet-50'; this should be 'ResNet-50'.
- [Section 5.2] The sentence 'From Table 1, the performance of R-CLIP-base...' refers to Table 1, but the classification results are in Table 2.
- [Section 4.2] The description of report preprocessing says newline symbols were removed and back translation was used for augmentation, but it is not stated whether the same preprocessing was applied to the reference reports used for metric computation; this affects the comparability of all reported BLEU/METEOR/ROUGE scores.
- [Section 1 and 3.2] The model is described as 'unsupervised' and 'self-supervised' interchangeably, but the object detection and region selection submodules are supervised, and the CLIP backbone is fine-tuned on paired image-text data; the terminology should be made precise.
Circularity Check
No significant circularity: evaluation uses held-out data; the only self-citation is background and not load-bearing.
full rationale
The paper makes no formal derivation claim that could reduce to its own inputs. The report-generation module is trained on MIMIC-CXR/ImaGenome report-image pairs and evaluated on held-out test reports with standard metrics; the classification module trains a CLIP backbone on image-report pairs and a linear head on RSNA with a 7:1.5:1.5 split. The use of R-CLIP-opt trained on RRG-generated reports to argue that generated reports resemble radiologist reports is an empirical downstream comparison, not a constructional equivalence: R-CLIP-base and R-CLIP-opt are separately trained models whose similar accuracy is an observed result, and the downstream classifier is evaluated on held-out RSNA data. The only self-citation, ref [41] by the first author, supports a background statement about supervised classification costs and is not load-bearing. Baseline comparability concerns (S&T/ADAATT trained on 100% of their datasets versus RRG on 3.7%, and ConVIRT's published AUC versus the authors' 1% split) are validity and comparison-threat issues rather than circularity; they do not show that any quantity is defined in terms of itself or that a fitted parameter is renamed as a prediction. Therefore no circular step is identified.
Assumptions & free parameters
free parameters (4)
- pos_weight =
2 (RRG-opt)
- token_num (max generation length) =
300
- training data fraction =
3.7% for generation, 1% for classification
- text_to_text_loss_weight =
0.5
assumptions (4)
- domain assumption BLEU, METEOR, ROUGE-L and CIDEr measure radiology report quality
- domain assumption The Faster R-CNN detector trained on Chest ImaGenome identifies all clinically relevant regions
- domain assumption Image-text alignment from CLIP pretraining on MIMIC reports transfers to RSNA pneumonia classification
- domain assumption Baseline numbers from prior papers are directly comparable across datasets
Cite this review
Pith. "Pith review of CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs." pith.science (2026). https://pith.science/paper/X6KU3RYD
@misc{pith2026250101989,
author = {Pith},
title = {Pith review of: CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs},
year = {2026},
howpublished = {\url{https://pith.science/paper/X6KU3RYD}},
note = {Machine review of arXiv:2501.01989}
}
read the original abstract
The complexity of stacked imaging and the massive number of radiographs make writing radiology reports complex and inefficient. Even highly experienced radiologists struggle to maintain accuracy and consistency in interpreting radiographs under prolonged high-intensity work. To address these issues, this work proposes the CRRG-CLIP Model (Chest Radiology Report Generation and Radiograph Classification Model), an end-to-end model for automated report generation and radiograph classification. The model consists of two modules: the radiology report generation module and the radiograph classification module. The generation module uses Faster R-CNN to identify anatomical regions in radiographs, a binary classifier to select key regions, and GPT-2 to generate semantically coherent reports. The classification module uses the unsupervised Contrastive Language Image Pretraining (CLIP) model, addressing the challenges of high-cost labelled datasets and insufficient features. The results show that the generation module performs comparably to high-performance baseline models on BLEU, METEOR, and ROUGE-L metrics, and outperformed the GPT-4o model on BLEU-2, BLEU-3, BLEU-4, and ROUGE-L metrics. The classification module significantly surpasses the state-of-the-art model in AUC and Accuracy. This demonstrates that the proposed model achieves high accuracy, readability, and fluency in report generation, while multimodal contrastive training with unlabelled radiograph-report pairs enhances classification performance.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Alec, R., Jeffrey, W., Rewon, C., David, L., Dario, A., Ilya, S.: Language Models are Unsupervised Multitask Learners | Enhanced Reader. OpenAI Blog1(8), 9 (2019)
work page 2019
-
[2]
Informatics in Medicine Unlocked 24, 100557 (jan 2021)
Alfarghaly, O., Khaled, R., Elkorany, A., Helal, M., Fahmy, A.: Automated radiology report generation using conditioned transformers. Informatics in Medicine Unlocked 24, 100557 (jan 2021)
work page 2021
-
[3]
In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization
Banerjee, S., Lavie, A.: Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In: Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization. pp. 65–72 (2005)
2005
-
[4]
Medical Image Analysis58, 101539 (2019)
Chen, L., Bentley, P., Mori, K., Misawa, K., Fujiwara, M., Rueckert, D.: Self- supervised learning for medical image analysis using image context restoration. Medical Image Analysis58, 101539 (2019)
work page 2019
-
[5]
Dalla Serra, F., Jacenków, G., Deligianni, F., Dalton, J., O’Neil, A.Q.: Improving Image Representations via MoCo Pre-training for Multimodal CXR Classification. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). vol. 13413 LNCS, pp. 623–635. Springer Science and Busine...
work page 2022
-
[6]
In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2009
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: A Large-Scale Hierarchical Image Database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2009. pp. 248–255 (2009) 12 J. Xu et al
work page 2009
-
[7]
In: Proceedings of the IEEE International Conference on Computer Vision
Girshick, R.: Fast R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 1440–1448 (2015)
2015
-
[8]
In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. vol. 2016-Decem, pp. 770–778 (2016)
work page 2016
Show all 48 references
-
[9]
Fahmy, A.: A Survey of Text Similarity Approaches
H.Gomaa, W., A. Fahmy, A.: A Survey of Text Similarity Approaches. International Journal of Computer Applications68(13), 13–18 (2013)
2013
-
[10]
Proceedings of the AAAI Conference on Artificial Intelligence pp
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., Seekins, J., Mong, D., Halabi, S., Sandberg, J., Jones, R., Larson, D., Langlotz, C., Patel, B., Lungren, M., Ng, A.: Chexpert: A large chest radiograph ...
2019
-
[11]
In: Proceedings - International Symposium on Biomedical Imaging
Jacenkow, G., O’Neil, A.Q., Tsaftaris, S.A.: Indication as Prior Knowledge for Multimodal Disease Classification in Chest Radiographs with Transformers. In: Proceedings - International Symposium on Biomedical Imaging. vol. 2022-March. IEEE Computer Society (feb 2022),https://a...
2022 arXiv
-
[12]
PhysioNet (2019)
Johnson, A., Lungren, M., Peng, Y., Lu, Z., Mark, R., Berkowitz, S., Horng, S.: MIMIC-CXR-JPG - chest radiographs with structured labels (version 2.0.0). PhysioNet (2019)
2019
-
[13]
Scientific Data6(1) (2019)
Johnson, A.E., Pollard, T.J., Berkowitz, S.J., Greenbaum, N.R., Lungren, M.P., ying Deng, C., Mark, R.G., Horng, S.: MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data6(1) (2019)
2019
-
[14]
Kaggle: Rsna pneumonia detection challenge (2018),https://www.kaggle.com/ competitions/rsna-pneumonia-detection-challenge
2018
-
[15]
Respirology26(S3), 458–459 (2021)
Kentaro Ito, Keisuke Ogaki, Dongyi Xue, Jumpei Ukita, Seiwa Honda, O.H.: P16- 27: A deep learning approach using chest X-ray data for screening drug-induced interstitial lung disease. Respirology26(S3), 458–459 (2021)
2021
-
[16]
In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Kisilev, P., Sason, E., Barkan, E., Hashoul, S.: Medical image description using multi-task-loss CNN. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). vol. 10008 LNCS, pp. 121–129. Springe...
2016
-
[17]
In: Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2021)
Kurisinkel, L.J., Aw, A.T., Chen, N.F.: Coherent and Concise Radiology Report Generation via Context Specific Image Representations and Orthogonal Sentence States. In: Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...
2021
-
[18]
In: Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017
Lu, J., Xiong, C., Parikh, D., Socher, R.: Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In: Proceedings - 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017. vol. 2017-Janua, pp. 3242–3250. Institute of Electrical...
2017
-
[19]
Journal of Applied Clinical Medical Physics 21(9), 235–243 (2020).https://doi.org/10.1002/acm2.13001
Ma, S., Huang, Y., Che, X., Gu, R.: Faster RCNN-based detection of cervical spinal cord injury and disc degeneration. Journal of Applied Clinical Medical Physics 21(9), 235–243 (2020).https://doi.org/10.1002/acm2.13001
2020 doi
-
[20]
In: 1st International Conference on Learning Representations, ICLR 2013 - Workshop Track Proceedings (2013)
Mikolov, T., Chen, K., Corrado, G., Dean, J.: Efficient estimation of word representa- tions in vector space. In: 1st International Conference on Learning Representations, ICLR 2013 - Workshop Track Proceedings (2013)
2013
-
[21]
IEEE Journal of Biomedical and Health Informatics26(12), 6070–6080 (dec 2022)
Moon, J.H., Lee, H., Shin, W., Kim, Y.H., Choi, E.: Multi-Modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training. IEEE Journal of Biomedical and Health Informatics26(12), 6070–6080 (dec 2022)
2022
-
[22]
Sensors (Switzerland)19(17) (2019) Title Suppressed Due to Excessive Length 13
Nasrullah, N., Sang, J., Alam, M.S., Mateen, M., Cai, B., Hu, H.: Automated lung nodule detection and classification using deep learning combined with multiple strategies. Sensors (Switzerland)19(17) (2019) Title Suppressed Due to Excessive Length 13
2019
-
[23]
In: Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020
Ni, J., Hsu, C.N., Gentili, A., McAuley, J.: Learning visual-semantic embeddings for reporting abnormal findings on chest x-rays. In: Findings of the Association for Computational Linguistics Findings of ACL: EMNLP 2020. pp. 1954–1960 (2020)
2020
-
[24]
OpenAI: Gpt-4o contributions (2024), https://openai.com/ gpt-4o-contributions/, accessed: 2024-08-07
2024
-
[25]
Turkish Thoracic Journal21(5), 314–321 (2020)
Ovacıllı, S., Atacan, S.E., Gökgöz, G., Yüksel, M., Koç, O., Yıldız, A.N.: Interna- tional classification of the pneumoconiosis radiograph reader training in Turkey. Turkish Thoracic Journal21(5), 314–321 (2020)
2020
-
[26]
In: Proceedings of the Annual Meeting of the Association for Computational Linguistics
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: BLEU: A method for automatic evaluation of machine translation. In: Proceedings of the Annual Meeting of the Association for Computational Linguistics. vol. 2002-July, pp. 311–318 (2002)
2002
-
[27]
In: Proceedings of Machine Learning Research
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning Transferable Visual Models From Natural Language Supervision. In: Proceedings of Machine Learning Research. vol. 139, pp....
2021
-
[28]
IEEE Transactions on Pattern Analysis and Machine Intelligence39(6), 1137–1149 (2017)
Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence39(6), 1137–1149 (2017)
2017
-
[29]
European Radiology 26(10), 3654–3659 (oct 2016)
Robinson,J.W.,Brennan,P.C.,Mello-Thoms,C.,Lewis,S.J.:Reportinginstructions significantly impact false positive rates when reading chest radiographs. European Radiology 26(10), 3654–3659 (oct 2016)
2016
-
[30]
IEEE Reviews in Biomedical Engineering (2024)
Sloan, P., Clatworthy, P., Simpson, E., Mirmehdi, M.: Automated Radiology Report Generation: A Review of Recent Advances. IEEE Reviews in Biomedical Engineering (2024)
2024
-
[31]
In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Tanida, T., Müller, P., Kaissis, G., Rueckert, D.: Interactive and Explainable Region- guided Radiology Report Generation. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. vol. 2023-June, pp. 7433–
2023
-
[32]
In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Vedantam, R., Zitnick, C.L., Parikh, D.: CIDEr: Consensus-based image description evaluation. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. vol. 07-12-June, pp. 4566–4575 (2015)
2015
-
[33]
In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: Show and tell: A neural image caption generator. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. vol. 07-12-June, pp. 3156–3164 (2015)
2015
-
[34]
European Journal of Radiology129 (2020)
Visser, J.J., de Vries, M., Kors, J.A.: Assessment of actionable findings in radiology reports. European Journal of Radiology129 (2020)
2020
-
[35]
AMIA Symposium (2022)
Wang, S., Tang, L., Lin, M., Shih, G., Ding, Y., Peng, Y.: Prior Knowledge Enhances Radiology Report Generation. AMIA Symposium (2022)
2022
-
[36]
In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Wang, Z., Zhou, L., Wang, L., Li, X.: A Self-boosting Framework for Automated Radiographic Report Generation. In: Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition. pp. 2433–2442 (2021)
2021
-
[37]
Wu, J., Agu, N., Lourentzou, I., Sharma, A., Paguio, J., Yao, J.S., Dee, E.C., Mitchell, W., Kashyap, S., Giovannini, A., Celi, L.A., Syeda-Mahmood, T., Moradi, M.: Chest imagenome dataset (version 1.0.0) (2021), version 1.0.0
2021
-
[38]
IEEE Transactions on Medical ImagingXX, 1 (2024)
Wu, J., Guo, D., Wang, G., Yue, Q., Yu, H., Li, K., Zhang, S.: FPL+: Filtered Pseudo Label-based Unsupervised Cross-Modality Adaptation for 3D Medical Image Segmentation. IEEE Transactions on Medical ImagingXX, 1 (2024)
2024
-
[39]
In: Proceedings - 2023 IEEE Winter Conference on Applications of Computer Vision, WACV 2023
Wu, T.W., Huang, J.H., Lin, J., Worring, M.: Expert-defined Keywords Improve Interpretability of Retinal Image Captioning. In: Proceedings - 2023 IEEE Winter Conference on Applications of Computer Vision, WACV 2023. pp. 1859–1868 (2023) 14 J. Xu et al
2023
-
[40]
In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Xiong, Y., Du, B., Yan, P.: Reinforced Transformer for Medical Image Captioning. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics). pp. 673–680. Springer (2019)
2019
-
[41]
International Journal of Image, Graphics and Signal Processing 13(4), 33–46 (aug 2021).https://doi.org/10.5815/ijigsp.2021.04.03
Xu, J.: A Review of Self-supervised Learning Methods in the Field of Medical Image Analysis. International Journal of Image, Graphics and Signal Processing 13(4), 33–46 (aug 2021).https://doi.org/10.5815/ijigsp.2021.04.03
2021 doi
-
[42]
IEEE Access (2019)
Yan, F., Huang, X., Yao, Y., Lu, M., Li, M.: Combining LSTM and DenseNet for Automatic Annotation and Classification of Chest X-Ray Images. IEEE Access (2019)
2019
-
[43]
In: Proceedings of the Annual Meeting of the Association for Computational Linguistics
Yan, S.: Memory-aligned Knowledge Graph for Clinically Accurate Radiology Image Report Generation. In: Proceedings of the Annual Meeting of the Association for Computational Linguistics. pp. 116–122 (2022)
2022
-
[44]
In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
You, K., Gu, J., Ham, J., Park, B., Kim, J., Hong, E.K., Baek, W., Roh, B.: CXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-training. In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinforma...
2023
-
[45]
Applied Sciences (21) (2022)
Zhang, D., Ren, A., Liang, J., Liu, Q., Wang, H., Ma, Y.: Improving Medical X-ray Report Generation by Using Knowledge Graph. Applied Sciences (21) (2022)
2022
-
[46]
Medical Image Analysis (2024)
Zhang, T., Wei, D., Zhu, M., Gu, S., Zheng, Y.: Self-supervised learning for medical image data with anatomy-oriented imaging planes. Medical Image Analysis (2024)
2024
-
[47]
Ziegelmayer, S., Marka, A.W., Lenhart, N., Nehls, N., Reischl, S., Harder, F., Sauter, A., Makowski, M., Graf, M., Gawlitza, J.: Evaluation of GPT-4’s Chest X-Ray Impression Generation: A Reader Study on Performance and Perception. Journal of Medical Internet Research25(1) (20...
2023
-
[7442]
IEEE Computer Society (apr 2023)
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.