REVIEW 3 major objections 6 minor 67 references
Chest X-ray model rankings depend heavily on whether labels come from images or radiology reports, and on which image-quality metric is used, so evaluation references should be chosen as part of clinical validity, not as an afterthought.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Chest X-ray AI model rankings and image-quality metric rankings change substantially with the choice of evaluation reference, so benchmark scores are not neutral.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection Solid, data-rich empirical study showing evaluation-reference choice can reorder CXR model rankings; the noisy image-reference is a real caveat but doesn't overturn the core message. the 3 major comments →
Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
Using 650 paired studies from a clinical cohort and 735 from a public dataset, the paper shows that image-derived and report-derived pathology labels disagree in a pathology-dependent way, and that this disagreement propagates to model rankings: for fine-tuned vision-language models, rankings under the two label sources can be nearly uncorrelated for Atelectasis (SRCC ≈ 0.21), while Pleural effusion remains stable. In image quality assessment, expert rankings of diagnostic usability correlate strongly with some full-reference metrics (HaarPSI, GMSD) but weakly with PSNR, CW-SSIM, and no-reference metrics. The authors conclude that image- and report-derived labels answer different clinical qu
What carries the argument
The central object is the paired evaluation-reference dataset XQA-Chest, with its image-only and report-only annotation protocols and an adjudicated (or automatically fused) consensus label. Rank correlation coefficients (SRCC, KRCC) and Cohen's kappa carry the argument by quantifying how label-source disagreement translates into model-ranking instability and how IQA metrics align with expert diagnostic-usability rankings.
Load-bearing premise
The paired image-derived labels (conflict-resolved on the clinical cohort, FUSE_OR-fused on the public dataset) accurately represent expert image-only judgment, and their disagreement with report-derived labels reflects systematic differences rather than idiosyncratic annotation noise.
What would settle it
An independent multi-institution annotation study with adjudicated image labels that found model rankings under image- and report-derived references to be highly correlated (e.g., SRCC > 0.9) for Atelectasis and Consolidation would directly contradict the claim that reference choice flips rankings. Alternatively, demonstrating that the ranking instability disappears when image labels are produced by a larger, more diverse panel of radiologists would indicate the effect is annotation noise rather than a structural property of the label sources.
If this is right
- Benchmark leaderboards in CXR classification are reference-dependent; a model's best status is only meaningful relative to the chosen label source.
- Evaluation against report-derived labels is internally reproducible but can diverge from image-derived evaluation; studies should state which report section (Findings vs Impression) is used.
- IQA metrics used for reconstruction or generative models should be validated against expert diagnostic-usability ratings before deployment, since PSNR and no-reference metrics can misrank outputs.
- The paired-references resource enables future work to test label-fusion methods and automated report-extraction strategies against a common expert benchmark.
Where Pith is reading between the lines
- If reference-dependence generalizes, single-reference benchmarks in medical imaging can mislead model selection; multi-reference evaluation may become standard practice for clinically meaningful comparisons.
- The low image inter-annotator agreement for some pathologies suggests image-derived labels are themselves noisy; a multi-institution adjudicated study could determine whether ranking instability is structural or partly annotation noise.
- The IQA results imply that clinically validated, task-specific quality metrics could be developed directly from expert preference data rather than borrowed from generic perceptual metrics.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper asks whether evaluation-reference choices—image-derived vs report-derived pathology labels, and generic IQA metrics vs expert ratings of diagnostic usability—change model rankings and hence scientific conclusions in chest X-ray machine learning. The authors introduce XQA-Chest, a paired dataset of expert image- and report-derived labels for 650 CUH studies, plus expert IQA ratings for 1,571 degraded images, and they annotate a 735-study MIMIC-CXR subset under similar protocols. They evaluate a broad set of supervised CNNs and vision-language models under different label references, measure agreement and rank correlations (SRCC/KRCC with bootstrap CIs), and similarly compare IQA metrics against expert z-MOS rankings. The central finding is that changing the evaluation reference can reorder models, sometimes dramatically (e.g., VLM-FT rankings for Atelectasis are nearly uncorrelated across references), and that commonly used IQA metrics such as PSNR and CW-SSIM align only weakly with expert judgments.
Significance. If the central claim holds, the paper makes an important methodological contribution: benchmark rankings in CXR machine learning are reference-dependent, and the choice of evaluation reference should be treated as a core component of clinical validity rather than an implementation detail. The paper's strengths include a new paired dataset with expert annotations, extensive agreement statistics with pessimistic/optimistic uncertainty mappings, bootstrap confidence intervals for pathology rank correlations, sensitivity analyses for label-fusion methods, and a qualitative radiology review of disagreements. These are appropriate tools for the claim and provide a useful resource for the community. The main risk is that the image-derived reference is not shown to be a stable construct independent of individual annotator noise; if the apparent reference effect is substantially driven by image-annotator disagreement, the interpretation as a systematic image-vs-report distinction would need to be tempered.
major comments (3)
- [§3.1.1, Table A.5] The conflict-resolved (CR) image labels for Atelectasis (165 positives), Cardiomegaly (21), and Support Devices (274) exactly match Img2's counts, while Img1's counts are very different (44, 153, 309). This strongly suggests that adjudication often deferred to one annotator rather than producing a genuine synthesis. The paper does not report model rankings induced by Img1 and Img2 separately. Without these controls, the central claim that changing the label source (image vs report) systematically changes rankings is confounded with image-annotator noise: the image-image rank correlation may be as low as the image-report rank correlation. For the VLM-FT group on Atelectasis, the CR-vs-report SRCC is 0.21±0.22 (Table E.26); if Img1-vs-Img2 SRCC is similar, the effect is not specific to the image/report distinction. Please add per-annotator ranking analyses for XQA-RR (and MIMIC where possi
- [§2.1.3, Table F.30] FUSE_OR is validated against CR on XQA-RR and then used as the image-derived reference on MIMIC-CXR, where no adjudication is available. This transfer is not fully supported. First, the high FUSE_OR-vs-CR agreement is partly a consequence of the CR pattern noted above: for pathologies where CR equals one annotator, an optimistic union rule will naturally reproduce the more positive annotator. Second, the MIMIC inter-annotator agreement profile differs substantially from XQA-RR (e.g., Lung opacity κ=0.11 on MIMIC vs 0.42 on XQA-RR; Table B.12 vs B.9), so a fusion rule validated on one institution's disagreement structure may behave differently elsewhere. The MIMIC comparisons labeled 'image-derived vs report-derived' are therefore better described as comparisons between a particular optimistic fusion rule and manually annotated report sections. Please add sensitivity analyses using Img1 a
- [§3.2, Table 4] The IQA rank correlations are reported as point estimates without confidence intervals, in contrast to the pathology-ranking results, which use 1000-sample bootstrap CIs. With 1,571 images, the differences among the better-performing metrics (e.g., HaarPSI 0.87, GMSD 0.86, HaarPSImed 0.85, MS-SSIM 0.84) may be within sampling variability. The paper's IQA conclusion—that metric selection determines which processed images appear best—depends on these metric-level differences being reliable. Please provide bootstrap CIs (or paired significance tests) for the SRCC/KRCC values in Table 4.
minor comments (6)
- [Abstract, §4] The abstract and Discussion state that SSIM and PSNR 'often fail to align' with expert judgment, but Table 4 gives SSIM-matlab SRCC=0.78, which is moderate agreement. Consider rephrasing to avoid overstatement, e.g., 'SSIM shows weaker alignment than top metrics, and PSNR aligns weakly.'
- [§2.1.2, Tables A.7/A.8] The handling of 'not mentioned' report labels is not stated. The tables list explicit negative, uncertain, and positive counts, while the positive prevalences in Table A.5 imply that 'not mentioned' was collapsed to negative in the binary analyses. This is a standard CheXpert-like assumption but should be stated explicitly, since it affects all report-derived agreement and ranking results.
- [Table 2] The symbols ✓/✗ in Table 2 are not defined consistently; for example, the 'Images' and 'Reports' columns use different combinations. Add a legend or use explicit 'Yes/No' entries.
- [§2.1.2] The text describes a 14-category schema and then excludes 'No finding' and 'Support devices' to obtain 12 pathologies. State this exclusions before listing the 14 categories to avoid apparent inconsistency.
- [Figures 10b–14b] The x-axis labels in Figures 10b–14b use the order 'Impression, FUSE_OR, Findings' while the text and the MIMIC tables usually list 'FUSE_OR, Findings, Impressions.' Use a consistent order for readability.
- [§4] The Discussion states that the observed divergence is 'systematic, clinically interpretable.' Given the single-institution annotation and the CR=Img2 pattern noted above, 'systematic' is too strong; please qualify this claim in light of the per-annotator analysis requested in the major comments.
Circularity Check
No significant circularity: ranking and IQA results are empirical comparisons from newly collected labels, not derived from fitted parameters or load-bearing self-citation.
full rationale
The paper's central claims are empirical: it collects paired expert image-derived and report-derived labels (XQA-RR and MIMIC-CXR subsets), evaluates a diverse set of classifiers and VLMs against each reference, and measures rank stability. No target quantity is defined in terms of its own input; no fitted parameter is renamed as a prediction. The FUSE_OR fusion rule is validated against conflict-resolved labels on XQA-RR and then applied to MIMIC-CXR; the paper explicitly acknowledges in Section 5 (Limitations) that this is "a choice validated against CUH but not equivalent to third-annotator adjudication," so this is a construct-validity caveat rather than a circular derivation. Low image inter-annotator kappa values (e.g., Atelectasis kappa = 0.15) and single-institution annotation are data-quality limitations, not evidence that conclusions are forced by construction. Self-citations exist: RadPert [17] supplies CUH fine-tuning labels, and Breger et al. [2,3] motivate IQA concerns. However, the ranking experiments and IQA correlations are recomputed from the newly annotated data and public models, so those citations are not load-bearing; the central finding would stand on the reported tables alone. The limitations that MIMIC labels come from automatic fusion rather than adjudication, and that both datasets come from one institution, affect generalizability but do not reduce the derivation to its inputs. No equation in the paper is identical to another by construction, and no 'prediction' is statistically forced by a fitted parameter. Score 1 reflects only the presence of minor, non-load-bearing self-citation; there is no material circularity.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Expert image-only annotation is a valid operationalization of what is visually recoverable from a radiograph.
- domain assumption Report-derived labels are a valid operationalization of the clinically contextualized interpretation, and Findings vs Impressions can be scored as separate references.
- ad hoc to paper FUSE_OR fusion transfers from XQA-RR to MIMIC-CXR as a proxy for adjudicated image labels.
- domain assumption ROC-AUC is a suitable cross-reference performance metric despite differing label prevalence.
- domain assumption A two-site sample with two image annotators per site is representative enough for ranking-stability conclusions.
Cite this review
Pith. "Pith review of Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance." pith.science (2026). https://pith.science/paper/3YY6HEGG
@misc{pith2026260726333,
author = {Pith},
title = {Pith review of: Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance},
year = {2026},
howpublished = {\url{https://pith.science/paper/3YY6HEGG}},
note = {Machine review of arXiv:2607.26333}
}
read the original abstract
Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.
Figures
Reference graph
Works this paper leans on
-
[1]
On model evalu- ation under non-constant class imbalance, in: International Conference on Computational Science, Springer
Brabec, J., Komárek, T., Franc, V., Machlica, L., 2020. On model evalu- ation under non-constant class imbalance, in: International Conference on Computational Science, Springer. pp. 74–87
2020
-
[2]
A study of why we need to reassess full reference image quality assessment with medical images
Breger, A., Biguri, A., Landman, M.S., Selby, I., Amberg, N., Brunner, E., Gröhl, J., Hatamikia, S., Karner, C., Ning, L., et al., 2025a. A study of why we need to reassess full reference image quality assessment with medical images. Journal of Imaging Informatics in Medicine 38, 3444–3469
-
[3]
A study on the adequacy of com- mon iqa measures for medical images, in: Su, R., Frangi, A.F., Zhang, Y
Breger, A., Karner, C., Selby, I., Gröhl, J., Dittmer, S., Lilley, E., Babar, J., Beckford, J., Else, T.R., Sadler, T.J., Shahipasand, S., Thavakumar, A., Roberts, M., Schönlieb, C.B., 2025b. A study on the adequacy of com- mon iqa measures for medical images, in: Su, R., Frangi, A.F., Zhang, Y. (Eds.), Proceedings of 2024 International Conference on Medi...
2024
-
[4]
Padch- est: A large chest x-ray image dataset with multi-label annotated reports
Bustos, A., Pertusa, A., Salinas, J.M., de la Iglesia-Vayá, M., 2020. Padch- est: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis 66, 101797. URL:https://www.sciencedirect. com/science/article/pii/S1361841520301614, doi:https://doi.org/ 10.1016/j.media.2020.101797
arXiv 2020
-
[5]
Why almost all ml models for medicine are wrong-and what we need for evidence-based medical ai
Cabitza, F., Jurman, G., Molinari, F., Bellazzi, R., 2026. Why almost all ml models for medicine are wrong-and what we need for evidence-based medical ai. International Journal of Medical Informatics , 106538
2026
-
[6]
A simple frame- work for contrastive learning of visual representations, in: International conference on machine learning, PmLR
Chen, T., Kornblith, S., Norouzi, M., Hinton, G., 2020. A simple frame- work for contrastive learning of visual representations, in: International conference on machine learning, PmLR. pp. 1597–1607
2020
-
[7]
Towards unifying medical vision-and-language pre-training via soft prompts, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp
Chen, Z., Diao, S., Wang, B., Li, G., Wan, X., 2023. Towards unifying medical vision-and-language pre-training via soft prompts, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23403–23413. 30
2023
-
[8]
A coefficient of agreement for nominal scales
Cohen, J., 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20, 37–46
1960
-
[9]
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Cohen, J., 1968. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin 70, 213
1968
-
[10]
On the limits of cross-domain generalization in automated x-ray prediction, in: Medical Imaging with Deep Learning, PMLR
Cohen, J.P., Hashir, M., Brooks, R., Bertrand, H., 2020. On the limits of cross-domain generalization in automated x-ray prediction, in: Medical Imaging with Deep Learning, PMLR. pp. 136–155
2020
-
[11]
Cohen, J.P., Viviano, J.D., Bertin, P., Morrison, P., Torabian, P., Guar- rera, M., Lungren, M.P., Chaudhari, A., Brooks, R., Hashir, M., et al.,
-
[12]
The relationship between precision-recall and roc curves, in: Proceedings of the 23rd international conference on Machine learning, pp
Davis, J., Goadrich, M., 2006. The relationship between precision-recall and roc curves, in: Proceedings of the 23rd international conference on Machine learning, pp. 233–240
2006
-
[13]
Maximum likelihood estimation of ob- server error-rates using the EM algorithm
Dawid, A.P., Skene, A.M., 1979. Maximum likelihood estimation of ob- server error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics) 28, 20–28
1979
-
[14]
Preparing a collection of radiology examinations for distribution and retrieval
Demner-Fushman, D., Kohli, M.D., Rosenman, M.B., Shooshan, S.E., Ro- driguez, L., Antani, S., Thoma, G.R., McDonald, C.J., 2015. Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association 23, 304–310
2015
-
[15]
An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Un- terthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkor- eit, J., Houlsby, N., 2021. An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations
2021
-
[16]
Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy
Efron, B., Tibshirani, R., 1986. Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical science , 54–75
1986
-
[17]
Fytas, P., Breger, A., Selby, I., Baker, S., Shahipasand, S., Korhonen, A.,
-
[18]
Digital Image Processing
Gonzalez, R.C., Woods, R.E., 2008. Digital Image Processing. 3rd ed., Pearson
2008
-
[19]
Evaluating the robustness and readiness of large frontier models in health ai applications
Gu, Y., Fu, J., Liu, X., Valanarasu, J.M.J., Codella, N.C., Tan, R., Liu, Q., Jin, Y., Zhang, S., Wang, J., et al., 2026. Evaluating the robustness and readiness of large frontier models in health ai applications. Nature Medicine , 1–9. 31
2026
-
[20]
Robust classification from noisy labels: Integrating additional knowledge for chest radiography abnormality assessment
Gündel, S., Setio, A.A., Ghesu, F.C., Grbic, S., Georgescu, B., Maier, A., Comaniciu, D., 2021. Robust classification from noisy labels: Integrating additional knowledge for chest radiography abnormality assessment. Med- ical Image Analysis 72, 102087
2021
-
[21]
The meaning and use of the area under a receiver operating characteristic (roc) curve
Hanley, J.A., McNeil, B.J., 1982. The meaning and use of the area under a receiver operating characteristic (roc) curve. Radiology 143, 29–36
1982
-
[22]
Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp
He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778
2016
-
[23]
Hovy, D., Berg-Kirkpatrick, T., Vaswani, A., Hovy, E., 2013. Learning whom to trust with MACE, in: Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp. 1120–1130
2013
-
[24]
Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp
Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q., 2017. Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708
2017
-
[25]
Huang, S.C., Shen, L., Lungren, M.P., Yeung, S., 2021. Gloria: A mul- timodal global-local representation learning framework for label-efficient medical image recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3942–3951
2021
-
[26]
Chexpert: A large chest radiograph dataset with uncertainty labels and expert compar- ison
Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R.L., Shpanskaya, K.S., Seekins, J., Mong, D.A., Halabi, S.S., Sandberg, J.K., Jones, R., Larson, D.B., Lan- glotz, C.P., Patel, B.N., Lungren, M.P., Ng, A.Y., 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert compar- i...
Pith/arXiv arXiv 2019
-
[27]
Jain, S., Smit, A., Truong, S.Q., Nguyen, C.D., Huynh, M.T., Jain, M., Young, V.A., Ng, A.Y., Lungren, M.P., Rajpurkar, P., 2021. Visu- alchexbert: addressing the discrepancy between radiology report labels and image labels, in: Proceedings of the Conference on Health, Infer- ence, and Learning, Association for Computing Machinery, New York, NY, USA. p. 1...
arXiv 2021
-
[28]
Johnson, A.E.W., Pollard, T., Mark, R., Berkowitz, S., Horng, S.,
-
[29]
Karner, C., Gröhl, J., Selby, I., Babar, J., Beckford, J., Else, T.R., Sadler, T.J., Shahipasand, S., Thavakumar, A., Roberts, M., Rudd, J.H., Schön- lieb, C.B., Weir-McCall, J.R., Breger, A., 2025. Parameter choices in 32 haarpsiforiqawithmedicalimages, in: 2025IEEE22ndInternationalSym- posium on Biomedical Imaging (ISBI), pp. 1–5. doi:10.1109/ISBI60581....
arXiv 2025
-
[30]
A new measure of rank corre- lation
KENDALL, M.G., 1938. A new measure of rank corre- lation. Biometrika 30, 81–93. URL:https://doi.org/ 10.1093/biomet/30.1-2.81, doi:10.1093/biomet/30.1-2.81, arXiv:https://academic.oup.com/biomet/article-pdf/30/1-2/81/423380/30-1-2-81.pdf
-
[31]
Klie, J.C., Bugert, M., Boullosa, B., Eckart de Castilho, R., Gurevych, I., 2018. The INCEpTION platform: Machine-assisted and knowledge- oriented interactive annotation, in: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, Santa Fe, New Mexico. pp. 5–9. URL:https://www.aclweb.org/anthology/ C18-2002
2018
-
[32]
Liu, C., Cheng, S., Chen, C., Qiao, M., Zhang, W., Shah, A., Bai, W., Ar- cucci, R., 2023. M-flag: Medical vision-language pre-training with frozen language models and latent space geometry optimization, in: Medical Im- age Computing and Computer Assisted Intervention – MICCAI 2023, pp. 637–647. doi:10.1007/978-3-031-43907-0_61
-
[33]
Chest radiograph interpretation with deep learning models: as- sessment with radiologist-adjudicated reference standards and population- adjusted evaluation
Majkowska, A., Mittal, S., Steiner, D.F., Reicher, J.J., McKinney, S.M., Duggan, G.E., Eswaran, K., Cameron Chen, P.H., Liu, Y., Kalidindi, S.R., et al., 2020. Chest radiograph interpretation with deep learning models: as- sessment with radiologist-adjudicated reference standards and population- adjusted evaluation. Radiology 294, 421–431
2020
-
[34]
completely blind
Mittal, A., Soundararajan, R., Bovik, A.C., 2012. Making a “completely blind” image quality analyzer. IEEE Signal processing letters 20, 209–212
2012
-
[35]
Vindr-cxr: An open dataset of chest x-rays with radiologist’s annotations
Nguyen, H.Q., Lam, K., Le, L.T., Pham, H.H., Tran, D.Q., Nguyen, D.B., Le, D.D., Pham, C.M., Tong, H.T., Dinh, D.H., et al., 2022. Vindr-cxr: An open dataset of chest x-rays with radiologist’s annotations. Scientific Data 9, 429
2022
-
[36]
Representation learning with con- trastive predictive coding
Oord, A.v.d., Li, Y., Vinyals, O., 2018. Representation learning with con- trastive predictive coding. arXiv preprint arXiv:1807.03748
Pith/arXiv arXiv 2018
-
[37]
Evaluation: fromprecision, recallandf-measuretoroc, informedness, markednessandcorrelation
Powers, D.M., 2020. Evaluation: fromprecision, recallandf-measuretoroc, informedness, markednessandcorrelation. arXivpreprintarXiv:2010.16061
Pith/arXiv arXiv 2020
-
[38]
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.,
-
[39]
Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists
Rajpurkar, P., Irvin, J., Ball, R.L., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C.P., et al., 2018. Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists. PLoS medicine 15, e1002686
2018
-
[40]
Reis, E.P., De Paiva, J.P., Da Silva, M.C., Ribeiro, G.A., Paiva, V.F., Bulgarelli, L., Lee, H.M., Santos, P.V., Brito, V.M., Amaral, L.T., et al.,
-
[41]
A haar wavelet-basedperceptualsimilarityindexforimagequalityassessment
Reisenhofer, R., Bosse, S., Kutyniok, G., Wiegand, T., 2018. A haar wavelet-basedperceptualsimilarityindexforimagequalityassessment. Sig- nal Processing: Image Communication 61, 33–43. doi:10.1016/j.image. 2017.11.001
doi:10.1016/j.image 2018
-
[42]
Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans
Roberts, M., Driggs, D., Thorpe, M., Gilbey, J., Yeung, M., Ursprung, S., Aviles-Rivero, A.I., Etmann, C., McCague, C., Beer, L., et al., 2021. Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans. Nature Machine Intelligence 3, 199–217
2021
-
[43]
The sankey diagram in energy and material flow man- agement: part ii: methodology and current applications
Schmidt, M., 2008. The sankey diagram in energy and material flow man- agement: part ii: methodology and current applications. Journal of indus- trial ecology 12, 173–185
2008
-
[44]
Speedy iqa for desktop: An image viewer and labeller for image quality assessment (iqa).https://github.com/selbs/speedy_iqa
Selby, I., 2026a. Speedy iqa for desktop: An image viewer and labeller for image quality assessment (iqa).https://github.com/selbs/speedy_iqa. GitHub repository, accessed 22 April 2026
2026
-
[45]
Scientific Data 9, 487
Brax, brazilian labeled chest x-ray dataset. Scientific Data 9, 487
-
[46]
Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia
Shih, G., Wu, C.C., Halabi, S.S., Kohli, M.D., Prevedello, L.M., Cook, T.S., Sharma, A., Amorosa, J.K., Arteaga, V., Galperin-Aizenberg, M., et al., 2019. Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia. Radiology: Artifi- cial Intelligence 1, e180041
2019
-
[47]
The proof and measurement of association between two things
Spearman, C., 1904. The proof and measurement of association between two things. The American Journal of Psychology 15, 72–101. URL:http: //www.jstor.org/stable/1412159
arXiv 1904
-
[48]
Multi-granularity cross-modal alignment for gener- alized medical visual representation learning, in: Advances in Neural Information Processing Systems, pp
Wang, F., Zhou, Y., Wang, S., Vardhanabhuti, V., Yu, L., 2022a. Multi-granularity cross-modal alignment for gener- alized medical visual representation learning, in: Advances in Neural Information Processing Systems, pp. 33536–33549. URL:https://papers.nips.cc/paper_files/paper/2022/hash/ d925bda407ada0df3190df323a212661-Abstract-Conference.html. 34
2022
-
[49]
Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M.,
-
[50]
Speedy qc: Customisable annotation tool for medical images.https://github.com/selbs/speedy_qc
Selby, I., 2026b. Speedy qc: Customisable annotation tool for medical images.https://github.com/selbs/speedy_qc. GitHub repository, ac- cessed 22 April 2026
2026
-
[51]
Wang, Z., Wu, Z., Agarwal, D., Sun, J., 2022b. MedCLIP: Con- trastive learning from unpaired medical images and text, in: Goldberg, Y., Kozareva, Z., Zhang, Y. (Eds.), Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Abu Dhabi, United Arab Emirates. pp. 3876–3887. URL:http...
-
[52]
The effect of class imbalance on precision-recall curves
Williams, C.K., 2021. The effect of class imbalance on precision-recall curves. Neural Computation 33, 853–857
2021
-
[53]
Medklip: Med- ical knowledge enhanced language-image pre-training for x-ray diagnosis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp
Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W., 2023. Medklip: Med- ical knowledge enhanced language-image pre-training for x-ray diagnosis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 21372–21383
2023
-
[54]
Weakly supervised lesion localization with probabilistic-cam pooling
Ye, W., Yao, J., Xue, H., Li, Y., 2020. Weakly supervised lesion localization with probabilistic-cam pooling. arXiv preprint arXiv:2005.14480
Pith/arXiv arXiv 2020
-
[55]
Fsim: A feature similarity indexforimagequalityassessment
Zhang, L., Zhang, L., Shen, X., Mou, X., 2011. Fsim: A feature similarity indexforimagequalityassessment. IEEETransactionsonImageProcessing 20, 2378–2386. doi:10.1109/TIP.2011.2109730
arXiv 2011
-
[56]
Image quality assessment: from error visibility to structural similarity
Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 600–612
2004
-
[57]
Gen- eralized radiograph representation learning via cross-supervision between images and free-text radiology reports
Zhou, H.Y., Chen, X., Zhang, Y., Luo, R., Wang, L., Yu, Y., 2022. Gen- eralized radiograph representation learning via cross-supervision between images and free-text radiology reports. Nature Machine Intelligence 4, 32–
2022
-
[58]
Advancing radiograph representation learning with masked record modeling, in: The Eleventh International Conference on Learning Representations (ICLR)
Zhou, H.Y., Lian, C., Wang, L., Yu, Y., 2023. Advancing radiograph representation learning with masked record modeling, in: The Eleventh International Conference on Learning Representations (ICLR). URL: https://openreview.net/forum?id=w-x7U26GM7j. 35
2023
-
[59]
Zhou, Y., Faith, T., Xu, Y., Leng, S., Xu, X., Liu, Y., Goh, R.S.M.,
-
[62]
Contrastive learning of medical visual representations from paired images and text, in: Proceedings of the 7th Machine Learning for Healthcare Con- ference, PMLR
Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., Langlotz, C.P., 2022. Contrastive learning of medical visual representations from paired images and text, in: Proceedings of the 7th Machine Learning for Healthcare Con- ference, PMLR. pp. 2–25. URL:https://proceedings.mlr.press/v182/ zhang22a.html
2022
-
[64]
doi:10.1038/s42256-021-00425-9
-
[67]
Advances in Neural Information Processing Systems 37, 6625–6647
Benchx: A unified benchmark framework for medical vision-language pretraining on chest x-rays. Advances in Neural Information Processing Systems 37, 6625–6647. 36 Appendix A. Dataset Label Prevalence Pathology CR Img1 Img2 Rep1 Rep2 No Finding 121 (18.6%) 91 (14.0%) 129 (19.8%) 16 (2.5%) 20 (3.1%) Enlarg. Card. 1 (0.2%) 0 (0.0%) 1 (0.2%) 2 (0.3%) 1 (0.2%)...
-
[2017]
Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax dis- eases. CoRR abs/1705.02315. URL:http://arxiv.org/abs/1705.02315, arXiv:1705.02315
-
[2019]
URL:https://physionet.org/content/ mimic-cxr/, doi:10.13026/C2JT1Q
The mimic-cxr database. URL:https://physionet.org/content/ mimic-cxr/, doi:10.13026/C2JT1Q
-
[2021]
Learning transferable visual models from natural language supervi- sion, in: ICML. 33
-
[2022]
Torchxrayvision: A library of chest x-ray datasets and models, in: International Conference on Medical Imaging with Deep Learning, PMLR. pp. 231–249
-
[2024]
Can rule-based insights enhance llms for radiology report classifi- cation? introducing the radprompt methodology., in: Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, pp. 212–235
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.