Pith. sign in

REVIEW 3 major objections 6 minor 67 references

Chest X-ray model rankings depend heavily on whether labels come from images or radiology reports, and on which image-quality metric is used, so evaluation references should be chosen as part of clinical validity, not as an afterthought.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review

2026-08-01 00:07 UTC pith:3YY6HEGG

load-bearing objection Solid, data-rich empirical study showing evaluation-reference choice can reorder CXR model rankings; the noisy image-reference is a real caveat but doesn't overturn the core message. the 3 major comments →

arxiv 2607.26333 v1 pith:3YY6HEGG submitted 2026-07-28 eess.IV cs.CVcs.LG

Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

classification eess.IV cs.CVcs.LG
keywords chest X-rayevaluation referencereport-derived labelsimage-derived labelsimage quality assessmentmodel rankingclinical validitypaired dataset
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper establishes that the choice of evaluation reference is not a technical detail: in chest X-ray machine learning, swapping image-derived labels for report-derived labels, or swapping a generic image-quality metric for expert judgment, can change which models and methods are judged best. The authors built a paired dataset in which expert radiologists independently labeled the same studies from the image alone and from the radiology report, and rated degraded images for diagnostic usability. They show that for pathologies such as Atelectasis and Consolidation, the top-ranked model under one reference can fall far down the ranking under another, and that commonly used quality metrics like PSNR and SSIM correlate weakly with expert assessments. The conclusion is that evaluation-reference selection should be treated as a central component of clinical validity in CXR machine learning, justified per pathology, task, and intended use.

Core claim

Using 650 paired studies from a clinical cohort and 735 from a public dataset, the paper shows that image-derived and report-derived pathology labels disagree in a pathology-dependent way, and that this disagreement propagates to model rankings: for fine-tuned vision-language models, rankings under the two label sources can be nearly uncorrelated for Atelectasis (SRCC ≈ 0.21), while Pleural effusion remains stable. In image quality assessment, expert rankings of diagnostic usability correlate strongly with some full-reference metrics (HaarPSI, GMSD) but weakly with PSNR, CW-SSIM, and no-reference metrics. The authors conclude that image- and report-derived labels answer different clinical qu

What carries the argument

The central object is the paired evaluation-reference dataset XQA-Chest, with its image-only and report-only annotation protocols and an adjudicated (or automatically fused) consensus label. Rank correlation coefficients (SRCC, KRCC) and Cohen's kappa carry the argument by quantifying how label-source disagreement translates into model-ranking instability and how IQA metrics align with expert diagnostic-usability rankings.

Load-bearing premise

The paired image-derived labels (conflict-resolved on the clinical cohort, FUSE_OR-fused on the public dataset) accurately represent expert image-only judgment, and their disagreement with report-derived labels reflects systematic differences rather than idiosyncratic annotation noise.

What would settle it

An independent multi-institution annotation study with adjudicated image labels that found model rankings under image- and report-derived references to be highly correlated (e.g., SRCC > 0.9) for Atelectasis and Consolidation would directly contradict the claim that reference choice flips rankings. Alternatively, demonstrating that the ranking instability disappears when image labels are produced by a larger, more diverse panel of radiologists would indicate the effect is annotation noise rather than a structural property of the label sources.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

If this is right

  • Benchmark leaderboards in CXR classification are reference-dependent; a model's best status is only meaningful relative to the chosen label source.
  • Evaluation against report-derived labels is internally reproducible but can diverge from image-derived evaluation; studies should state which report section (Findings vs Impression) is used.
  • IQA metrics used for reconstruction or generative models should be validated against expert diagnostic-usability ratings before deployment, since PSNR and no-reference metrics can misrank outputs.
  • The paired-references resource enables future work to test label-fusion methods and automated report-extraction strategies against a common expert benchmark.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If reference-dependence generalizes, single-reference benchmarks in medical imaging can mislead model selection; multi-reference evaluation may become standard practice for clinically meaningful comparisons.
  • The low image inter-annotator agreement for some pathologies suggests image-derived labels are themselves noisy; a multi-institution adjudicated study could determine whether ranking instability is structural or partly annotation noise.
  • The IQA results imply that clinically validated, task-specific quality metrics could be developed directly from expert preference data rather than borrowed from generic perceptual metrics.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper asks whether evaluation-reference choices—image-derived vs report-derived pathology labels, and generic IQA metrics vs expert ratings of diagnostic usability—change model rankings and hence scientific conclusions in chest X-ray machine learning. The authors introduce XQA-Chest, a paired dataset of expert image- and report-derived labels for 650 CUH studies, plus expert IQA ratings for 1,571 degraded images, and they annotate a 735-study MIMIC-CXR subset under similar protocols. They evaluate a broad set of supervised CNNs and vision-language models under different label references, measure agreement and rank correlations (SRCC/KRCC with bootstrap CIs), and similarly compare IQA metrics against expert z-MOS rankings. The central finding is that changing the evaluation reference can reorder models, sometimes dramatically (e.g., VLM-FT rankings for Atelectasis are nearly uncorrelated across references), and that commonly used IQA metrics such as PSNR and CW-SSIM align only weakly with expert judgments.

Significance. If the central claim holds, the paper makes an important methodological contribution: benchmark rankings in CXR machine learning are reference-dependent, and the choice of evaluation reference should be treated as a core component of clinical validity rather than an implementation detail. The paper's strengths include a new paired dataset with expert annotations, extensive agreement statistics with pessimistic/optimistic uncertainty mappings, bootstrap confidence intervals for pathology rank correlations, sensitivity analyses for label-fusion methods, and a qualitative radiology review of disagreements. These are appropriate tools for the claim and provide a useful resource for the community. The main risk is that the image-derived reference is not shown to be a stable construct independent of individual annotator noise; if the apparent reference effect is substantially driven by image-annotator disagreement, the interpretation as a systematic image-vs-report distinction would need to be tempered.

major comments (3)
  1. [§3.1.1, Table A.5] The conflict-resolved (CR) image labels for Atelectasis (165 positives), Cardiomegaly (21), and Support Devices (274) exactly match Img2's counts, while Img1's counts are very different (44, 153, 309). This strongly suggests that adjudication often deferred to one annotator rather than producing a genuine synthesis. The paper does not report model rankings induced by Img1 and Img2 separately. Without these controls, the central claim that changing the label source (image vs report) systematically changes rankings is confounded with image-annotator noise: the image-image rank correlation may be as low as the image-report rank correlation. For the VLM-FT group on Atelectasis, the CR-vs-report SRCC is 0.21±0.22 (Table E.26); if Img1-vs-Img2 SRCC is similar, the effect is not specific to the image/report distinction. Please add per-annotator ranking analyses for XQA-RR (and MIMIC where possi
  2. [§2.1.3, Table F.30] FUSE_OR is validated against CR on XQA-RR and then used as the image-derived reference on MIMIC-CXR, where no adjudication is available. This transfer is not fully supported. First, the high FUSE_OR-vs-CR agreement is partly a consequence of the CR pattern noted above: for pathologies where CR equals one annotator, an optimistic union rule will naturally reproduce the more positive annotator. Second, the MIMIC inter-annotator agreement profile differs substantially from XQA-RR (e.g., Lung opacity κ=0.11 on MIMIC vs 0.42 on XQA-RR; Table B.12 vs B.9), so a fusion rule validated on one institution's disagreement structure may behave differently elsewhere. The MIMIC comparisons labeled 'image-derived vs report-derived' are therefore better described as comparisons between a particular optimistic fusion rule and manually annotated report sections. Please add sensitivity analyses using Img1 a
  3. [§3.2, Table 4] The IQA rank correlations are reported as point estimates without confidence intervals, in contrast to the pathology-ranking results, which use 1000-sample bootstrap CIs. With 1,571 images, the differences among the better-performing metrics (e.g., HaarPSI 0.87, GMSD 0.86, HaarPSImed 0.85, MS-SSIM 0.84) may be within sampling variability. The paper's IQA conclusion—that metric selection determines which processed images appear best—depends on these metric-level differences being reliable. Please provide bootstrap CIs (or paired significance tests) for the SRCC/KRCC values in Table 4.
minor comments (6)
  1. [Abstract, §4] The abstract and Discussion state that SSIM and PSNR 'often fail to align' with expert judgment, but Table 4 gives SSIM-matlab SRCC=0.78, which is moderate agreement. Consider rephrasing to avoid overstatement, e.g., 'SSIM shows weaker alignment than top metrics, and PSNR aligns weakly.'
  2. [§2.1.2, Tables A.7/A.8] The handling of 'not mentioned' report labels is not stated. The tables list explicit negative, uncertain, and positive counts, while the positive prevalences in Table A.5 imply that 'not mentioned' was collapsed to negative in the binary analyses. This is a standard CheXpert-like assumption but should be stated explicitly, since it affects all report-derived agreement and ranking results.
  3. [Table 2] The symbols ✓/✗ in Table 2 are not defined consistently; for example, the 'Images' and 'Reports' columns use different combinations. Add a legend or use explicit 'Yes/No' entries.
  4. [§2.1.2] The text describes a 14-category schema and then excludes 'No finding' and 'Support devices' to obtain 12 pathologies. State this exclusions before listing the 14 categories to avoid apparent inconsistency.
  5. [Figures 10b–14b] The x-axis labels in Figures 10b–14b use the order 'Impression, FUSE_OR, Findings' while the text and the MIMIC tables usually list 'FUSE_OR, Findings, Impressions.' Use a consistent order for readability.
  6. [§4] The Discussion states that the observed divergence is 'systematic, clinically interpretable.' Given the single-institution annotation and the CR=Img2 pattern noted above, 'systematic' is too strong; please qualify this claim in light of the per-annotator analysis requested in the major comments.

Circularity Check

0 steps flagged

No significant circularity: ranking and IQA results are empirical comparisons from newly collected labels, not derived from fitted parameters or load-bearing self-citation.

full rationale

The paper's central claims are empirical: it collects paired expert image-derived and report-derived labels (XQA-RR and MIMIC-CXR subsets), evaluates a diverse set of classifiers and VLMs against each reference, and measures rank stability. No target quantity is defined in terms of its own input; no fitted parameter is renamed as a prediction. The FUSE_OR fusion rule is validated against conflict-resolved labels on XQA-RR and then applied to MIMIC-CXR; the paper explicitly acknowledges in Section 5 (Limitations) that this is "a choice validated against CUH but not equivalent to third-annotator adjudication," so this is a construct-validity caveat rather than a circular derivation. Low image inter-annotator kappa values (e.g., Atelectasis kappa = 0.15) and single-institution annotation are data-quality limitations, not evidence that conclusions are forced by construction. Self-citations exist: RadPert [17] supplies CUH fine-tuning labels, and Breger et al. [2,3] motivate IQA concerns. However, the ranking experiments and IQA correlations are recomputed from the newly annotated data and public models, so those citations are not load-bearing; the central finding would stand on the reported tables alone. The limitations that MIMIC labels come from automatic fusion rather than adjudication, and that both datasets come from one institution, affect generalizability but do not reduce the derivation to its inputs. No equation in the paper is identical to another by construction, and no 'prediction' is statistically forced by a fitted parameter. Score 1 reflects only the presence of minor, non-load-bearing self-citation; there is no material circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

No fitted free parameters appear in the central claims. The analysis is empirical; the main assumptions concern how label sources are operationalized, the transfer of FUSE_OR to MIMIC-CXR, and the representativeness of the small clinical cohort. No new theoretical entities are introduced; XQA-Chest is a dataset resource rather than an invented entity.

axioms (5)
  • domain assumption Expert image-only annotation is a valid operationalization of what is visually recoverable from a radiograph.
    The whole comparison defines one reference as conflict-resolved image-derived labels. Low inter-annotator kappa for Atelectasis (0.15) shows this reference is noisy; if that noise is idiosyncratic rather than source-typical, rank differences are partly annotation noise. Section 2.1.2, Figure 5.
  • domain assumption Report-derived labels are a valid operationalization of the clinically contextualized interpretation, and Findings vs Impressions can be scored as separate references.
    Used to define report references. Within-report disagreement between Findings and Impressions on MIMIC-CXR (macro kappa 0.31) shows this reference is internally variable. Section 2.1.2, Figure 7.
  • ad hoc to paper FUSE_OR fusion transfers from XQA-RR to MIMIC-CXR as a proxy for adjudicated image labels.
    FUSE_OR was selected after comparing fusion methods on XQA-RR and then used to create the MIMIC image reference. The paper explicitly notes this is 'not equivalent to third-annotator adjudication' (Section 5).
  • domain assumption ROC-AUC is a suitable cross-reference performance metric despite differing label prevalence.
    The paper argues prevalence invariance, but this relies on fixed positive/negative score distributions per reference. If score distributions are not comparable across references, AUC differences may conflate reference effects with calibration. Section 2.1.1.
  • domain assumption A two-site sample with two image annotators per site is representative enough for ranking-stability conclusions.
    The limitations acknowledge single-institution annotation style and small sample for rare pathologies. If annotation culture differs elsewhere, the observed rank-instability magnitudes may not generalize. Section 5.

reviewed 2026-08-01 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance." pith.science (2026). https://pith.science/paper/3YY6HEGG

@misc{pith2026260726333,
  author       = {Pith},
  title        = {Pith review of: Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3YY6HEGG}},
  note         = {Machine review of arXiv:2607.26333}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Chest X-ray (CXR) machine learning relies heavily on automated evaluation using reference standards that aim to approximate clinical judgment. However, commonly used report-derived labels for pathology classification or generic image quality metrics for reconstruction may not reliably reflect clinical judgment. We systematically investigate how evaluation-reference choices affect model performance and ranking in both pathology classification and image quality assessment (IQA). To enable controlled comparison across evaluation references, we collected paired expert image- and report-derived labels for thoracic findings from a clinical cohort at Cambridge University Hospitals (CUH) and curated a subset of the public MIMIC-CXR dataset, along with expert ratings of diagnostic image quality. We show that for supervised image classifiers (ResNet, DenseNet), several zero-shot and fine-tuned vision-language models (e.g., MedKLIP, GLoRIA, and ConVIRT), changing the label source leads to substantial differences not only in performance estimates but also in model rankings. In parallel, alignment of IQA measures with expert judgment depends heavily on the choice of measure, and commonly used IQA metrics such as SSIM and PSNR often fail to align with expert assessments of diagnostic usability. Our results demonstrate that evaluation choices are crucial: they can determine which models and methods appear best and are therefore selected for further development or deployment. The selection of evaluation references should therefore be treated as a central component of clinical validity in CXR machine learning, and justified with respect to the pathology, imaging task, and intended downstream clinical use.

Figures

Figures reproduced from arXiv: 2607.26333 by Alex Sawer, Anna Breger, Anna Korhonen, Arthikkaa Thavakumar, Carola-Bibiane Sch\"onlieb, Clemens Karner, Ian Selby, Jake Beckford, J. H. F. Rudd, John Li Chen, Jonathan Weir-McCall, Judith Babar, Michael Roberts, Panagiotis Fytas, Shahab Shahipasand, Simon Baker, Timothy J. Sadler.

Figure 1
Figure 1. Figure 1: Overview of pathology classification with deep learning models. There are two [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of full-reference, no-reference, and clinician-preference-based IQA. The [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview XQA-Chest. examine disagreement between image-derived and report-derived labels on paired studies, they do not release such paired expertly-annotated labels as a pub￾lic benchmark or systematically analyze how evaluation reference choice af￾fects downstream conclusions; instead, expert image-derived labels are typically treated as the evaluation reference, despite the underlying assumption of thei… view at source ↗
Figure 4
Figure 4. Figure 4: Sankey diagrams showing agreement and omission patterns between conflict-resolved [PITH_FULL_IMAGE:figures/full_fig_p015_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: XQA-RR inter-annotator and cross-source label agreement summarized using [PITH_FULL_IMAGE:figures/full_fig_p016_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Agreement between each image-derived label fusion method and the XQA-RR [PITH_FULL_IMAGE:figures/full_fig_p016_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: MIMIC-CXR image inter-annotator agreement, agreement across labels derived [PITH_FULL_IMAGE:figures/full_fig_p017_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: XQA-RR rank correlation coefficients between model rankings induced by each [PITH_FULL_IMAGE:figures/full_fig_p019_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: MIMIC-CXR rank correlation coefficients between model rankings induced by each [PITH_FULL_IMAGE:figures/full_fig_p020_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Model rankings across label sources for Atelectasis. of evaluation variability and can affect conclusions about which models perform best. 3.2. IQA [PITH_FULL_IMAGE:figures/full_fig_p021_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Model rankings across label sources for Consolidation. Rep1 Rep2 CR Evaluation references 5 10 15 20 25 30 35 Rank (1 is best) XRV VLM-FT VLM-0-SHOT (a) XQA-RR Impression FUSE_OR Findings Evaluation references 5 10 15 20 25 Rank (1 is best) XRV VLM-FT VLM-0-SHOT (b) MIMIC-CXR [PITH_FULL_IMAGE:figures/full_fig_p022_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Model rankings across label sources for Lung opacity. 22 [PITH_FULL_IMAGE:figures/full_fig_p022_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Model rankings across label sources for Pleural effusion. Rep1 Rep2 CR Evaluation references 5 10 15 20 25 30 35 Rank (1 is best) XRV VLM-FT VLM-0-SHOT (a) XQA-RR Impression FUSE_OR Findings Evaluation references 5 10 15 20 25 Rank (1 is best) XRV VLM-FT VLM-0-SHOT (b) MIMIC-CXR [PITH_FULL_IMAGE:figures/full_fig_p023_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Model rankings across label sources, macro-averaged across the four selected [PITH_FULL_IMAGE:figures/full_fig_p023_14.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 1 canonical work pages

  1. [1]

    On model evalu- ation under non-constant class imbalance, in: International Conference on Computational Science, Springer

    Brabec, J., Komárek, T., Franc, V., Machlica, L., 2020. On model evalu- ation under non-constant class imbalance, in: International Conference on Computational Science, Springer. pp. 74–87

  2. [2]

    A study of why we need to reassess full reference image quality assessment with medical images

    Breger, A., Biguri, A., Landman, M.S., Selby, I., Amberg, N., Brunner, E., Gröhl, J., Hatamikia, S., Karner, C., Ning, L., et al., 2025a. A study of why we need to reassess full reference image quality assessment with medical images. Journal of Imaging Informatics in Medicine 38, 3444–3469

  3. [3]

    A study on the adequacy of com- mon iqa measures for medical images, in: Su, R., Frangi, A.F., Zhang, Y

    Breger, A., Karner, C., Selby, I., Gröhl, J., Dittmer, S., Lilley, E., Babar, J., Beckford, J., Else, T.R., Sadler, T.J., Shahipasand, S., Thavakumar, A., Roberts, M., Schönlieb, C.B., 2025b. A study on the adequacy of com- mon iqa measures for medical images, in: Su, R., Frangi, A.F., Zhang, Y. (Eds.), Proceedings of 2024 International Conference on Medi...

  4. [4]

    Padch- est: A large chest x-ray image dataset with multi-label annotated reports

    Bustos, A., Pertusa, A., Salinas, J.M., de la Iglesia-Vayá, M., 2020. Padch- est: A large chest x-ray image dataset with multi-label annotated reports. Medical Image Analysis 66, 101797. URL:https://www.sciencedirect. com/science/article/pii/S1361841520301614, doi:https://doi.org/ 10.1016/j.media.2020.101797

  5. [5]

    Why almost all ml models for medicine are wrong-and what we need for evidence-based medical ai

    Cabitza, F., Jurman, G., Molinari, F., Bellazzi, R., 2026. Why almost all ml models for medicine are wrong-and what we need for evidence-based medical ai. International Journal of Medical Informatics , 106538

  6. [6]

    A simple frame- work for contrastive learning of visual representations, in: International conference on machine learning, PmLR

    Chen, T., Kornblith, S., Norouzi, M., Hinton, G., 2020. A simple frame- work for contrastive learning of visual representations, in: International conference on machine learning, PmLR. pp. 1597–1607

  7. [7]

    Towards unifying medical vision-and-language pre-training via soft prompts, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp

    Chen, Z., Diao, S., Wang, B., Li, G., Wan, X., 2023. Towards unifying medical vision-and-language pre-training via soft prompts, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23403–23413. 30

  8. [8]

    A coefficient of agreement for nominal scales

    Cohen, J., 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20, 37–46

  9. [9]

    Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit

    Cohen, J., 1968. Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin 70, 213

  10. [10]

    On the limits of cross-domain generalization in automated x-ray prediction, in: Medical Imaging with Deep Learning, PMLR

    Cohen, J.P., Hashir, M., Brooks, R., Bertrand, H., 2020. On the limits of cross-domain generalization in automated x-ray prediction, in: Medical Imaging with Deep Learning, PMLR. pp. 136–155

  11. [11]

    Cohen, J.P., Viviano, J.D., Bertin, P., Morrison, P., Torabian, P., Guar- rera, M., Lungren, M.P., Chaudhari, A., Brooks, R., Hashir, M., et al.,

  12. [12]

    The relationship between precision-recall and roc curves, in: Proceedings of the 23rd international conference on Machine learning, pp

    Davis, J., Goadrich, M., 2006. The relationship between precision-recall and roc curves, in: Proceedings of the 23rd international conference on Machine learning, pp. 233–240

  13. [13]

    Maximum likelihood estimation of ob- server error-rates using the EM algorithm

    Dawid, A.P., Skene, A.M., 1979. Maximum likelihood estimation of ob- server error-rates using the EM algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics) 28, 20–28

  14. [14]

    Preparing a collection of radiology examinations for distribution and retrieval

    Demner-Fushman, D., Kohli, M.D., Rosenman, M.B., Shooshan, S.E., Ro- driguez, L., Antani, S., Thoma, G.R., McDonald, C.J., 2015. Preparing a collection of radiology examinations for distribution and retrieval. Journal of the American Medical Informatics Association 23, 304–310

  15. [15]

    An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Un- terthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkor- eit, J., Houlsby, N., 2021. An image is worth 16x16 words: Transformers for image recognition at scale, in: International Conference on Learning Representations

  16. [16]

    Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy

    Efron, B., Tibshirani, R., 1986. Bootstrap methods for standard errors, confidence intervals, and other measures of statistical accuracy. Statistical science , 54–75

  17. [17]

    Fytas, P., Breger, A., Selby, I., Baker, S., Shahipasand, S., Korhonen, A.,

  18. [18]

    Digital Image Processing

    Gonzalez, R.C., Woods, R.E., 2008. Digital Image Processing. 3rd ed., Pearson

  19. [19]

    Evaluating the robustness and readiness of large frontier models in health ai applications

    Gu, Y., Fu, J., Liu, X., Valanarasu, J.M.J., Codella, N.C., Tan, R., Liu, Q., Jin, Y., Zhang, S., Wang, J., et al., 2026. Evaluating the robustness and readiness of large frontier models in health ai applications. Nature Medicine , 1–9. 31

  20. [20]

    Robust classification from noisy labels: Integrating additional knowledge for chest radiography abnormality assessment

    Gündel, S., Setio, A.A., Ghesu, F.C., Grbic, S., Georgescu, B., Maier, A., Comaniciu, D., 2021. Robust classification from noisy labels: Integrating additional knowledge for chest radiography abnormality assessment. Med- ical Image Analysis 72, 102087

  21. [21]

    The meaning and use of the area under a receiver operating characteristic (roc) curve

    Hanley, J.A., McNeil, B.J., 1982. The meaning and use of the area under a receiver operating characteristic (roc) curve. Radiology 143, 29–36

  22. [22]

    Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

    He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778

  23. [23]

    Hovy, D., Berg-Kirkpatrick, T., Vaswani, A., Hovy, E., 2013. Learning whom to trust with MACE, in: Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), pp. 1120–1130

  24. [24]

    Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp

    Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q., 2017. Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708

  25. [25]

    Huang, S.C., Shen, L., Lungren, M.P., Yeung, S., 2021. Gloria: A mul- timodal global-local representation learning framework for label-efficient medical image recognition, in: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3942–3951

  26. [26]

    Chexpert: A large chest radiograph dataset with uncertainty labels and expert compar- ison

    Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R.L., Shpanskaya, K.S., Seekins, J., Mong, D.A., Halabi, S.S., Sandberg, J.K., Jones, R., Larson, D.B., Lan- glotz, C.P., Patel, B.N., Lungren, M.P., Ng, A.Y., 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert compar- i...

  27. [27]

    Jain, S., Smit, A., Truong, S.Q., Nguyen, C.D., Huynh, M.T., Jain, M., Young, V.A., Ng, A.Y., Lungren, M.P., Rajpurkar, P., 2021. Visu- alchexbert: addressing the discrepancy between radiology report labels and image labels, in: Proceedings of the Conference on Health, Infer- ence, and Learning, Association for Computing Machinery, New York, NY, USA. p. 1...

  28. [28]

    Johnson, A.E.W., Pollard, T., Mark, R., Berkowitz, S., Horng, S.,

  29. [29]

    Parameter choices in 32 haarpsiforiqawithmedicalimages, in: 2025IEEE22ndInternationalSym- posium on Biomedical Imaging (ISBI), pp

    Karner, C., Gröhl, J., Selby, I., Babar, J., Beckford, J., Else, T.R., Sadler, T.J., Shahipasand, S., Thavakumar, A., Roberts, M., Rudd, J.H., Schön- lieb, C.B., Weir-McCall, J.R., Breger, A., 2025. Parameter choices in 32 haarpsiforiqawithmedicalimages, in: 2025IEEE22ndInternationalSym- posium on Biomedical Imaging (ISBI), pp. 1–5. doi:10.1109/ISBI60581....

  30. [30]

    A new measure of rank corre- lation

    KENDALL, M.G., 1938. A new measure of rank corre- lation. Biometrika 30, 81–93. URL:https://doi.org/ 10.1093/biomet/30.1-2.81, doi:10.1093/biomet/30.1-2.81, arXiv:https://academic.oup.com/biomet/article-pdf/30/1-2/81/423380/30-1-2-81.pdf

  31. [31]

    Klie, J.C., Bugert, M., Boullosa, B., Eckart de Castilho, R., Gurevych, I., 2018. The INCEpTION platform: Machine-assisted and knowledge- oriented interactive annotation, in: Proceedings of the 27th International Conference on Computational Linguistics: System Demonstrations, Santa Fe, New Mexico. pp. 5–9. URL:https://www.aclweb.org/anthology/ C18-2002

  32. [32]

    Liu, C., Cheng, S., Chen, C., Qiao, M., Zhang, W., Shah, A., Bai, W., Ar- cucci, R., 2023. M-flag: Medical vision-language pre-training with frozen language models and latent space geometry optimization, in: Medical Im- age Computing and Computer Assisted Intervention – MICCAI 2023, pp. 637–647. doi:10.1007/978-3-031-43907-0_61

  33. [33]

    Chest radiograph interpretation with deep learning models: as- sessment with radiologist-adjudicated reference standards and population- adjusted evaluation

    Majkowska, A., Mittal, S., Steiner, D.F., Reicher, J.J., McKinney, S.M., Duggan, G.E., Eswaran, K., Cameron Chen, P.H., Liu, Y., Kalidindi, S.R., et al., 2020. Chest radiograph interpretation with deep learning models: as- sessment with radiologist-adjudicated reference standards and population- adjusted evaluation. Radiology 294, 421–431

  34. [34]

    completely blind

    Mittal, A., Soundararajan, R., Bovik, A.C., 2012. Making a “completely blind” image quality analyzer. IEEE Signal processing letters 20, 209–212

  35. [35]

    Vindr-cxr: An open dataset of chest x-rays with radiologist’s annotations

    Nguyen, H.Q., Lam, K., Le, L.T., Pham, H.H., Tran, D.Q., Nguyen, D.B., Le, D.D., Pham, C.M., Tong, H.T., Dinh, D.H., et al., 2022. Vindr-cxr: An open dataset of chest x-rays with radiologist’s annotations. Scientific Data 9, 429

  36. [36]

    Representation learning with con- trastive predictive coding

    Oord, A.v.d., Li, Y., Vinyals, O., 2018. Representation learning with con- trastive predictive coding. arXiv preprint arXiv:1807.03748

  37. [37]

    Evaluation: fromprecision, recallandf-measuretoroc, informedness, markednessandcorrelation

    Powers, D.M., 2020. Evaluation: fromprecision, recallandf-measuretoroc, informedness, markednessandcorrelation. arXivpreprintarXiv:2010.16061

  38. [38]

    Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.,

  39. [39]

    Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists

    Rajpurkar, P., Irvin, J., Ball, R.L., Zhu, K., Yang, B., Mehta, H., Duan, T., Ding, D., Bagul, A., Langlotz, C.P., et al., 2018. Deep learning for chest radiograph diagnosis: A retrospective comparison of the chexnext algorithm to practicing radiologists. PLoS medicine 15, e1002686

  40. [40]

    Reis, E.P., De Paiva, J.P., Da Silva, M.C., Ribeiro, G.A., Paiva, V.F., Bulgarelli, L., Lee, H.M., Santos, P.V., Brito, V.M., Amaral, L.T., et al.,

  41. [41]

    A haar wavelet-basedperceptualsimilarityindexforimagequalityassessment

    Reisenhofer, R., Bosse, S., Kutyniok, G., Wiegand, T., 2018. A haar wavelet-basedperceptualsimilarityindexforimagequalityassessment. Sig- nal Processing: Image Communication 61, 33–43. doi:10.1016/j.image. 2017.11.001

  42. [42]

    Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans

    Roberts, M., Driggs, D., Thorpe, M., Gilbey, J., Yeung, M., Ursprung, S., Aviles-Rivero, A.I., Etmann, C., McCague, C., Beer, L., et al., 2021. Common pitfalls and recommendations for using machine learning to detect and prognosticate for covid-19 using chest radiographs and ct scans. Nature Machine Intelligence 3, 199–217

  43. [43]

    The sankey diagram in energy and material flow man- agement: part ii: methodology and current applications

    Schmidt, M., 2008. The sankey diagram in energy and material flow man- agement: part ii: methodology and current applications. Journal of indus- trial ecology 12, 173–185

  44. [44]

    Speedy iqa for desktop: An image viewer and labeller for image quality assessment (iqa).https://github.com/selbs/speedy_iqa

    Selby, I., 2026a. Speedy iqa for desktop: An image viewer and labeller for image quality assessment (iqa).https://github.com/selbs/speedy_iqa. GitHub repository, accessed 22 April 2026

  45. [45]

    Scientific Data 9, 487

    Brax, brazilian labeled chest x-ray dataset. Scientific Data 9, 487

  46. [46]

    Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia

    Shih, G., Wu, C.C., Halabi, S.S., Kohli, M.D., Prevedello, L.M., Cook, T.S., Sharma, A., Amorosa, J.K., Arteaga, V., Galperin-Aizenberg, M., et al., 2019. Augmenting the national institutes of health chest radiograph dataset with expert annotations of possible pneumonia. Radiology: Artifi- cial Intelligence 1, e180041

  47. [47]

    The proof and measurement of association between two things

    Spearman, C., 1904. The proof and measurement of association between two things. The American Journal of Psychology 15, 72–101. URL:http: //www.jstor.org/stable/1412159

  48. [48]

    Multi-granularity cross-modal alignment for gener- alized medical visual representation learning, in: Advances in Neural Information Processing Systems, pp

    Wang, F., Zhou, Y., Wang, S., Vardhanabhuti, V., Yu, L., 2022a. Multi-granularity cross-modal alignment for gener- alized medical visual representation learning, in: Advances in Neural Information Processing Systems, pp. 33536–33549. URL:https://papers.nips.cc/paper_files/paper/2022/hash/ d925bda407ada0df3190df323a212661-Abstract-Conference.html. 34

  49. [49]

    Wang, X., Peng, Y., Lu, L., Lu, Z., Bagheri, M., Summers, R.M.,

  50. [50]

    Speedy qc: Customisable annotation tool for medical images.https://github.com/selbs/speedy_qc

    Selby, I., 2026b. Speedy qc: Customisable annotation tool for medical images.https://github.com/selbs/speedy_qc. GitHub repository, ac- cessed 22 April 2026

  51. [51]

    MedCLIP: Con- trastive learning from unpaired medical images and text, in: Goldberg, Y., Kozareva, Z., Zhang, Y

    Wang, Z., Wu, Z., Agarwal, D., Sun, J., 2022b. MedCLIP: Con- trastive learning from unpaired medical images and text, in: Goldberg, Y., Kozareva, Z., Zhang, Y. (Eds.), Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, Association for Computational Linguistics, Abu Dhabi, United Arab Emirates. pp. 3876–3887. URL:http...

  52. [52]

    The effect of class imbalance on precision-recall curves

    Williams, C.K., 2021. The effect of class imbalance on precision-recall curves. Neural Computation 33, 853–857

  53. [53]

    Medklip: Med- ical knowledge enhanced language-image pre-training for x-ray diagnosis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp

    Wu, C., Zhang, X., Zhang, Y., Wang, Y., Xie, W., 2023. Medklip: Med- ical knowledge enhanced language-image pre-training for x-ray diagnosis, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 21372–21383

  54. [54]

    Weakly supervised lesion localization with probabilistic-cam pooling

    Ye, W., Yao, J., Xue, H., Li, Y., 2020. Weakly supervised lesion localization with probabilistic-cam pooling. arXiv preprint arXiv:2005.14480

  55. [55]

    Fsim: A feature similarity indexforimagequalityassessment

    Zhang, L., Zhang, L., Shen, X., Mou, X., 2011. Fsim: A feature similarity indexforimagequalityassessment. IEEETransactionsonImageProcessing 20, 2378–2386. doi:10.1109/TIP.2011.2109730

  56. [56]

    Image quality assessment: from error visibility to structural similarity

    Wang, Z., Bovik, A.C., Sheikh, H.R., Simoncelli, E.P., 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 600–612

  57. [57]

    Gen- eralized radiograph representation learning via cross-supervision between images and free-text radiology reports

    Zhou, H.Y., Chen, X., Zhang, Y., Luo, R., Wang, L., Yu, Y., 2022. Gen- eralized radiograph representation learning via cross-supervision between images and free-text radiology reports. Nature Machine Intelligence 4, 32–

  58. [58]

    Advancing radiograph representation learning with masked record modeling, in: The Eleventh International Conference on Learning Representations (ICLR)

    Zhou, H.Y., Lian, C., Wang, L., Yu, Y., 2023. Advancing radiograph representation learning with masked record modeling, in: The Eleventh International Conference on Learning Representations (ICLR). URL: https://openreview.net/forum?id=w-x7U26GM7j. 35

  59. [59]

    Zhou, Y., Faith, T., Xu, Y., Leng, S., Xu, X., Liu, Y., Goh, R.S.M.,

  60. [62]

    Contrastive learning of medical visual representations from paired images and text, in: Proceedings of the 7th Machine Learning for Healthcare Con- ference, PMLR

    Zhang, Y., Jiang, H., Miura, Y., Manning, C.D., Langlotz, C.P., 2022. Contrastive learning of medical visual representations from paired images and text, in: Proceedings of the 7th Machine Learning for Healthcare Con- ference, PMLR. pp. 2–25. URL:https://proceedings.mlr.press/v182/ zhang22a.html

  61. [64]

    doi:10.1038/s42256-021-00425-9

  62. [67]

    Advances in Neural Information Processing Systems 37, 6625–6647

    Benchx: A unified benchmark framework for medical vision-language pretraining on chest x-rays. Advances in Neural Information Processing Systems 37, 6625–6647. 36 Appendix A. Dataset Label Prevalence Pathology CR Img1 Img2 Rep1 Rep2 No Finding 121 (18.6%) 91 (14.0%) 129 (19.8%) 16 (2.5%) 20 (3.1%) Enlarg. Card. 1 (0.2%) 0 (0.0%) 1 (0.2%) 2 (0.3%) 1 (0.2%)...

  63. [2017]

    CoRR abs/1705.02315

    Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax dis- eases. CoRR abs/1705.02315. URL:http://arxiv.org/abs/1705.02315, arXiv:1705.02315

  64. [2019]

    URL:https://physionet.org/content/ mimic-cxr/, doi:10.13026/C2JT1Q

    The mimic-cxr database. URL:https://physionet.org/content/ mimic-cxr/, doi:10.13026/C2JT1Q

  65. [2021]

    Learning transferable visual models from natural language supervi- sion, in: ICML. 33

  66. [2022]

    Torchxrayvision: A library of chest x-ray datasets and models, in: International Conference on Medical Imaging with Deep Learning, PMLR. pp. 231–249

  67. [2024]

    Can rule-based insights enhance llms for radiology report classifi- cation? introducing the radprompt methodology., in: Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, pp. 212–235

This paper was first reviewed by deepseek-v4-flash on August 1, 2026.