REVIEW 2 cited by
Demographic Bias of Expert-Level Vision-Language Foundation Models in Medical Imaging
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Advances in artificial intelligence (AI) have achieved expert-level performance in medical imaging applications. Notably, self-supervised vision-language foundation models can detect a broad spectrum of pathologies without relying on explicit training annotations. However, it is crucial to ensure that these AI models do not mirror or amplify human biases, thereby disadvantaging historically marginalized groups such as females or Black patients. The manifestation of such biases could systematically delay essential medical care for certain patient subgroups. In this study, we investigate the algorithmic fairness of state-of-the-art vision-language foundation models in chest X-ray diagnosis across five globally-sourced datasets. Our findings reveal that compared to board-certified radiologists, these foundation models consistently underdiagnose marginalized groups, with even higher rates seen in intersectional subgroups, such as Black female patients. Such demographic biases present over a wide range of pathologies and demographic attributes. Further analysis of the model embedding uncovers its significant encoding of demographic information. Deploying AI systems with these biases in medical imaging can intensify pre-existing care disparities, posing potential challenges to equitable healthcare access and raising ethical questions about their clinical application.
Forward citations
Cited by 2 Pith papers
-
Understanding Dataset Bias in Medical Imaging: A Case Study on Chest X-rays
Classifiers can identify the origin of chest X-rays across NIH, CheXpert, MIMIC-CXR, and PadChest with F1 scores up to about 99%, and the bias appears driven mainly by pixel intensity and texture.
-
Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?
An empirical comparison of three ViT backbones with ResNet for few-shot demographic face authentication reports Swin Transformer as best, but the fairness conclusion is not supported by the experimental design.
Discussion (0). Sign in to comment.