Pith. sign in

REVIEW 2 cited by

CheXclusion: Fairness gaps in deep chest X-ray classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.00827 v2 pith:3PRNJELY submitted 2020-02-14 cs.CV cs.AIcs.LGeess.IVstat.ML

classification cs.CVcs.AIcs.LGeess.IVstat.ML
keywords clinicaldisparitiesclassifiersdatasetsx-rayattributeschestchexclusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Machine learning systems have received much attention recently for their ability to achieve expert-level performance on clinical tasks, particularly in medical imaging. Here, we examine the extent to which state-of-the-art deep learning classifiers trained to yield diagnostic labels from X-ray images are biased with respect to protected attributes. We train convolution neural networks to predict 14 diagnostic labels in 3 prominent public chest X-ray datasets: MIMIC-CXR, Chest-Xray8, CheXpert, as well as a multi-site aggregation of all those datasets. We evaluate the TPR disparity -- the difference in true positive rates (TPR) -- among different protected attributes such as patient sex, age, race, and insurance type as a proxy for socioeconomic status. We demonstrate that TPR disparities exist in the state-of-the-art classifiers in all datasets, for all clinical tasks, and all subgroups. A multi-source dataset corresponds to the smallest disparities, suggesting one way to reduce bias. We find that TPR disparities are not significantly correlated with a subgroup's proportional disease burden. As clinical models move from papers to products, we encourage clinical decision makers to carefully audit for algorithmic disparities prior to deployment. Our code can be found at, https://github.com/LalehSeyyed/CheXclusion

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Performance Claims for Chest X-Ray Models Using Clinical Context

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Chest X-ray models' apparent accuracy drops significantly when evaluated on cases matched to remove clinical context from prior notes, suggesting much of their performance relies on context rather than image evidence.

  2. Fairness of Deep Ensembles: On the interplay between per-group task difficulty and under-representation

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Homogeneous deep ensembles shrink accuracy gaps between demographic groups without lowering overall accuracy, and the optimal training-data balance shifts toward the harder group when per-group task difficulty differs.

Pith tools