Pith. sign in

REVIEW 1 cited by

Explaining medical AI performance disparities across sites with confounder Shapley value analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.08168 v1 pith:DV6TGT2X submitted 2021-11-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords performanceacrossdisparitiesmodelsiteswhenalgorithmsbiases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Medical AI algorithms can often experience degraded performance when evaluated on previously unseen sites. Addressing cross-site performance disparities is key to ensuring that AI is equitable and effective when deployed on diverse patient populations. Multi-site evaluations are key to diagnosing such disparities as they can test algorithms across a broader range of potential biases such as patient demographics, equipment types, and technical parameters. However, such tests do not explain why the model performs worse. Our framework provides a method for quantifying the marginal and cumulative effect of each type of bias on the overall performance difference when a model is evaluated on external data. We demonstrate its usefulness in a case study of a deep learning model trained to detect the presence of pneumothorax, where our framework can help explain up to 60% of the discrepancy in performance across different sites with known biases like disease comorbidities and imaging parameters.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. "Who experiences large model decay and why?" A Hierarchical Framework for Diagnosing Heterogeneous Performance Drift

    cs.LG 2025-05 conditional novelty 7.0 of 10

    SHIFT is a hierarchical hypothesis-testing method that detects subgroups with large model performance decay under distribution shift and explains the decay via variable-subset-specific covariate or outcome shifts.

Pith tools