REVIEW 4 major objections 6 minor 69 references
CAPRA recovers auditable medical-imaging subgroup failures from calibrated image-derived proxy axes when deployment metadata are missing.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 05:23 UTC pith:S6DCC7BW
load-bearing objection Solid methods paper that packages proxy axes, calibration, and reweighting into a reusable missing-metadata audit interface; multi-domain evidence is real, external shift evidence is thin. the 4 major comments →
Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Hidden subgroup analysis under missing metadata can be recast as calibrated proxy-axis risk estimation: image-derived semantic axes, after patient-level cross-fitted calibration on a small labeled split, form an interpretable interface that both exposes clinically meaningful failure modes and supplies group structure to robust learners without subgroup labels at deployment.
What carries the argument
CAPRA (Calibrated Proxy-Axis Risk Auditing): the calibrated subgroup interface of proxy-axis posteriors (temperature scaling plus confusion-matrix shrinkage under patient-level cross-fitting), reliability scores, and trust-weighted axis selection that concentrates robustness mass on axes that are both well calibrated and failure-relevant.
Load-bearing premise
A small metadata-labeled calibration set, handled with patient-level cross-fitting, is enough to turn imperfect image-based proxy predictions into posteriors that stay faithful under real deployment shift.
What would settle it
On a fresh multi-site external cohort never used for calibration or selection, CAPRA partitions show no better agreement with known failure axes and no higher support-filtered worst-group accuracy than image-only clustering or ExMap, even when the source calibration cohort still looks well calibrated.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CAPRA, a calibrated proxy-axis framework for hidden-subgroup analysis when demographic, acquisition, and quality metadata are unavailable at deployment. It learns image-derived semantic proxy axes, calibrates their posteriors on a small metadata-labeled cohort with patient-level cross-fitting (temperature scaling plus confusion-matrix shrinkage), and packages tokens, posteriors, reliability scores, and trust×disparity axis weights into a reusable interface I={t,q,r,α}. That interface is used both for standalone failure-aware adaptation and as input to downstream robust learners (GroupDRO, JTT, DFR, DPE, GSR). Experiments on BRSET, external handheld mBRSET, HAM10000, and CheXpert address six RQs covering downstream transfer, external shift, failure-axis portability, semantic alignment versus image-only/ExMap partitions, calibration budget, and interface ablations. The authors report improved support-filtered worst-group accuracy in most settings, stronger alignment with explicit failure axes than latent baselines, and domain-dependent reuse gains, while acknowledging that proxy axes are not ground-truth metadata and that external validation remains limited.
Significance. If the calibrated interface remains a faithful enough surrogate for hidden subgroup risk under realistic missing-metadata deployment, CAPRA would fill a genuine gap between group-robust optimization (which usually assumes known groups) and latent slice discovery (which often yields clinically opaque clusters). The problem setting is clinically relevant: metadata routinely disappear in imaging workflows, and aggregate metrics can mask failure modes. Strengths that should be credited include a clear axis-selective risk formulation that avoids unstable full Cartesian products; explicit leakage control (external audit never used for calibration/selection); patient-level splits; multi-modality evaluation on public datasets; systematic RQs with means±std; and an honest Limitations section on domain-dependent gains and non-universal failure taxonomies. The reusable-interface framing (audit plus optional robust transfer) is more useful than yet another standalone robust learner. The main significance risk is that the dual claim—deployment-time audit under shift and reusable robust transfer—currently rests on thin external evidence and mixed average-case deltas.
major comments (4)
- [§6.2 RQ2, Table 2] RQ2 / Table 2: The abstract and contributions claim that CAPRA “remains informative under dataset shift,” but the only external deployment-shift result is BRSET→mBRSET (one handheld fundus acquisition change). On matched BRSET the simpler CAPRA anchor is stronger on BA and WGA; full CAPRA wins only under that single external set. One modality-matched shift is insufficient to underwrite a general shift claim for a multi-domain methods paper. Either add at least one further external/protocol shift (e.g., multi-site CXR or dermoscopy device shift) with the same no-leakage protocol, or substantially soften the abstract/intro/conclusion language to “informative under the tested handheld fundus shift.”
- [§6.1 RQ1, Table 1] RQ1 / Table 1: CAPRA improves WGA in 14/15 comparisons, but BA deltas are mixed and sometimes large and negative (DFR/CheXpert −8.5 BA, −3.2 WGA; GSR/BRSET −4.3 BA; JTT/CheXpert −2.4 BA). The text correctly notes domain dependence, yet the reusable-transfer claim still needs a clearer failure analysis: under what measurable conditions does attaching I hurt average performance or even WGA? Without that, readers cannot decide when to deploy the interface versus leave a strong baseline (e.g., DFR on CheXpert) alone. A short diagnostic—e.g., relating negative transfer to axis-trust u_k, calibration ECE, or overlap between proxy buckets and true error mass—would make the central “reusable interface” claim operational rather than post-hoc.
- [§4.1 Calibrated Proxy-Axis Interface, Alg. 1] §4.1 and Algorithm 1: Proxy teachers T_k are load-bearing for every subsequent object (tokens, q, r, α), but the manuscript only says CAPRA “trains or obtains” teachers that produce logits over V_k. Architecture, training data, whether teachers see the same D_tr images, label source for each axis (age/sex vs quality/focus/artifacts), and teacher accuracy/ECE on D_cal are not reported. Without this, the faithfulness assumption behind Eq. (7) cannot be audited, and the pipeline is not reproducible. Please specify teacher construction per dataset/axis and report teacher calibration quality before and after the temperature+shrinkage step.
- [§6.5 RQ5, Fig. 4, Eq. (9)] RQ5 / Figure 4 and Eq. (9): The calibration-budget sweep (gains leveling after ~2–5% on BRSET/HAM10000) and the QUALITY-trust / AGE-weight pattern are only shown in-domain. The skeptic concern is exactly whether small-cohort calibrated posteriors remain faithful once acquisition/protocol shift changes which slices are hard. Because axis weights α_k ∝ u_k Δ_k(h_0) are fit on D_cal from the source, a source-specific shortcut that D_cal cannot correct would propagate into both audit and reweighting. A minimal stress test—re-estimate or freeze α_k under the mBRSET (or another) shift and report posterior reliability / worst-axis BA as a function of calibration fraction under shift—would directly address the weakest assumption of the paper.
minor comments (6)
- [Abstract] Abstract line “reveals disparity patterns missed by metadata-only slicing” is stronger than the main evidence, which primarily shows alignment with explicit failure axes and different dominant axes across domains (Fig. 2, Table 3). Clarify what metadata-only slicing would have missed in each dataset.
- [§3 Eq. (4), §5.2] Eq. (4) and §5.2: π_min and the default support rule for WGA versus the fixed @20 threshold in Fig. 2 should be stated numerically per dataset so support-filtered metrics are reproducible.
- [Fig. 3, §6.4] Figure 3 is described as “aligned 2D subgroup partitions” but the embedding method (e.g., UMAP/t-SNE on which features) is not specified in the main text; a one-sentence caption or methods note would help.
- [Table 1] Table 1 footnote marks GroupDRO* as a proxy-group variant; consider also stating in the table header or caption that no method receives oracle subgroup labels at training, to prevent misreading the large +CAPRA jumps as oracle GroupDRO gains.
- [§2 Related Work, References] Several related-work citations appear only loosely connected to medical imaging subgroup audit (e.g., control/PHD-filter and federated quantization references). Trim or relocate non-essential citations to keep the related-work narrative focused.
- [§3–§4] Notation: Z_k is introduced as proxy semantic axes sharing vocabulary with A_k, but hard predictions â_k and soft q^(k) are both used; a short glossary of I={t,q,r,α} early in §4 would reduce reader load.
Circularity Check
No significant circularity: CAPRA is an empirical interface-construction pipeline evaluated against held-out external shift and independent robust-learning baselines, not a first-principles derivation that reduces to its inputs.
full rationale
The paper builds a calibrated proxy-axis interface (proxy teachers → patient-level cross-fitted posteriors Eq. 7 → trust×disparity axis weights Eq. 9 → optional reweighting or transfer) and evaluates it empirically. Calibration and axis weights use a small held-out D_cal; external mBRSET is never used for calibration, selection, or tuning; downstream tables compare against image-only anchors and standard learners (GroupDRO/JTT/DFR/DPE/GSR). Higher agreement of CAPRA partitions with explicit metadata axes is expected once proxies are supervised/calibrated to those axes, but the paper also reports independent Gap, WGA, and external BA/Macro-F1 results that are not forced by that fit. Self-citations ([12],[23],[24],[34]) are peripheral and not load-bearing for the CAPRA claims. No equation equates a claimed prediction to its fitted input by construction; no uniqueness theorem or ansatz is imported from the authors to forbid alternatives. Fragility of small-cohort calibration under shift is an assumption/evidence-boundary issue, not circularity.
Axiom & Free-Parameter Ledger
free parameters (6)
- temperature τ_k per axis
- confusion-matrix shrinkage λ_k
- support threshold π_min
- excess-risk scale λ and weight clip ρ
- calibration budget fraction (e.g., 2–5% on BRSET)
- axis weights α_k via trust × disparity
axioms (5)
- domain assumption Image-derived proxy axes Z_k share vocabulary with explicit audit metadata A_k and can be predicted well enough from X to support audit after calibration.
- domain assumption Patient-level cross-fitting on a small D_cal yields calibrated posteriors without leakage into deployment audit.
- ad hoc to paper Support-filtered worst-axis risk is a valid primary robustness objective when full Cartesian subgroup products are statistically unstable.
- domain assumption Post-hoc temperature scaling plus confusion-matrix shrinkage is an adequate calibration model for discrete proxy axes.
- domain assumption Standard supervised risk and balanced accuracy metrics on public datasets reflect clinically meaningful subgroup failure.
invented entities (2)
-
Calibrated CAPRA subgroup interface I = {t, q, r, α}
independent evidence
-
Proxy semantic axes Z_k with reliability r_i^(k)
independent evidence
read the original abstract
Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.
Figures
Reference graph
Works this paper leans on
-
[1]
Marcus A Badgeley, John R Zech, Luke Oakden-Rayner, Benjamin S Glicksberg, Manway Liu, William Gale, Michael V McConnell, Bethany Percha, Thomas M Snyder, and Joel T Dudley. 2019. Deep learning predicts hip fracture using confounding patient and healthcare variables.NPJ digital medicine2, 1 (2019), 31
2019
-
[2]
Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Perez-Garcia, Maxim- ilian Ilse, Daniel C Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, et al. 2023. Learning to exploit temporal structure for biomedical vision-language processing. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 15016–15027
2023
-
[3]
Alceu Bissoto, Trung-Dung Hoang, Tim Flühmann, Susu Sun, Christian F Baum- gartner, and Lisa M Koch. 2025. Subgroup Performance Analysis in Hidden Stratifications. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 594–603
2025
-
[4]
Rwiddhi Chakraborty, Adrian Sletten, and Michael C Kampffmeyer. 2024. Exmap: Leveraging explainability heatmaps for unsupervised group robustness to spuri- ous correlations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12017–12026
2024
-
[5]
Jiawei Du, Jia Guo, Weihang Zhang, Shengzhu Yang, Hanruo Liu, Huiqi Li, and Ningli Wang. 2024. Ret-clip: A retinal image foundation model pre-trained with clinical diagnostic reports. InInternational conference on medical image computing and computer-assisted intervention. Springer, 709–719
2024
-
[6]
Grant Duffy, Shoa L Clarke, Matthew Christensen, Bryan He, Neal Yuan, Su- san Cheng, and David Ouyang. 2022. Confounders mediate AI prediction of demographics in medical imaging.NPJ digital medicine5, 1 (2022), 188
2022
-
[7]
Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Ré. 2022. Domino: Discovering systematic errors with cross-modal embeddings.arXiv preprint arXiv:2203.14960(2022)
Pith/arXiv arXiv 2022
-
[8]
Maria Galanty, Dieuwertje Luitse, Sijm H Noteboom, Philip Croon, Alexander P Vlaar, Thomas Poell, Clara I Sanchez, Tobias Blanke, and Ivana Išgum. 2024. Assessing the documentation of publicly available medical image and signal datasets and their impact on bias using the BEAMRAD tool.Scientific Reports14, 1 (2024), 31846
2024
-
[9]
Ben Glocker, Charles Jones, Mélanie Bernhardt, and Stefan Winzeck. 2023. Algo- rithmic encoding of protected characteristics in chest X-ray disease detection models.EBioMedicine89 (2023)
2023
-
[10]
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. InInternational conference on machine learning. PMLR, 1321–1330
2017
-
[11]
Shih-Cheng Huang, Liyue Shen, Matthew P Lungren, and Serena Yeung. 2021. Gloria: A multimodal global-local representation learning framework for label- efficient medical image recognition. InProceedings of the IEEE/CVF international conference on computer vision. 3942–3951
2021
-
[12]
Cuiying Huo, Di Jin, Yawen Li, Dongxiao He, Yu-Bin Yang, and Lingfei Wu
-
[13]
InProceedings of the AAAI Conference on Artificial Intelligence, Vol
T2-gnn: Graph neural networks for graphs with incomplete features and structure via teacher-student distillation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 4339–4346
-
[14]
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al
-
[15]
InProceedings of the AAAI conference on artificial intelligence, Vol
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. InProceedings of the AAAI conference on artificial intelligence, Vol. 33. 590–597
-
[16]
Saad M Khan, Xiaoxuan Liu, Siddharth Nath, Edward Korot, Livia Faes, Siegfried K Wagner, Pearse A Keane, Neil J Sebire, Matthew J Burton, and Alastair K Dennis- ton. 2021. A global review of publicly available datasets for ophthalmological imaging: barriers to access, usability, and generalisability.The Lancet Digital Health3, 1 (2021), e51–e66
2021
-
[17]
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson. 2022. Last layer re-training is sufficient for robustness to spurious correlations.arXiv preprint arXiv:2204.02937(2022)
Pith/arXiv arXiv 2022
-
[18]
Lisa M Koch, Christian F Baumgartner, and Philipp Berens. 2024. Distribution shift detection for the postmarket surveillance of medical AI algorithms: a retro- spective simulation study.NPJ Digital Medicine7, 1 (2024), 120
2024
-
[19]
Lisa M Koch, Christian M Schürch, Arthur Gretton, and Philipp Berens. 2022. Hidden in plain sight: Subgroup shifts escape OOD detection. InMedical Imaging with Deep Learning
2022
-
[20]
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pier- son, Been Kim, and Percy Liang. 2020. Concept bottleneck models. InInternational conference on machine learning. PMLR, 5338–5348
2020
-
[21]
Amar Kumar, Anita Kriz, Barak Pertzov, and Tal Arbel. 2025. Leveraging Vision- Language Foundation Models to Reveal Hidden Image-Attribute Relationships in Medical Imaging. InProceedings of the Computer Vision and Pattern Recognition Conference. 4840–4845
2025
-
[22]
Tyler LaBonte, Vidya Muthukumar, and Abhishek Kumar. 2023. Towards last- layer retraining for group robustness with fewer annotations.Advances in Neural Information Processing Systems36 (2023), 11552–11579
2023
-
[23]
Lin Li, Yingmin Jia, Junping Du, and Shiying Yuan. 2008. Robust L2–L ∞ control for uncertain singular systems with time-varying delay.Progress in Natural science18, 8 (2008), 1015–1021
2008
-
[24]
Wenling Li, Yingmin Jia, Junping Du, and Fashan Yu. 2013. Gaussian mixture PHD filter for multi-sensor multi-target tracking with registration errors.Signal Processing93, 1 (2013), 86–99
2013
-
[25]
Yawen Li, Wenling Li, and Zhe Xue. 2022. Federated learning with stochastic quantization.International Journal of Intelligent Systems37, 12 (2022), 11600– 11621
2022
-
[26]
Yawen Li, Liu Yang, Bohan Yang, Ning Wang, and Tian Wu. 2019. Application of interpretable machine learning models for the intelligent decision.Neurocomput- ing333 (2019), 273–283
2019
-
[27]
Hu Linmei, Tianchi Yang, Chuan Shi, Houye Ji, and Xiaoli Li. 2019. Heteroge- neous graph attention networks for semi-supervised short text classification. InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). 4821–4830
2019
-
[28]
Evan Z Liu, Behzad Haghgoo, Annie S Chen, Aditi Raghunathan, Pang Wei Koh, Shiori Sagawa, Percy Liang, and Chelsea Finn. 2021. Just train twice: Improving group robustness without training group information. InInternational Conference on Machine Learning. PMLR, 6781–6792
2021
-
[29]
William Lotter. 2024. Acquisition parameters influence AI recognition of race in chest x-rays and mitigating these factors reduces underdiagnosis bias.Nature communications15, 1 (2024), 7465
2024
-
[30]
Yan Luo, Min Shi, Muhammad Osama Khan, Muhammad Muneeb Afzal, Hao Huang, Shuaihang Yuan, Yu Tian, Luo Song, Ava Kouhana, Tobias Elze, et al
-
[31]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
Fairclip: Harnessing fairness in vision-language learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 12289–12301
-
[32]
Luis Filipe Nakayama, Mariana Goncalves, L Zago Ribeiro, Helen Santos, Daniel Ferraz, Fernando Malerbi, Leo Anthony Celi, and Caio Regatieri. 2023. A Brazilian multilabel ophthalmological dataset (BRSET).PhysioNet13026 (2023), 2
2023
-
[33]
Luis Filipe Nakayama, L Zago Ribeiro, D Restrepo, et al. 2024. mBRSET, a mobile Brazilian retinal dataset.PhysioNet https://doi. org/10.13026/QXPD-1Y65(2024)
-
[34]
Junhyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee, and Jinwoo Shin. 2020. Learning from failure: De-biasing classifier from biased classifier.Advances in Neural Information Processing Systems33 (2020), 20673–20684
2020
-
[35]
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré
-
[36]
InProceedings of the ACM conference on health, inference, and learning
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. InProceedings of the ACM conference on health, inference, and learning. 151–159
-
[37]
Vincent Olesen, Nina Weng, Aasa Feragen, and Eike Petersen. 2024. Slicing through bias: explaining performance gaps in medical image analysis using slice discovery methods. InMICCAI Workshop on Fairness of AI in Medical Imaging. Springer, 3–13
2024
-
[38]
Shilong Ou, Zhe Xue, Yawen Li, Meiyu Liang, Yuanqiang Cai, and Junjiang Wu
-
[39]
InProceedings of the IEEE/CVF conference on computer vision and pattern recognition
View-category interactive sharing transformer for incomplete multi-view multi-label learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 27467–27476
-
[40]
H Yi Paul, Tae Kyung Kim, Eliot Siegel, and Noushin Yahyavi-Firouz-Abadi
-
[41]
Journal of the American College of Radiology19, 1 (2022), 192–200
Demographic reporting in publicly available chest radiograph data sets: opportunities for mitigating sex and racial disparities in deep learning models. Journal of the American College of Radiology19, 1 (2022), 192–200
2022
-
[42]
Rui Qiao, Zhaoxuan Wu, Jingtan Wang, Pang Wei Koh, and Bryan Kian Hsiang Low. 2025. Group-robust sample reweighting for subpopulation shifts via influ- ence functions.arXiv preprint arXiv:2503.07315(2025)
Pith/arXiv arXiv 2025
-
[43]
Shikai Qiu, Andres Potapczynski, Pavel Izmailov, and Andrew Gordon Wilson
-
[44]
In International Conference on Machine Learning
Simple and fast group robustness by automatic feature reweighting. In International Conference on Machine Learning. PMLR, 28448–28467
-
[45]
María Agustina Ricci Lara, Rodrigo Echeveste, and Enzo Ferrante. 2022. Address- ing fairness in artificial intelligence for medical imaging.nature communications 13, 1 (2022), 4581
2022
-
[46]
Tim GJ Rudner, Ya Shi Zhang, Andrew Gordon Wilson, and Julia Kempe. 2024. Mind the gap: Improving robustness to subpopulation shifts with group-aware priors. InInternational Conference on Artificial Intelligence and Statistics. PMLR, 127–135
2024
-
[47]
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang. 2019. Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization.arXiv preprint arXiv:1911.08731 (2019)
Pith/arXiv arXiv 2019
-
[48]
Jarrel Seah, Cyril Tang, Quinlan D Buchlak, Michael Robert Milne, Xavier Holt, Hassan Ahmad, John Lambert, Nazanin Esmaili, Luke Oakden-Rayner, Peter Brotchie, et al. 2021. Do comprehensive deep learning algorithms suffer from hidden stratification? A retrospective study on pneumothorax detection in chest radiography.BMJ open11, 12 (2021), e053024
2021
-
[49]
Laleh Seyyed-Kalantari, Haoran Zhang, Matthew BA McDermott, Irene Y Chen, and Marzyeh Ghassemi. 2021. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Conference’17, July 2017, Washington, DC, USA Yawen Li, Yan Li, Zhe Xue, Yingxia Shao, Meiyu Liang, Guanhua Ye * Nature medicine27,...
2021
-
[50]
Danli Shi, Weiyi Zhang, Jiancheng Yang, Siyu Huang, Xiaolan Chen, Pusheng Xu, Kai Jin, Shan Lin, Jin Wei, Mayinuer Yusufu, et al . 2025. A multimodal visual–language foundation model for computational ophthalmology.npj digital medicine8, 1 (2025), 381
2025
-
[51]
Julio Silva-Rodriguez, Hadi Chakor, Riadh Kobbi, Jose Dolz, and Ismail Ben Ayed
-
[52]
A foundation language-image model of the retina (flair): Encoding expert knowledge in text supervision.Medical Image Analysis99 (2025), 103357
2025
-
[53]
Nimit Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu, and Christopher Ré. 2020. No subclass left behind: Fine-grained robustness in coarse-grained classification problems.Advances in Neural Information Processing Systems33 (2020), 19339–19352
2020
-
[54]
Weimin Tan, Qiaoling Wei, Zhen Xing, Hao Fu, Hongyu Kong, Yi Lu, Bo Yan, and Chen Zhao. 2024. Fairer AI in ophthalmology via implicit fairness learning for mitigating sexism and ageism.Nature communications15, 1 (2024), 4750
2024
-
[55]
Minh Nguyen Nhat To, Paul F RWilson, Viet Nguyen, Mohamed Harmanani, Michael Cooper, Fahimeh Fooladgar, Purang Abolmaesumi, Parvin Mousavi, and Rahul G Krishnan. 2025. Diverse Prototypical Ensembles Improve Robustness to Subpopulation Shift.arXiv preprint arXiv:2505.23027(2025)
Pith/arXiv arXiv 2025
-
[56]
Christopher J Tosh and Daniel Hsu. 2022. Simple and near-optimal algorithms for hidden stratification and multi-group learning. InInternational Conference on Machine Learning. PMLR, 21633–21657
2022
-
[57]
Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. 2018. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions.Scientific data5, 1 (2018), 180161
2018
-
[58]
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. 2022. Medclip: Contrastive learning from unpaired medical images and text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 3876–3887
2022
-
[59]
Shitao Xiao, Zheng Liu, Yingxia Shao, and Zhao Cao. 2022. RetroMAE: Pre- training retrieval-oriented language models via masked auto-encoder. InProceed- ings of the 2022 Conference on Empirical Methods in Natural Language Processing. 538–548
2022
-
[60]
Liang Xu, Junping Du, and Qingping Li. 2013. Image fusion based on nonsubsam- pled contourlet transform and saliency-motivated pulse coupled neural networks. Mathematical Problems in Engineering2013, 1 (2013), 135182
2013
-
[61]
Zikang Xu, Jun Li, Qingsong Yao, Han Li, Mingyue Zhao, and S Kevin Zhou
-
[62]
Addressing fairness issues in deep learning-based medical image analysis: a systematic review.npj Digital Medicine7, 1 (2024), 286
2024
-
[63]
Yuzhe Yang, Yujia Liu, Xin Liu, Avanti Gulhane, Domenico Mastrodicasa, Wei Wu, Edward J Wang, Dushyant Sahani, and Shwetak Patel. 2025. Demographic bias of expert-level vision-language foundation models in medical imaging.Science Advances11, 13 (2025), eadq0305
2025
-
[64]
Yuzhe Yang, Haoran Zhang, Judy W Gichoya, Dina Katabi, and Marzyeh Ghassemi
-
[65]
The limits of fair medical imaging AI in real-world generalization.Nature medicine30, 10 (2024), 2838–2848
2024
-
[66]
John R Zech, Marcus A Badgeley, Manway Liu, Anthony B Costa, Joseph J Titano, and Eric Karl Oermann. 2018. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: a cross-sectional study. PLoS medicine15, 11 (2018), e1002683
2018
-
[67]
Kai Zhang, Rong Zhou, Eashan Adhikarla, Zhiling Yan, Yixin Liu, Jun Yu, Zhengliang Liu, Xun Chen, Brian D Davison, Hui Ren, et al . 2024. A gener- alist vision–language foundation model for diverse biomedical tasks.Nature medicine30, 11 (2024), 3129–3141
2024
-
[68]
Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, et al . 2023. Biomed- clip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.arXiv preprint arXiv:2303.00915(2023)
Pith/arXiv arXiv 2023
-
[69]
Yongshuo Zong, Yongxin Yang, and Timothy Hospedales. 2022. MEDFAIR: bench- marking fairness for medical imaging.arXiv preprint arXiv:2210.01725(2022)
Pith/arXiv arXiv 2022
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.