Pith. sign in

REVIEW 4 major objections 6 minor 60 references

A source-only module improves eye-disease classification on unseen populations by damping edge shortcuts and scoring disease by angle, not activation strength.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 09:19 UTC pith:ABN5GWTG

load-bearing objection Solid source-only plug-in with consistent external F1 gains on a strict White-only FairVision protocol; compositional novelty, modest absolute numbers, and an under-checked edge prior, but still worth a referee. the 4 major comments →

arxiv 2607.10777 v1 pith:ABN5GWTG submitted 2026-07-12 cs.CV

RED-Sphere: Hyperspherical Residual Edge Debiasing for Cross-Population Fundus Disease Domain Generalization

classification cs.CV
keywords domain generalizationsource-only robustnessfundus imagingshortcut debiasinghyperspherical prototypesAMDdiabetic retinopathymedical image classification
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Medical classifiers trained on one patient group often fail when appearance, lighting, or disease mix shift in other groups. Fairness methods usually need demographic labels or target samples that clinics do not have at training time. RED-Sphere addresses a strict source-only setting: only one population is used for training, validation, and model selection; external groups appear only at test. The module estimates where population-linked low-level cues (edges and feature energy) are strong, soft-gates them while leaving a residual path so lesion structure is not erased, regularizes a masked nuisance view with consistency and separation losses, and classifies with normalized spherical prototypes so decisions rest on angular disease evidence rather than raw magnitude. On White-only training for AMD and DR fundus images, it raises held-out macro-F1 in every one of 20 backbone–task comparisons, with average gains of about 1.3 and 3.0 points, plus better ranking metrics and more aligned external embeddings.

Core claim

Under a strict White-only Harvard-FairVision 2D SLO protocol—external Asian and Black cohorts unseen for optimization, validation, scheduling, hyperparameters, and model selection—attaching RED-Sphere to a visual backbone improves held-out macro-F1 across all 20 task–backbone comparisons (average +1.28 F1 on AMD, +2.98 on DR), with gains in AUC and PR-AUC and visual evidence of stronger source–external semantic overlap and more coherent angular disease geometry.

What carries the argument

RED-Sphere: residual soft gating of an edge-and-feature-energy nuisance mask (preserving a direct residual path), counterfactual-inspired consistency and separation losses on the masked nuisance view during training only, and spherical prototype classification by cosine similarity of normalized embeddings, so prediction favors angular disease semantics over population-correlated activation magnitude.

Load-bearing premise

The method assumes that a Sobel edge map plus local feature-energy map, softly gated with a residual path, can damp population-linked appearance shortcuts without systematically erasing the same low-level edge structure that carries diagnostic lesions.

What would settle it

If, under the same White-only selection rule and matched backbones, RED-Sphere fails to raise held-out Asian/Black macro-F1 (and AUC/PR-AUC) relative to the base classifier on AMD or DR, or if lesion-bearing edges are visibly suppressed in the semantic stream while external disease ranking collapses, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes RED-Sphere, a plug-and-play module for source-only cross-population medical image classification. It estimates shortcut-sensitive responses via a Sobel edge map and local feature-energy prior, attenuates them with residual soft gating, regularizes a masked nuisance feature view with KL consistency and directional separation losses, and classifies with class-aware spherical prototypes. Under a strict White-only Harvard-FairVision 2D SLO protocol (Asian/Black cohorts fully held out from optimization, validation, scheduling, and model selection), the method is reported to improve held-out macro-F1 in all 20 task–backbone comparisons (average +1.28 F1 on AMD, +2.98 on DR), with supporting AUC/PR-AUC gains, ablations, sensitivity analyses, and semantic-manifold/sphere/margin visualizations.

Significance. Source-only cross-population robustness is a practically important and under-served setting in medical imaging, where target demographics are often unavailable before deployment. The experimental protocol is carefully designed and stronger than many fairness/DG evaluations: matched Base vs Ours comparisons across ten architectural families, three seeds, and fully held-out external cohorts. If the gains are real and the residual edge mechanism is doing what is claimed, RED-Sphere is a useful, backbone-agnostic robustness layer and a clear template for modality-specific nuisance priors. The combination of residual gating, feature-space nuisance regularization, and angular prototypes is incremental rather than foundational, but the controlled multi-backbone evidence and strict selection discipline are genuine strengths of the submission.

major comments (4)
  1. §3.3 (mask M=σ(h_φ([E;Q])) and residual gate eF=F⊙(1+ρ−αM)) and the abstract/conclusion claim that the method “preserves lesion structure” while attenuating population-correlated edges/textures: the paper itself states that lesions and shortcuts share edge/texture structure, yet there is no quantitative check that high-M regions are predominantly non-lesion, or that F_sem retains lesion-local evidence on Asian/Black images (e.g., lesion-mask overlap, expert-annotated ROI activation, or controlled lesion-edge retention metrics). Without this, the uniform Table 2 gains and Figure 4–6 diagnostics could largely reflect spherical prototypes, adaptive margins, and consistency losses rather than safe residual shortcut isolation. Either add such evidence or substantially soften the lesion-preservation and “modality-specific nuisance prior” transfer claims.
  2. §4.3 / Table 2: the only controlled comparison is matched Base vs RED-Sphere on the same backbone. Under the same White-only protocol, the paper does not report other source-only robustness/DG baselines (e.g., strong augmentation/MixStyle-style feature perturbation, RSC, SWAD, spherical prototypes alone, or simple magnitude normalization). Given that Table 3 shows non-trivial drops when removing the spherical classifier or adaptive margin, it is currently unclear how much of the all-20 F1 claim is due to residual edge debiasing versus angular geometry and training regularization. At least a small set of competitive source-only alternatives on the same splits is needed for the central methodological claim.
  3. Table 3 and §4.5–4.6: ablations remove the counterfactual branch, adaptive margin, and spherical classifier, and sensitivity varies ρ and α0, but there is no ablation of the edge–energy prior itself (edge-only, energy-only, random/uniform mask, or learned mask without E/Q conditioning). This is load-bearing for the paper’s distinctive mechanism. If a random residual gate plus spherical classification recovers most of the gain, the interpretation of RED-Sphere as residual edge debiasing is not supported. Please add this control on the same three representative backbones.
  4. Table 2 absolute levels and mixed ranking metrics: external AMD macro-F1 remains roughly 23–29% on a four-class task, and DR ConvNeXt-Tiny gains +5.68 F1 while AUC and PR-AUC decrease. The manuscript treats macro-F1 as primary and notes class imbalance, but does not adequately discuss clinical meaningfulness of ~1–3 average F1 points, calibration/operating-point shifts, or when ranking metrics disagree with F1. A short analysis of per-class external F1 (especially rare AMD grades / VTDR) and explicit discussion of these tradeoffs is needed before the “stronger external semantic alignment” claim can be accepted at face value.
minor comments (6)
  1. Inconsistent hyphenation and spacing throughout (e.g., “source only” vs “source-only”, “plug and play”, “T yne”, “T echnology”, “Y an Lin”). Normalize terminology and clean author/affiliation formatting.
  2. Figure 2 is described as a teaser of all 20 comparisons, but the main text does not state whether points are seed-averaged or best-seed; clarify and match Table 2 reporting.
  3. §3.5 margin formula uses source class counts n_c; state explicitly that this uses only White training counts and does not leak external prevalence.
  4. Implementation details defer full hyperparameters to supplementary material; for reproducibility, list λ_cf, λ_orth, λ_mask, α_min/α_max, m_min/m_max, and D in the main text or a compact table.
  5. Related work cites several fairness/DG methods; a short table contrasting required supervision (group labels, target data, image synthesis) would make the source-only positioning clearer.
  6. Figures 4–6 are informative but dense; define Mix/MMD/centroid-drift formulas in the caption or appendix and ensure color/marker legends remain readable in grayscale.

Circularity Check

0 steps flagged

No significant circularity: empirical plug-and-play module whose external F1/AUC gains are measured on cohorts excluded from every selection stage.

full rationale

RED-Sphere is an engineering framework (edge+energy prior M=σ(h_φ([E;Q])), residual gate eF=F⊙(1+ρ−αM), KL consistency + orthogonality on nuisance views, spherical prototypes with class-aware margin) whose central claim is empirical improvement under a strict source-only protocol. Asian and Black cohorts are never used for optimization, validation, scheduling, hyperparameter selection or model selection (§4.1, Table 1); White validation macro-F1 alone controls all selection. Table 2 therefore reports genuine held-out numbers, not quantities fitted to the reported external metrics. Ablations (Table 3) and sensitivity checks (§4.6) vary components or ρ/α0 and report signed changes; they do not redefine the target F1 by construction. Hyperparameters are ordinary White-validation tuning, not self-definitional fits that force the external scores. No uniqueness theorem, load-bearing self-citation chain, or ansatz-smuggled derivation appears; related-work citations supply background, not the result. The residual-gate and spherical-classifier equations are design choices, not reductions of a claimed first-principles prediction to its own inputs. Hence the derivation chain contains none of the enumerated circular patterns.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

The claim rests on standard CV building blocks plus domain assumptions about fundus shortcuts and several hand-chosen coefficients that control residual preservation, initial suppression, and loss weights. No new physical entity is postulated; the invented pieces are architectural constructs whose only evidence is the paper’s own ablations and held-out tables.

free parameters (5)
  • residual preservation coefficient ρ = 0.50
    Fixed to 0.50 after sensitivity analysis; controls how much original feature evidence bypasses the mask and is load-bearing for not erasing lesions.
  • initial suppression coefficient α0 (and clip bounds) = 0.20
    Initialized at 0.20 (best in 5/6 settings); sets how strongly high-mask responses are attenuated.
  • nuisance loss weights λ_cf, λ_orth, λ_mask = fixed (values in supplementary)
    Fixed coefficients balancing consistency, separation, and mean-mask sparsity; chosen on source validation and held fixed across tasks/backbones.
  • hyperspherical temperature s and class margins m_min/m_max = s=16; m from source class counts
    Temperature 16 and count-based adaptive margins shape angular decision geometry; not derived from first principles.
  • semantic embedding dimension D = 256
    Set to 256 for projection heads and prototypes; architectural free choice.
axioms (5)
  • domain assumption Retinal images encode population-correlated information via pigmentation, luminance, color, vessel maps, and edge/texture strength that can act as disease shortcuts.
    Invoked in Introduction and §3.3 with citations to FairVision and related ophthalmic studies; grounds the choice of edge/energy nuisance prior.
  • domain assumption Disease evidence (drusen, hemorrhages, exudates, lesion boundaries) shares edge/texture structure with those shortcuts, so hard removal is unsafe and residual gating is required.
    Central design premise of residual soft gating in §3.3; not independently proven, but clinically motivated.
  • domain assumption Angular similarity on normalized embeddings is less sensitive to illumination/contrast/population-correlated magnitude than raw feature norms.
    Justifies spherical prototype classifier §3.5 via CosFace/ArcFace/neural-collapse literature.
  • ad hoc to paper Feature-space masked nuisance views with KL consistency and directional separation are a safe substitute for image-level demographic counterfactuals under source-only constraints.
    §3.4 explicitly avoids image translation due to hallucination risk and defines L_cf and L_orth as the training-time substitute.
  • standard math Standard deep learning optimization (AdamW, focal loss, ImageNet-normalized 224 inputs) yields comparable Base vs Ours comparisons when only the RED-Sphere module differs.
    Implementation protocol §4.2; ordinary experimental control assumption.
invented entities (3)
  • Edge–feature-energy shortcut mask M = σ(h_φ([E;Q])) no independent evidence
    purpose: Estimate spatial locations of population-sensitive low-level responses without demographic labels.
    Constructed in §3.3 from Sobel magnitude and channel-averaged Gram energy; evidence is internal ablations and external F1 gains only.
  • Residual soft-gated semantic stream F_sem with refiners r_φ, c_φ no independent evidence
    purpose: Attenuate mask-dominant responses while preserving lesion-bearing residual features for deployment prediction.
    Core architectural object of RED-Sphere; no external validation outside this paper’s tables/figures.
  • Feature-space counterfactual-inspired nuisance view F_nuis with L_cf and L_orth no independent evidence
    purpose: Expose the classifier to shortcut-dominant directions during training without synthesizing images or using target populations.
    Defined only for training regularization in §3.4; ablation shows largest average drop when removed.

pith-pipeline@v1.1.0-grok45 · 23165 in / 4005 out tokens · 57995 ms · 2026-07-14T09:19:50.374129+00:00 · methodology

0 comments
read the original abstract

Medical image classifiers are often trained within one source population, yet clinical deployment requires robustness to patients whose appearance, acquisition style, and disease prevalence differ from the source cohort. Existing fairness and robustness methods often require group supervision or treat appearance variation as an undifferentiated nuisance, which is insufficient when population-correlated low-level cues and lesion evidence share edge and texture structure. We study a strict source-only cross-population setting, where external populations are unseen during optimization, validation, scheduling, hyperparameter and model selection. We propose RED-Sphere, a plug-and-play robustness framework for image classification under unseen population shifts. It estimates shortcut-sensitive nuisance responses with an edge and feature energy prior, attenuates dominant responses through residual soft gating, regularizes masked nuisance views with counterfactual-inspired consistency and separation losses, and predicts labels with normalized spherical prototypes. It favours angular semantic evidence over source-correlated activation magnitude while preserving lesion structure. Although demonstrated on 2D Scanning Laser Ophthalmoscopy (SLO) fundus classification for Age-Related Macular Degeneration (AMD) and Diabetic Retinopathy (DR), RED-Sphere is not tied to retinal anatomy: the same principle can be adapted with modality-specific nuisance priors wherever appearance shortcuts and semantic evidence are entangled. Under a strict White-only Harvard-FairVision protocol, RED-Sphere improves held-out macro-F1 across all 20 task and backbone comparisons, with average gains of 1.28 and 2.98 F1 points on AMD and DR. Gains in AUC and PR-AUC, visual diagnostics, ablations, and sensitivity analyses further support stronger external semantic alignment and more stable angular disease geometry.

Figures

Figures reproduced from arXiv: 2607.10777 by Amir Atapour-Abarghouei, Shuang Chen, Stephen McGough, Yan Lin, Ziheng Wang.

Figure 1
Figure 1. Figure 1: Source only cross population fundus disease classification. Training and model [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Teaser comparison between matched source only baselines and RED-Sphere on [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Overview of the proposed RED-Sphere framework. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Cross-cohort semantic manifold overlap under RED-Sphere on ConvNeXt-Tiny, [PITH_FULL_IMAGE:figures/full_fig_p012_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Spherical decision geometry on AMD and DR for ConvNeXt-Tiny, Swin-T, and [PITH_FULL_IMAGE:figures/full_fig_p013_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Logit margin densities under RED-Sphere across White, Asian, and Black cohorts. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

60 extracted references · 9 canonical work pages · 1 internal anchor

  1. [1]

    Distribution matching losses can hallucinate features in medical image translation

    Joseph Paul Cohen, Margaux Luck, and Sina Honari. Distribution matching losses can hallucinate features in medical image translation. In Alejandro F. Frangi, Ju- lia A. Schnabel, Christos Davatzikos, Carlos Alberola-López, and Gabor Fichtinger, editors,Medical Image Computing and Computer Assisted Intervention - MICCAI 2018 - 21st International Conference...

  2. [2]

    URLhttps://doi.org/10

    doi: 10.1007/978-3-030-00928-1\_60. URLhttps://doi.org/10. 1007/978-3-030-00928-1_60

  3. [3]

    Association of biomarker-based artificial intelligence with risk of racial bias in retinal images.JAMA ophthalmology, 141(6):543–552, 2023

    Aaron S Coyner, Praveer Singh, James M Brown, Susan Ostmo, RV Paul Chan, Michael F Chiang, Jayashree Kalpathy-Cramer, and J Peter Campbell. Association of biomarker-based artificial intelligence with risk of racial bias in retinal images.JAMA ophthalmology, 141(6):543–552, 2023

  4. [4]

    Arcface: Additive angular margin loss for deep face recognition.IEEE Trans

    Jiankang Deng, Jia Guo, Jing Yang, Niannan Xue, Irene Kotsia, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition.IEEE Trans. Pattern Anal. Mach. Intell., 44(10):5962–5979, 2022. doi: 10.1109/TPAMI.2021.3087709. URLhttps://doi.org/10.1109/TPAMI.2021.3087709

  5. [5]

    An image is worth 16x16 words: Trans- formers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Trans- formers for image recognition at scale. In9th International Conference on Learning Representations, ICLR 202...

  6. [6]

    URLhttps://openreview.net/forum?id=YicbFdNTTy

  7. [7]

    Tsaftaris, and Timothy M

    Raman Dutt, Ondrej Bohdal, Sotirios A. Tsaftaris, and Timothy M. Hospedales. Fair- tune: Optimizing parameter efficient fine tuning for fairness in medical image anal- ysis. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps: //openreview.net/forum?id=ArpwmicoYW

  8. [8]

    Wich- mann, and Wieland Brendel

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wich- mann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; in- creasing shape bias improves accuracy and robustness. In7th International Con- ference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9,

  9. [9]

    URLhttps://openreview.net/forum?id= Bygh9j09KX

    OpenReview.net, 2019. URLhttps://openreview.net/forum?id= Bygh9j09KX. 16LIN ET AL.: RED-SPHERE

  10. [10]

    Ai recognition of patient race in medical imaging: a modelling study.The Lancet Digital Health, 4(6):e406–e414, 2022

    Judy Wawira Gichoya, Imon Banerjee, Ananth Reddy Bhimireddy, John L Burns, Leo Anthony Celi, Li-Ching Chen, Ramon Correa, Natalie Dullerud, Marzyeh Ghas- semi, Shih-Cheng Huang, et al. Ai recognition of patient race in medical imaging: a modelling study.The Lancet Digital Health, 4(6):e406–e414, 2022

  11. [11]

    Algorith- mic encoding of protected characteristics in chest x-ray disease detection models

    Ben Glocker, Charles Jones, Mélanie Bernhardt, and Stefan Winzeck. Algorith- mic encoding of protected characteristics in chest x-ray disease detection models. EBioMedicine, 89, 2023

  12. [12]

    In search of lost domain generalization

    Ishaan Gulrajani and David Lopez-Paz. In search of lost domain generalization. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021. URLhttps://openreview. net/forum?id=lQdXeXDoWtI

  13. [14]

    Equality of opportunity in super- vised learning

    Moritz Hardt, Eric Price, and Nati Srebro. Equality of opportunity in super- vised learning. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors,Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Process- ing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages...

  14. [15]

    URLhttps://proceedings.neurips.cc/paper/2016/hash/ 9d2682367c3935defcb1f9e247a97c0d-Abstract.html

  15. [16]

    Mambavision: A hybrid mamba-transformer vision backbone

    Ali Hatamizadeh and Jan Kautz. Mambavision: A hybrid mamba-transformer vision backbone. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pages 25261–25270. Com- puter Vision Foundation / IEEE, 2025. doi: 10.1109/CVPR52734.2025.02352. URLhttps://openaccess.thecvf.com/content/CVPR2025/ html/H...

  16. [17]

    Haibo He and Edwardo A. Garcia. Learning from imbalanced data.IEEE Trans. Knowl. Data Eng., 21(9):1263–1284, 2009. doi: 10.1109/TKDE.2008.239. URLhttps: //doi.org/10.1109/TKDE.2008.239

  17. [18]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV , USA, June 27-30, 2016, pages 770–778. IEEE Computer Society, 2016. doi: 10.1109/CVPR.2016.90. URLhttps://doi. org/10.1109/CVPR.2016.90

  18. [19]

    Weinberger

    Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger. Densely connected convolutional networks. In2017 IEEE Conference on Computer Vision and LIN ET AL.: RED-SPHERE17 Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 2261–

  19. [20]

    doi: 10.1109/CVPR.2017.243

    IEEE Computer Society, 2017. doi: 10.1109/CVPR.2017.243. URLhttps: //doi.org/10.1109/CVPR.2017.243

  20. [21]

    Re- visiting masked image modeling with standardized color space for domain generalized fundus photography classification

    Eojin Jang, Myeongkyun Kang, Soopil Kim, Min Sagong, and Sang Hyun Park. Re- visiting masked image modeling with standardized color space for domain generalized fundus photography classification. In James C. Gee, Daniel C. Alexander, Jaesung Hong, Juan Eugenio Iglesias, Carole H. Sudre, Archana Venkataraman, Polina Gol- land, Jong Hyo Kim, and Jinah Park,...

  21. [22]

    Kevin Zhou, and Xiaoxiao Li

    Ruinan Jin, Zikang Xu, Yuan Zhong, Qingsong Yao, Qi Dou, S. Kevin Zhou, and Xiaoxiao Li. Fairmedfm: Fairness benchmarking for medical imaging foundation models. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors,Advances in Neural Information Processing Systems 37: Annual Conference ...

  22. [23]

    PRISM: High-Resolution & Precise Counterfactual Medical Image Generation using Language-guided Stable Diffusion

    Amar Kumar, Anita Kriz, Mohammad Havaei, and Tal Arbel. PRISM: high-resolution & precise counterfactual medical image generation using language-guided stable dif- fusion.CoRR, abs/2503.00196, 2025. doi: 10.48550/ARXIV .2503.00196. URL https://doi.org/10.48550/arXiv.2503.00196

  23. [24]

    Kusner, Joshua R

    Matt J. Kusner, Joshua R. Loftus, Chris Russell, and Ricardo Silva. Counterfactual fairness. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wal- lach, Rob Fergus, S. V . N. Vishwanathan, and Roman Garnett, editors,Advances in Neural Information Processing Systems 30: Annual Conference on Neural In- formation Processing Systems 2017, December...

  24. [25]

    Let samples speak: Mitigating spurious correlation by exploiting the clusterness of samples

    Weiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng, Yahao Liu, and Wen Li. Let samples speak: Mitigating spurious correlation by exploiting the clusterness of samples. InIEEE/CVF Conference on Computer Vision and Pattern Recog- nition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025, pages 15486–15496. Computer Vision Foundation / IEEE, 2025. doi: 10.1109...

  25. [26]

    Un- certainty modeling for out-of-distribution generalization

    Xiaotong Li, Yongxing Dai, Yixiao Ge, Jun Liu, Ying Shan, and Lingyu Duan. Un- certainty modeling for out-of-distribution generalization. InThe Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 18LIN ET AL.: RED-SPHERE

  26. [27]

    URLhttps://openreview.net/forum?id= 6HN7LHyzGgC

    OpenReview.net, 2022. URLhttps://openreview.net/forum?id= 6HN7LHyzGgC

  27. [28]

    Incomplete modality disentangled representa- tion for ophthalmic disease grading and diagnosis

    Chengzhi Liu, Zile Huang, Zhe Chen, Feilong Tang, Yu Tian, Zhongxing Xu, Zihong Luo, Yalin Zheng, and Yanda Meng. Incomplete modality disentangled representa- tion for ophthalmic disease grading and diagnosis. In Toby Walsh, Julie Shah, and Zico Kolter, editors,Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty- Seventh Conference on Innovati...

  28. [29]

    Sphereface revived: Unifying hyperspherical face recognition.IEEE Trans

    Weiyang Liu, Yandong Wen, Bhiksha Raj, Rita Singh, and Adrian Weller. Sphereface revived: Unifying hyperspherical face recognition.IEEE Trans. Pattern Anal. Mach. Intell., 45(2):2458–2474, 2023. doi: 10.1109/TPAMI.2022.3159732. URLhttps: //doi.org/10.1109/TPAMI.2022.3159732

  29. [30]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 9992–10002. IEEE,

  30. [31]

    URLhttps://doi.org/10

    doi: 10.1109/ICCV48922.2021.00986. URLhttps://doi.org/10. 1109/ICCV48922.2021.00986

  31. [32]

    A convnet for the 2020s

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 11966–11976. IEEE, 2022. doi: 10.1109/CVPR52688.2022.01167. URLhttps: //doi.org/10.1109/CVPR52688.2022.01167

  32. [33]

    Fairvision: Equitable deep learning for eye disease screening via fair identity scaling, 2024

    Yan Luo, Muhammad Osama Khan, Yu Tian, Min Shi, Zehao Dou, Tobias Elze, Yi Fang, and Mengyu Wang. Fairvision: Equitable deep learning for eye disease screening via fair identity scaling, 2024. URLhttps://arxiv.org/abs/2310. 02492

  33. [34]

    Fairclip: Harnessing fairness in vision-language learning

    Yan Luo, Min Shi, Muhammad Osama Khan, Muhammad Muneeb Afzal, Hao Huang, Shuaihang Yuan, Yu Tian, Luo Song, Ava Kouhana, Tobias Elze, Yi Fang, and Mengyu Wang. Fairclip: Harnessing fairness in vision-language learning. InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 12289–12301. IEEE...

  34. [35]

    Shufflenet V2: prac- tical guidelines for efficient CNN architecture design

    Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun. Shufflenet V2: prac- tical guidelines for efficient CNN architecture design. In Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss, editors,Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Pro- ceedings, Part XIV, Lecture Notes in Co...

  35. [36]

    URLhttps://doi.org/10.1007/ 978-3-030-01264-9_8

    doi: 10.1007/978-3-030-01264-9\_8. URLhttps://doi.org/10.1007/ 978-3-030-01264-9_8. LIN ET AL.: RED-SPHERE19

  36. [37]

    Hyperspherical prototype net- works

    Pascal Mettes, Elise van der Pol, and Cees Snoek. Hyperspherical prototype net- works. In Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché- Buc, Emily B. Fox, and Roman Garnett, editors,Advances in Neural Informa- tion Processing Systems 32: Annual Conference on Neural Information Process- ing Systems 2019, NeurIPS 2019, December 8-1...

  37. [38]

    Slicing through bias: Explaining performance gaps in medical image analysis using slice discovery meth- ods

    Vincent Olesen, Nina Weng, Aasa Feragen, and Eike Petersen. Slicing through bias: Explaining performance gaps in medical image analysis using slice discovery meth- ods. In Esther Puyol-Antón, Ghada Zamzmi, Aasa Feragen, Andrew P. King, Veronika Cheplygina, Melanie Ganz-Benjaminsen, Enzo Ferrante, Ben Glocker, Eike Petersen, John S. H. Baxter, Islem Rekik,...

  38. [39]

    Vardan Papyan, X. Y . Han, and David L. Donoho. Prevalence of neural collapse during the terminal phase of deep learning training.CoRR, abs/2008.08186, 2020. URL https://arxiv.org/abs/2008.08186

  39. [40]

    Girshick, Kaiming He, and Piotr Dollár

    Ilija Radosavovic, Raj Prateek Kosaraju, Ross B. Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020, pages 10425–10433. Computer Vision Foundation / IEEE, 2020. doi: 10.1109/CVPR42600.2020.01044. URLhttps://openaccess....

  40. [41]

    Machine learning derived retinal pigment score from ophthalmic imaging shows ethnicity is not biology.Nature communications, 16(1):60, 2025

    Anand E Rajesh, Abraham Olvera-Barrios, Alasdair N Warwick, Yue Wu, Kelsey V Stuart, Mahantesh I Biradar, Chuin Ying Ung, Anthony P Khawaja, Robert Luben, Paul J Foster, et al. Machine learning derived retinal pigment score from ophthalmic imaging shows ethnicity is not biology.Nature communications, 16(1):60, 2025. doi: 10.1038/s41467-024-55198-7

  41. [42]

    Hashimoto, and Percy Liang

    Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang. Distri- butionally robust neural networks for group shifts: On the importance of regular- ization for worst-case generalization.CoRR, abs/1911.08731, 2019. URLhttp: //arxiv.org/abs/1911.08731

  42. [43]

    The precision-recall plot is more informative than the roc plot when evaluating binary classifiers on imbalanced datasets.PloS one, 10 (3):e0118432, 2015

    Takaya Saito and Marc Rehmsmeier. The precision-recall plot is more informative than the roc plot when evaluating binary classifiers on imbalanced datasets.PloS one, 10 (3):e0118432, 2015

  43. [44]

    vmfcoop: Towards equilibrium on a unified hyperspherical manifold for prompt- ing biomedical vlms

    Minye Shao, Sihan Guo, Xinrun Li, Xingyu Miao, Haoran Duan, and Yang Long. vmfcoop: Towards equilibrium on a unified hyperspherical manifold for prompt- ing biomedical vlms. In Sven Koenig, Chad Jenkins, and Matthew E. Taylor, ed- itors,Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference 20LIN ET AL.: RED-SPHERE on Innovative App...

  44. [45]

    A 3x3 isotropic gradient operator for image process- ing.a talk at the Stanford Artificial Project in, 1968:271–272, 1968

    Irwin Sobel, Gary Feldman, et al. A 3x3 isotropic gradient operator for image process- ing.a talk at the Stanford Artificial Project in, 1968:271–272, 1968

  45. [46]

    Mingxing Tan and Quoc V . Le. Efficientnet: Rethinking model scaling for convolu- tional neural networks. In Kamalika Chaudhuri and Ruslan Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, Proceedings of Machine Learning Research, pages 6105–6114. PMLR, 201...

  46. [47]

    Fairdomain: Achiev- ing fairness in cross-domain medical image segmentation and classification

    Yu Tian, Congcong Wen, Min Shi, Muhammad Muneeb Afzal, Hao Huang, Muham- mad Osama Khan, Yan Luo, Yi Fang, and Mengyu Wang. Fairdomain: Achiev- ing fairness in cross-domain medical image segmentation and classification. In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, editors,Computer Vision - ECCV 2024 - 18th...

  47. [48]

    Bovik, and Yinxiao Li

    Zhengzhong Tu, Hossein Talebi, Han Zhang, Feng Yang, Peyman Milanfar, Alan C. Bovik, and Yinxiao Li. Maxvit: Multi-axis vision transformer. In Shai Avidan, Gabriel J. Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner, editors,Computer Vision - ECCV 2022: 17th European Conference, Tel Aviv, Is- rael, October 23-27, 2022, Proceedings, Part...

  48. [49]

    Gilligan-Lee

    Athanasios Vlontzos, Bernhard Kainz, and Ciarán M. Gilligan-Lee. Estimating cat- egorical counterfactuals via deep twin networks.Nat. Mac. Intell., 5(2):159–168,

  49. [50]

    URLhttps://doi.org/10.1038/ s42256-023-00611-x

    doi: 10.1038/S42256-023-00611-X. URLhttps://doi.org/10.1038/ s42256-023-00611-x

  50. [51]

    Cosface: Large margin cosine loss for deep face recog- nition

    Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. Cosface: Large margin cosine loss for deep face recog- nition. In2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 5265–5274. Com- puter Vision Foundation / IEEE Computer Society, 2018....

  51. [52]

    Navigate beyond shortcuts: Debiased learning through the lens of neural collapse

    Yining Wang, Junjie Sun, Chenyue Wang, Mi Zhang, and Min Yang. Navigate beyond shortcuts: Debiased learning through the lens of neural collapse. InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pages 12322–12331. IEEE, 2024. doi: 10.1109/CVPR52733.2024.01171. URLhttps://doi.org/10.1109/CVPR...

  52. [53]

    Fast diffusion-based counterfactuals for shortcut removal and generation

    Nina Weng, Paraskevas Pegios, Eike Petersen, Aasa Feragen, and Siavash Arjomand Bigdeli. Fast diffusion-based counterfactuals for shortcut removal and generation. In Ales Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, editors,Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October ...

  53. [54]

    Generalizing to unseen domains in diabetic retinopa- thy with disentangled representations

    Peng Xia, Ming Hu, Feilong Tang, Wenxue Li, Wenhao Zheng, Lie Ju, Peibo Duan, Huaxiu Yao, and Zongyuan Ge. Generalizing to unseen domains in diabetic retinopa- thy with disentangled representations. In Marius George Linguraru, Qi Dou, Aasa Feragen, Stamatia Giannarou, Ben Glocker, Karim Lekadir, and Julia A. Schnabel, editors,Medical Image Computing and C...

  54. [55]

    URLhttps://doi.org/10

    doi: 10.1007/978-3-031-72117-5\_40. URLhttps://doi.org/10. 1007/978-3-031-72117-5_40

  55. [56]

    The limits of fair medical imaging ai in real-world generalization.Nature medicine, 30 (10):2838–2848, 2024

    Yuzhe Yang, Haoran Zhang, Judy W Gichoya, Dina Katabi, and Marzyeh Ghassemi. The limits of fair medical imaging ai in real-world generalization.Nature medicine, 30 (10):2838–2848, 2024

  56. [57]

    Mitigating unwanted bi- ases with adversarial learning

    Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. Mitigating unwanted bi- ases with adversarial learning. In Jason Furman, Gary E. Marchant, Huw Price, and Francesca Rossi, editors,Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, AIES 2018, New Orleans, LA, USA, February 02-03, 2018, pages 335–

  57. [58]

    doi: 10.1145/3278721.3278779

    ACM, 2018. doi: 10.1145/3278721.3278779. URLhttps://doi.org/10. 1145/3278721.3278779

  58. [59]

    Exact feature distri- bution matching for arbitrary style transfer and domain generalization

    Yabin Zhang, Minghan Li, Ruihuang Li, Kui Jia, and Lei Zhang. Exact feature distri- bution matching for arbitrary style transfer and domain generalization. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 8025–8035. IEEE, 2022. doi: 10.1109/CVPR52688. 2022.00787. URLhttps://doi.org/...

  59. [60]

    Prototype correlation matching and class- relation reasoning for few-shot medi- cal image segmentation.IEEE Trans

    Yumin Zhang, Hongliu Li, Yajun Gao, Haoran Duan, Yawen Huang, and Yefeng Zheng. Prototype correlation matching and class- relation reasoning for few-shot medi- cal image segmentation.IEEE Trans. Medical Imaging, 43(11):4041–4054, 2024. doi: 10.1109/TMI.2024.3412420. URLhttps://doi.org/10.1109/TMI.2024. 3412420

  60. [61]

    Discover and mit- igate multiple biased subgroups in image classifiers

    Zeliang Zhang, Mingqian Feng, Zhiheng Li, and Chenliang Xu. Discover and mit- igate multiple biased subgroups in image classifiers. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16- 22, 2024, pages 10906–10915. IEEE, 2024. doi: 10.1109/CVPR52733.2024.01037. URLhttps://doi.org/10.1109/CVPR52733.2024.01037