REVIEW 4 major objections 5 minor 49 references
TrustSkin: A Fairness Pipeline for Trustworthy Facial Affect Analysis Across Skin Tone
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The skin-tone scale chosen for an emotion-AI audit can hide bias against darker-skinned faces.
desk verdict A useful first comparison of skin-tone taxonomies in FAA, but the central F1-disparity claim reverses when recomputed from the paper's own Table IV. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a pair of skin-tone classifiers and the fairness metrics that compare them. ITA, defined as $\mathrm{ITA} = \arctan((L^*-50)/b^*) \cdot 180/\pi$, uses lightness and blue-yellow chroma but drops the green-red axis ($a^*$). The alternative $H^*$-$L^*$ method uses lightness $L^*$ plus hue angle $\mathrm{Hue} = \arctan(b^*/a^*)$, with thresholds (Light $L^*>67$, Medium $37 \leq L^* \leq 67$, Dark $L^*<37$) and a brown-tone override rule that reclassifies pixels with certain RGB ranges and hue between $20^\circ$ and $50^\circ$ as Dark regardless of $L^*$. That override is what lets the method catch dark-skinned faces whose lightness was raised by lighting, and the fairness metrics—F1 gap, accuracy range, and Equal Opportunity difference—turn each taxonomy into quantitative audit conclusions.
What would settle it
Collect a sample of faces with measured skin tone (e.g., spectrophotometer or self-report), compare ITA and hue-lightness labels against those measurements, and recompute the fairness gaps using only faces both methods label correctly; if ITA agrees with measured tone as often as hue-lightness on dark faces, the larger hue-lightness disparity would be a grouping artifact rather than evidence of hidden bias.
Extended reading notes
Core claim
The paper's central discovery is that the skin-tone taxonomy used for fairness evaluation is itself a source of variation in the conclusions. On the AffectNet test set, the $H^*$-$L^*$ grouping produces an F1-score disparity of $0.080$ and a true-positive-rate disparity of $0.106$, while ITA grouping produces $0.050$ and $0.091$, respectively, and the Equal Opportunity heatmaps differ qualitatively. The paper attributes the difference to ITA pulling medium and light faces into the dark group under illumination changes, so that ITA's dark-group recall is partly inflated; by including hue and a brown-tone override, $H^*$-$L^*$ isolates true dark faces and thereby exposes larger, previously masked disparities. The paper concludes that $H^*$-$L^*$ gives more consistent subgrouping and that ITA-based evaluations may overlook disparities affecting darker-skinned individuals.
Load-bearing premise
The whole conclusion depends on the assumption that the hue-lightness method labels dark skin more accurately than ITA does, because the rule that overrides lightness for brown tones was tuned on two borderline cases from the same dataset rather than checked against real ground-truth skin tone.
Editorial extensions
If this is right
- Fairness audits of facial affect systems that use only ITA should be rechecked with a hue-aware measure, because the choice of taxonomy changes which groups look disadvantaged.
- The severe underrepresentation of dark skin in AffectNet (roughly 2% of samples) means reported dark-group metrics are statistically fragile and should be treated as diagnostic hints rather than stable estimates.
- The larger F1 and TPR disparities seen with $H^*$-$L^*$ should be read as increased visibility of bias, not as evidence that the model treats dark skin worse than ITA suggested.
- Equal Opportunity diagnostics can help separate real subgroup gaps from grouping artifacts, since ITA's dark-group recall is partly inflated by misclassified medium and light faces.
- Rule-based corrections like the brown-tone override improve subgroup fidelity but need dataset-specific tuning before being reused.
Reading between the lines
- Editorial inference: if hue-aware measures become standard, previously published ITA-based fairness conclusions in emotion recognition may need to be revisited, and some audits that looked acceptable could flip to showing bias.
- Editorial inference: the same measurement-choice effect likely extends beyond emotion recognition to face recognition and dermatology datasets that rely on colorimetric skin-tone proxies, so the masking may be widespread.
- Editorial inference: a direct test of the paper's interpretation would be to collect spectrophotometer or self-reported skin-tone labels for a sample of faces and compare ITA and $H^*$-$L^*$ agreement; if ITA labels are equally accurate for dark skin, the larger $H^*$-$L^*$ disparity would be a grouping artifact rather than a revelation.
- Editorial inference: re-running the audit on a balanced test set with many more dark-skin samples would show whether the larger $H^*$-$L^*$ disparity persists or shrinks into a small-sample artifact.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper compares two objective skin-tone stratification methods—ITA and an L*–H* (Hue–Lightness) method—for fairness evaluation of facial affect analysis. Using AffectNet and a MobileNet classifier, it reports underrepresentation of dark skin tones, per-group emotion metrics, F1/TPR disparities, Equal Opportunity analysis, and Grad-CAM visualizations. The central claim is that the choice of skin-tone measurement changes fairness conclusions and that ITA-based audits may underestimate disparities affecting darker-skinned individuals.
Significance. The paper addresses a timely and under-explored question: whether the operationalization of skin tone changes fairness audits in facial affect analysis. Its proposed modular pipeline, explicit comparison of ITA and H*L*, and use of Equal Opportunity and Grad-CAM are useful ingredients. The authors are candid about limitations such as empirical thresholds and the tiny dark-skin sample. However, the quantitative support for the headline claim is undermined by an internal inconsistency between Figure 6 and Table IV, and by the lack of external validation for the labeling rule. If the numerical issues are resolved, the paper could make a modest but useful contribution; in its current form the central quantitative claim is not reproducible from the supplied data.
major comments (4)
- [Section IV, Figure 6 vs. Table IV] The F1-disparity values in Figure 6 do not reproduce from the paper's own Table IV under the definition in Section II.C. Averaging the eight per-emotion F1 scores in Table IV gives H*L* Light=0.414, Medium=0.388, Dark=0.398 (gap 0.026), while ITA Light=0.403, Medium=0.394, Dark=0.478 (gap 0.084). Thus Table IV implies the opposite ordering from Figure 6: ITA has the larger macro-F1 gap, and the Dark group is the best ITA group, not a disadvantaged one. This directly affects the abstract's 'up to 0.08' claim and the Section IV interpretation that H*L* has a more pronounced F1 disparity. Please recompute and correct Figure 6 and the accompanying text, or report the exact computation that produced the printed values.
- [Section II.B] The brown-tone override is introduced with thresholds for RGB ranges and hue, and the only evidence offered is that it 'correctly reassigning two borderline cases' in the same dataset. Since the same override is then used to argue that H*L* gives more accurate dark-skin labels and hence better fairness diagnostics, this is in-sample validation. An independent ground truth (spectrophotometer, self-report, or manually labeled held-out set) or at least a sensitivity analysis over the override thresholds is needed before the claim that ITA underestimates dark-skin disparities can be supported.
- [Section III.A, Tables II-IV] The Dark group contains only 52 test images, yet the paper reports per-emotion precision/recall/F1 without confidence intervals. Values such as Dark H*L* Disgust precision 1.00, Surprise precision 0.12, and Neutral recall 1.00 are likely based on very small denominators and should be accompanied by counts, confidence intervals, or bootstrap estimates. Without this, the strong qualitative claims about Dark-group performance in Section III.A and Section IV.A are not statistically grounded.
- [Abstract / Section I vs. Section V] The abstract and Section I state that 'targeted data augmentation failed to improve fairness,' but Section V and the Ethical Impact Statement say that no dataset rebalancing was applied in this work. These statements cannot both describe the executed experiments. Please clarify whether augmentation experiments were run and, if so, report them; otherwise remove the augmentation claim from the abstract and introduction.
minor comments (5)
- [Tables II-IV] The emotion row uses 'Neutre' instead of 'Neutral'; fix the typo for consistency with the rest of the paper.
- [Figure 6] The caption appears incomplete ('ITA ITAH*-L* H*-L*') and the panels are unlabeled; add a clear caption and axis labels so the reader can distinguish the F1 and TPR panels.
- [Section II.B] The Hue formula is written as arctan(b*/a*); specify that a four-quadrant arctangent (atan2) is used, since a* can be negative and the angle range matters for the stated 20–50 degree interval.
- [Throughout] The notation 'H*L*' and 'L*H*' is used interchangeably; choose one form and use it consistently.
- [Section II.A] The paper refers to 'validated thresholds from prior work [16], [31]' but does not list the actual segmentation thresholds; a short appendix with the YCrCb and HSV thresholds would improve reproducibility.
Circularity Check
Brown-tone override is fit to two same-dataset cases and then cited as evidence that H*L* is more reliable; the fairness-gap comparison itself is not circularly fit.
-
fitted input called prediction
[Section II.B (brown-tone override) and Section IV.A (interpretation of Dark-group recall)]
"To address this, we apply a brown-tone override that reclassifies pixels as Dark based on RGB ranges and Hue, independently of their initial L∗ label. Specifically, a pixel qualifies if: 100 ≤ R ≤ 170, 60 ≤ G ≤ 110, 40 ≤ B ≤ 85, with R > G > B, and (R − G) < 30, (G − B) < 25. The Hue angle must fall within 20◦ ≤ H ∗ ≤ 50◦, capturing typical brown hues. This rule improved classification in underrepresented groups, correctly reassigning two borderline cases (see last row of Figure 4)."
The RGB and hue thresholds of the brown-tone override were explicitly chosen so that two AffectNet borderline cases would be reassigned to the Dark group. The paper then presents this same reassignment as evidence that the H*L* method 'improves tone estimation' and 'more reliably captures chromatic characteristics, particularly brown tones.' The evidence for H*L* superiority on dark skin is therefore the in-sample fit used to define the method: the rule is constructed to classify those cases as Dark, and the same cases are then displayed as proof of improved classification.
full rationale
The paper's central contribution is an empirical comparison of fairness metrics under two skin-tone taxonomies. That comparison is not circular: the F1 and TPR disparities are computed from model predictions and the two taxonomies, and the paper does not fit the taxonomy to produce a desired fairness gap. The clearest circular element is the brown-tone override, whose thresholds were tuned to two borderline cases in the same dataset and then cited as evidence that the H*L* method is more accurate on dark skin. The authors acknowledge this limitation, calling the thresholds 'empirical' and 'dataset-specific,' which makes the practice transparent but does not remove the in-sample nature of that particular claim. No load-bearing self-citation was found; the Mery citations are only offered as future explainability alternatives. The reported inconsistency between Figure 6's F1 disparity values and Table IV's per-group F1 scores, and the abstract's claim about augmentation versus the statement that no rebalancing was applied, are reproducibility and internal-consistency concerns rather than circularity, and I do not count them as circular steps.
Assumptions & free parameters
free parameters (3)
- ITA thresholds =
Light >55 degrees, Medium 30-55 degrees, Dark <30 degrees
- H*L* L* thresholds =
Light >67.0, Medium 37.0-67.0, Dark <37.0
- Brown-tone override thresholds =
R 100-170, G 60-110, B 40-85, R>G>B, (R-G)<30, (G-B)<25, Hue 20-50 degrees
assumptions (4)
- domain assumption The skin pixel segmentation (Otsu plus chrominance filters in YCrCb and HSV) correctly isolates facial skin regions.
- domain assumption Average RGB of segmented skin pixels, converted to Lab, is a reliable proxy for skin tone.
- domain assumption The AffectNet test set is an appropriate real-world distribution for fairness evaluation.
- domain assumption Grad-CAM attention differences by group indicate variation in feature encoding relevant to fairness.
Cite this review
Pith. "Pith review of TrustSkin: A Fairness Pipeline for Trustworthy Facial Affect Analysis Across Skin Tone." pith.science (2026). https://pith.science/paper/UBAUTR4W
@misc{pith2026250520637,
author = {Pith},
title = {Pith review of: TrustSkin: A Fairness Pipeline for Trustworthy Facial Affect Analysis Across Skin Tone},
year = {2026},
howpublished = {\url{https://pith.science/paper/UBAUTR4W}},
note = {Machine review of arXiv:2505.20637}
}
abstract
Understanding how facial affect analysis (FAA) systems perform across different demographic groups requires reliable measurement of sensitive attributes such as ancestry, often approximated by skin tone, which itself is highly influenced by lighting conditions. This study compares two objective skin tone classification methods: the widely used Individual Typology Angle (ITA) and a perceptually grounded alternative based on Lightness ($L^*$) and Hue ($H^*$). Using AffectNet and a MobileNet-based model, we assess fairness across skin tone groups defined by each method. Results reveal a severe underrepresentation of dark skin tones ($\sim 2 \%$), alongside fairness disparities in F1-score (up to 0.08) and TPR (up to 0.11) across groups. While ITA shows limitations due to its sensitivity to lighting, the $H^*$-$L^*$ method yields more consistent subgrouping and enables clearer diagnostics through metrics such as Equal Opportunity. Grad-CAM analysis further highlights differences in model attention patterns by skin tone, suggesting variation in feature encoding. To support future mitigation efforts, we also propose a modular fairness-aware pipeline that integrates perceptual skin tone estimation, model interpretability, and fairness evaluation. These findings emphasize the relevance of skin tone measurement choices in fairness assessment and suggest that ITA-based evaluations may overlook disparities affecting darker-skinned individuals.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Demographic Bias in Biometrics: A Survey on an Emerging Chal- lenge,
P. Drozdowski, C. Rathgeb, A. Dantcheva, N. Damer, and C. Busch, “Demographic Bias in Biometrics: A Survey on an Emerging Chal- lenge,” IEEE Transactions on Technology and Society , vol. 1, no. 2, pp. 89–103, 2020
work page 2020
-
[2]
The Impact of Racial Distribution in Training Data on Face Recognition Bias: A Closer Look,
M. Kolla and A. Savadamuthu, “The Impact of Racial Distribution in Training Data on Face Recognition Bias: A Closer Look,” in IEEE/CVF Winter Conference on Applications of Computer Vision Workshops, WACVW 2023, pp. 313–322, 2023
work page 2023
-
[3]
NIST special publication 1270. Towards a standard for identifying and managing bias in AI,
R. Schwartz, A. Vassilev, K. Greene, L. Perine, A. Burt, P. Hall, and K. Greene, “NIST special publication 1270. Towards a standard for identifying and managing bias in AI,” NIST special, pp. 1–86, 2022
work page 2022
-
[4]
M. Mattioli and F. Cabitza, “Not in My Face: Challenges and Ethical Considerations in Automatic Face Emotion Recognition Technology,” Machine Learning and Knowledge Extraction, vol. 6, no. 4, pp. 2201– 2231, 2024
work page 2024
-
[5]
Investigating Bias and Fairness in Facial Expression Recognition,
T. Xu, J. White, S. Kalkan, and H. Gunes, “Investigating Bias and Fairness in Facial Expression Recognition,” in Computer Vision – ECCV 2020 Workshops (A. Bartoli and A. Fusiello, eds.), (Cham), pp. 506–523, Springer International Publishing, 2020
work page 2020
-
[6]
Balancing the Scales: Enhanc- ing Fairness in Facial Emotion Recognition with Latent Alignment,
S. S. A. Rizvi, A. Seth, and P. Narang, “Balancing the Scales: Enhanc- ing Fairness in Facial Emotion Recognition with Latent Alignment,” in Pattern Recognition (A. Antonacopoulos, S. Chaudhuri, R. Chellappa, C.-L. Liu, S. Bhattacharya, and U. Pal, eds.), (Cham), pp. 113–128, Springer Nature Switzerland, 2025
work page 2025
-
[7]
Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Mod- els,
M. M. Hosseini, A. P. Fard, and M. H. Mahoor, “Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Mod- els,” pp. 1–14, 2025
work page 2025
-
[8]
Notes from the AI frontier : Tackling bias in AI ( and in humans ),
J. Silberg and J. Manyika, “Notes from the AI frontier : Tackling bias in AI ( and in humans ),” Mckinsey Global Institute , pp. 1–8, 2019
work page 2019
Show all 49 references
-
[9]
Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,
J. Buolamwini and T. Gebru, “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” Proceedings of Machine Learning Research , vol. 81, pp. 77–91, 2018
2018
-
[10]
How We’ve Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial Analysis,
M. K. Scheuerman, K. Wade, C. Lustig, and J. R. Brubaker, “How We’ve Taught Algorithms to See Identity: Constructing Race and Gender in Image Databases for Facial Analysis,” Proceedings of the ACM on Human-Computer Interaction , vol. 4, no. CSCW1, pp. 1–35, 2020
2020
-
[11]
One label, one billion faces: Usage and consistency of racial categories in computer vision,
Z. Khan and Y . Fu, “One label, one billion faces: Usage and consistency of racial categories in computer vision,” FAccT 2021 - Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, pp. 587–597, 2021
2021
-
[12]
Consensus and Subjectivity of Skin Tone Annotation for ML Fairness,
C. Schumann, G. O. Olanubi, A. Wright, E. Monk, C. Heldreth, and S. Ricco, “Consensus and Subjectivity of Skin Tone Annotation for ML Fairness,” Advances in Neural Information Processing Systems , vol. 36, no. NeurIPS, 2023
2023
-
[13]
Which Skin Tone Measures Are the Most Inclusive? An Investigation of Skin Tone Measures for Artificial Intelligence,
C. M. Heldreth, E. P. Monk, A. T. Clark, C. Schumann, X. Eyee, and S. Ricco, “Which Skin Tone Measures Are the Most Inclusive? An Investigation of Skin Tone Measures for Artificial Intelligence,” ACM Journal on Responsible Computing , vol. 1, no. 1, pp. 1–21, 2024
2024
-
[14]
The validity and practicality of sun-reactive skin types I through VI.,
T. B. Fitzpatrick, “The validity and practicality of sun-reactive skin types I through VI.,” Archives of dermatology, vol. 124, pp. 869–871, jun 1988
1988
-
[15]
Fitz- patrick Skin Type, Individual Typology Angle, and Melanin Index in an African Population: Steps Toward Universally Applicable Skin Photosensitivity Assessments.,
M. Wilkes, C. Y . Wright, J. L. du Plessis, and A. Reeder, “Fitz- patrick Skin Type, Individual Typology Angle, and Melanin Index in an African Population: Steps Toward Universally Applicable Skin Photosensitivity Assessments.,” JAMA dermatology, vol. 151, pp. 902– 903, aug 2015
2015
-
[16]
Diversity in Faces,
M. Merler, N. Ratha, R. S. Feris, and J. R. Smith, “Diversity in Faces,” pp. 1–29, 2019
2019
-
[17]
FairFace: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation,
K. Karkkainen and J. Joo, “FairFace: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation,” Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision, WACV 2021 , pp. 1547–1557, 2021
2021
-
[18]
Beyond Skin Tone: A Multidi- mensional Measure of Apparent Skin Color,
W. Thong, P. Joniak, and A. Xiang, “Beyond Skin Tone: A Multidi- mensional Measure of Apparent Skin Color,” Proceedings of the IEEE International Conference on Computer Vision , pp. 4880–4890, 2023
2023
-
[19]
Racial Bias within Face Recognition : A Survey,
T. P. Breckon, “Racial Bias within Face Recognition : A Survey,” vol. 1, no. 1, 2023
2023
-
[20]
Evaluating Group Fairness in News Recommendations : A Comparative Study of Algorithms and Metrics,
B. Huebner, T. E. Kolb, and J. Neidhardt, “Evaluating Group Fairness in News Recommendations : A Comparative Study of Algorithms and Metrics,” pp. 337–346
-
[21]
Meta Balanced Network for Fair Face Recognition,
M. Wang, Y . Zhang, and W. Deng, “Meta Balanced Network for Fair Face Recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, pp. 8433–8448, 2021
2021
-
[22]
Towards Measuring Fairness in AI: The Casual Conversations Dataset,
C. Hazirbas, J. Bitton, B. Dolhansky, J. Pan, A. Gordo, and C. C. Ferrer, “Towards Measuring Fairness in AI: The Casual Conversations Dataset,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 4, no. 3, pp. 324–332, 2022
2022
-
[23]
Racial Influence on Automated Perceptions of Emotions,
L. Rhue, “Racial Influence on Automated Perceptions of Emotions,” SSRN Electronic Journal , 2018
2018
-
[24]
Addressing Racial Bias in Facial Emotion Recognition,
A. Fan, X. Xiao, and P. Washington, “Addressing Racial Bias in Facial Emotion Recognition,” 2023
2023
-
[25]
Understanding perception of algorithmic decisions: Fair- ness, trust, and emotion in response to algorithmic management,
M. K. Lee, “Understanding perception of algorithmic decisions: Fair- ness, trust, and emotion in response to algorithmic management,” Big data & society , vol. 5, no. 1, p. 2053951718756684, 2018
2018
-
[26]
Counterfactual fairness for facial expression recognition,
J. Cheong, S. Kalkan, and H. Gunes, “Counterfactual fairness for facial expression recognition,” in European Conference on Computer Vision, pp. 245–261, Springer, 2022
2022
-
[27]
Bias and fairness on multimodal emotion detection algorithms,
M. Schmitz, R. Ahmed, and J. Cao, “Bias and fairness on multimodal emotion detection algorithms,” arXiv preprint arXiv:2205.08383 , 2022
2022 arXiv
-
[28]
Domain adaptation for bias mitigation in affective computing: use cases for facial emotion recognition and sentiment analysis systems,
P. Singhal, S. Gokhale, A. Shah, D. K. Jain, R. Walambe, A. Ekart, and K. Kotecha, “Domain adaptation for bias mitigation in affective computing: use cases for facial emotion recognition and sentiment analysis systems,” Discover Applied Sciences , vol. 7, no. 4, p. 229, 2025
2025
-
[29]
Be- yond Accuracy: Fairness, Scalability, and Uncertainty Considerations in Facial Emotion Recognition,
L. Fromberg, T. Nielsen, F. D. Frumosu, and L. H. Clemmensen, “Be- yond Accuracy: Fairness, Scalability, and Uncertainty Considerations in Facial Emotion Recognition,” in Northern Lights Deep Learning Conference, pp. 67–74, PMLR, 2024
2024
-
[30]
Com- putational methods for pigmented skin lesion classification in images: review and future trends,
R. B. Oliveira, J. P. Papa, A. S. Pereira, and J. M. R. Tavares, “Com- putational methods for pigmented skin lesion classification in images: review and future trends,”Neural Computing and Applications, vol. 29, no. 3, pp. 613–636, 2018
2018
-
[31]
Human Skin Detection Using RGB, HSV and YCbCr Color Models,
S. Kolkur, D. Kalbande, P. Shimpi, C. Bapat, and J. Jatakia, “Human Skin Detection Using RGB, HSV and YCbCr Color Models,” vol. 137, pp. 324–332, 2017
2017
-
[32]
Colorimetric skin tone scale for improved accuracy of human skin tone annotations.,
C. Cook, J. Howard, L. Rabbitt, I. Shuggi, Y . Sirotin, J. Tipton, and A. Vemury, “Colorimetric skin tone scale for improved accuracy of human skin tone annotations.,” ACM J. Responsib. Comput. , May
-
[33]
Fairness and Bias Mitigation in Computer Vision: A Survey,
S. Dehdashtian, R. He, Y . Li, G. Balakrishnan, N. Vasconcelos, V . Ordonez, and V . N. Boddeti, “Fairness and Bias Mitigation in Computer Vision: A Survey,” 2024
2024
-
[34]
AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild,
A. Mollahosseini, B. Hasani, and M. H. Mahoor, “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild,” IEEE Transactions on Affective Computing , vol. 10, no. 1, pp. 18–31, 2019
2019
-
[35]
Exploring strategies to generate Fitz- patrick skin type metadata for dermoscopic images using individual ty- pology angle techniques,
A. Corbin and O. Marques, “Exploring strategies to generate Fitz- patrick skin type metadata for dermoscopic images using individual ty- pology angle techniques,” Multimedia Tools and Applications, vol. 82, no. 15, pp. 23771–23795, 2023
2023
-
[37]
Variations in skin colour and the biological consequences of ultraviolet radiation exposure,
S. Del Bino and F. Bernerd, “Variations in skin colour and the biological consequences of ultraviolet radiation exposure,” British Journal of Dermatology , vol. 169, no. SUPPL. 3, pp. 33–40, 2013
2013
-
[38]
Achieve fairness without demographics for dermatological disease diagnosis,
C. H. Chiu, Y . J. Chen, Y . Wu, Y . Shi, and T. Y . Ho, “Achieve fairness without demographics for dermatological disease diagnosis,” Medical Image Analysis, vol. 95, no. December 2023, p. 103188, 2024
2023
-
[39]
Estimating Skin Tone and Effects on Classification Performance in Dermatology Datasets,
N. M. Kinyanjui, T. Odonga, C. Cintas, N. C. F. Codella, R. Panda, P. Sattigeri, and K. R. Varshney, “Estimating Skin Tone and Effects on Classification Performance in Dermatology Datasets,” pp. 1–11, 2019
2019
-
[40]
Towards Fairness in AI for Melanoma Detection: Systemic Review and Recommenda- tions,
L. N. Montoya, J. S. Roberts, and B. S. Hidalgo, “Towards Fairness in AI for Melanoma Detection: Systemic Review and Recommenda- tions,” 2024
2024
-
[41]
AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias,
R. K. Bellamy, A. Mojsilovic, S. Nagar, K. N. Ramamurthy, J. Richards, D. Saha, P. Sattigeri, M. Singh, K. R. Varshney, Y . Zhang, K. Dey, M. Hind, S. C. Hoffman, S. Houde, K. Kannan, P. Lohia, J. Martino, and S. Mehta, “AI Fairness 360: An extensible toolkit for detecting and...
2019
-
[42]
FACET: Fairness in Computer Vision Evalua- tion Benchmark,
L. Gustafson, C. Rolland, N. Ravi, Q. Duval, A. Adcock, C. Y . Fu, M. Hall, and C. Ross, “FACET: Fairness in Computer Vision Evalua- tion Benchmark,” Proceedings of the IEEE International Conference on Computer Vision , pp. 20313–20325, 2023
2023
-
[43]
Equality of opportunity in super- vised learning,
M. Hardt, E. Price, and N. Srebro, “Equality of opportunity in super- vised learning,” Advances in Neural Information Processing Systems , pp. 3323–3331, 2016
2016
-
[44]
Coroama and A
L. Coroama and A. Groza, Evaluation Metrics in Explainable Artificial Intelligence (XAI), vol. 1675 CCIS. Springer International Publishing, 2022
2022
-
[45]
Explaining Face Recognition Through SHAP-Based Pixel-Level Face Image Quality Assessment,
C. Biagi, L. Rethfeld, A. Kuijper, and P. Terh ¨orst, “Explaining Face Recognition Through SHAP-Based Pixel-Level Face Image Quality Assessment,” in 2023 IEEE International Joint Conference on Bio- metrics (IJCB), pp. 1–10, 2023
2023
-
[46]
On Black-Box Explanation for Face Ver- ification,
D. Mery and B. Morris, “On Black-Box Explanation for Face Ver- ification,” 2022 IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2022 , pp. 1194–1203, 2022
2022
-
[47]
True Black-Box Explanation in Facial Analysis,
D. Mery, “True Black-Box Explanation in Facial Analysis,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1595–1604, 2022
2022
-
[48]
Grad-cam: Why did you say that? visual explanations from deep networks via gradient-based localization,
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: Why did you say that? visual explanations from deep networks via gradient-based localization,” Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization , vol. 17, pp....
2016
-
[49]
On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation,
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. M ¨uller, and W. Samek, “On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation,”PLOS ONE, vol. 10, no. 7, pp. 1–46, 2015
2015
-
[50]
Focused LRP: Explainable AI for Face Morphing Attack Detection,
C. Seibold, A. Hilsmann, and P. Eisert, “Focused LRP: Explainable AI for Face Morphing Attack Detection,” Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision Workshops, WACVW 2021, pp. 88–96, 2021
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.