Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

Review of Demographic Fairness in Face Recognition

T0 review · 4 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Demographic fairness in face recognition is not one problem with one fix: causes interact, no single dataset or metric captures it, and many group disparities may be driven by correlated traits such as hairstyle and makeup rather than…

desk verdict A current, well-organized review of fairness in face recognition that is more useful for its taxonomy and emphasis on soft-biometric confounds than for any new result; deserves review despite overclaiming its systematicity. read the letter →

arxiv 2502.02309 v3 pith:RAQNWKD2 submitted 2025-02-04 cs.CV cs.CR

classification cs.CVcs.CR
keywords demographicfairnessfacerecognitionbiascausessoft-biometricattributesmetricsmitigationdatasetsintersectionality
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Demographic fairness in face recognition, this review argues, is not a single defect with a single remedy. The paper systematically organizes the literature into four interacting dimensions—causes, datasets, assessment metrics, and mitigation strategies—and draws two overarching conclusions. First, observed accuracy differences across race, gender, and age typically arise from multiple overlapping factors, so attributing a disparity to the demographic attribute alone is often unsafe. Second, recent studies show that many apparent demographic gaps, especially by gender, can shrink or disappear when groups are matched on non-demographic attributes such as hairstyle, makeup, and facial hair, which means what looks like demographic bias may partly be an artifact of correlated social and cultural appearance norms. A sympathetic reader would take the paper as a structured map of the field and a caution against simplistic bias attributions and one-number fairness scores.

What carries the argument

The machinery of the review is a four-part taxonomy—causes, datasets, metrics, and mitigations—organized around the distinction between differential performance (differences in genuine and impostor score distributions, independent of thresholds) and differential outcome (threshold-dependent differences in FMR and FNMR). Within this structure, the load-bearing concept is the soft-biometric attribute: a non-demographic, often culturally entangled trait such as hairstyle, facial hair, makeup, or occlusion that can shift score distributions and mimic demographic bias. On the measurement side, the review catalogs threshold-based indices such as Inequity Ratio, Fairness Discrepancy Rate, GARBE, MAPE, and SEDG, alongside threshold-agnostic measures such as d-prime, Kolmogorov–Smirnov distance, and the Separation, Compactness, and Distribution Fairness Indices, arguing that threshold choice (global, yoked, or majority-group) materially changes fairness conclusions.

What would settle it

Take a face-recognition model and a test set whose demographic cells are exactly matched on hairstyle, makeup, facial hair, brightness, pose, and resolution (by resampling or image synthesis), then measure per-group FMR and FNMR at a fixed threshold; if substantial group differences persist under such attribute-matched conditions, the paper's caution that soft attributes may explain observed disparities is weakened for that model, whereas vanishing differences would support it.

Watch

Extended reading notes

Core claim

This review's central claim is that demographic fairness in face recognition is best understood as a multifaceted problem whose causes interact. It consolidates evidence that training-data imbalance, skin-tone and skin-reflectance effects, image quality and acquisition conditions, algorithmic choices, and soft-biometric attributes such as hairstyle, makeup, facial hair, and occlusion all contribute to performance differences, and that no single dataset, metric, or mitigation technique captures or resolves the issue. The authors specifically highlight work showing that gender-related accuracy gaps can vanish when men and women share the same facial attributes, suggesting that many reported disparities may be driven by demographically correlated non-demographic factors rather than by the demographic attribute itself. They therefore urge caution before concluding that a face-recognition system is biased toward a particular group, and call for multi-attribute annotations and controlled isolation of factors to make causal claims reliable.

Load-bearing premise

The synthesis assumes that the demographic labels and experimental protocols in the surveyed studies are reliable and comparable; the paper itself acknowledges that labels are noisy, datasets are skewed, and thresholds and yoking choices differ, so if those measurement conditions are systematically wrong, many synthesized conclusions about which groups are disadvantaged could be artifacts.

Editorial extensions

If this is right

  • Fairness evaluations that report only global FMR/FNMR gaps risk misattributing causes; the review implies that evaluations should pair threshold-based metrics with threshold-agnostic distribution measures.
  • Datasets annotated with both demographic and non-demographic attributes become a prerequisite for isolating why disparities occur, since most current public datasets lack the multi-attribute labels needed for controlled comparisons.
  • Mitigation methods that target one demographic attribute may shift disparities onto another, so intersectional evaluation (e.g., race by gender by age) should accompany any debiasing claim.
  • Because fairness gains often come from added model capacity or architecture changes rather than a true resolution of the fairness–accuracy trade-off, comparisons across mitigation methods are only meaningful when complexity and training data are accounted for.
  • Deployment niches such as lightweight models, low-resolution surveillance imagery, lossy compression, and remote identity verification need dedicated fairness evaluation, since compression, quantization, and resolution loss can amplify demographic disparities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if soft attributes substantially drive observed gaps, then 'demographic bias' in many operational systems is better described as a socially mediated appearance effect, and interventions at the acquisition stage (lighting, standardization, attribute normalization) could be more effective than demographic-label-based retraining.
  • Editorial inference: the review's caution suggests that third-party fairness audits should include attribute-matched probe sets and statistical uncertainty intervals before stating that a system is biased against a group; a single error-rate table is not enough.
  • Editorial inference: a natural testable extension is a standardized 'matched-attribute audit protocol' in which every demographic cell is balanced on a fixed list of soft attributes; adopting such a protocol across benchmarks would make cross-study fairness comparisons meaningful for the first time.
  • Editorial inference: synthetic data with precisely controlled attribute distributions could operationalize the 'partial derivative' experiment the paper calls for, letting researchers vary one attribute at a time; the current realism gap in synthetic faces limits this, but the direction is directly implied by the review.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. This paper is a survey of demographic fairness in face recognition (FR), organized into four main areas: causes of performance differences (Section III), datasets for fairness research (Section IV), fairness assessment metrics (Section V), and bias mitigation methods (Section VI), followed by future directions (Section VII) and a conclusion (Section VIII). The central thesis is that demographic fairness is multifaceted: causes interact, no single dataset or metric captures the problem, and many observed disparities may be driven by correlated non-demographic (soft-biometric) attributes rather than by the demographic attribute itself. The paper catalogs a large number of recent works and provides summary tables for causes, datasets, metrics, and mitigation approaches.

Significance. If the paper's synthesis is correct, it offers a structured map of the field and a valuable caution against attributing observed FR performance gaps directly to demographic factors. Its strengths include a broad and current bibliography (including 2024–2025 works), useful taxonomy tables (Tables I–IV), explicit acknowledgment of label noise and threshold sensitivity, and a nuanced discussion of soft-biometric confounds such as hairstyle, makeup, and facial hair. The conclusion's warning that 'attributing lower FR performance to demographic bias may be misleading' is an important, falsifiable stance. However, the review's central claims depend on comparing findings across studies with heterogeneous demographic labels, thresholds, and evaluation protocols; this dependency is acknowledged but not stress-tested, which weakens the evidential basis for several synthesized conclusions.

major comments (4)
  1. [Sec. III (A, D, F) and Sec. V] The survey synthesizes directional claims about which demographic groups are disadvantaged from studies with different protocols, thresholds, and demographic label definitions. For example, Section III-A cites [26] for 'African-American cohorts exhibit higher FMRs, while Caucasian cohorts face higher FNMR,' Section III-D cites [41] for lighter skin tones 'consistently outperforming medium-dark tones,' and Section III-F cites [65] for the gender gap vanishing with matched attributes. The paper itself notes in Section V that 'threshold setting significantly affects all threshold-based fairness metrics' and in Section VII that demographic labels are noisy and discretized, yet it does not perform a sensitivity audit or protocol compatibility analysis across the cited studies. This is load-bearing because the survey's structural claims about which disparities are real and which are confounded could change if labels or yoking conventions differ systematically. I recommend adding a table that records, for each cited empirical study, the dataset, demographic label source (self-report, classifier, manual), threshold/yoking procedure, and whether within-group or cross-group impostor pairs were used, and then qualifying synthesized claims accordingly or restricting them to studies with commensurable protocols.
  2. [Abstract and Sec. I] The paper claims to 'systematically examine' the literature and to provide a 'comprehensive' review, but no literature selection protocol is described: there is no search strategy, inclusion/exclusion criteria, or quality assessment. This matters because the central claim that causes are multifaceted and interacting rests on the representativeness of the included works. Without a documented protocol, the possibility of selection bias cannot be ruled out, especially given that the authors' own metrics and mitigation methods are prominently featured. I recommend either adding a methodology paragraph describing how sources were identified and screened, or softening the 'comprehensive/systematic' claims to 'broad narrative review.'
  3. [Sec. V, Eq. (8); Sec. VI, [117], [107]] The authors' own proposed fairness measures and mitigation methods are presented without critical comparison to alternatives or discussion of their limitations. For instance, Eq. (8) defines DFI using the KL divergence to a reference distribution P_ref, but the choice of P_ref is application-dependent, KL divergence is asymmetric, and the index can be sensitive to distribution support; none of these limitations is discussed. Similarly, the Demographic Fairness Transformer (DeFT) [117] and the regularized score calibration approach [107] are described with favorable results, but no failure modes, computational costs, or comparisons to label-free baselines are provided. Since the review aims to be comprehensive, a balanced treatment that situates these methods among alternatives would increase confidence in the survey's objectivity.
  4. [Sec. V and Sec. VIII] The review asserts that no single metric captures demographic fairness and that 'future fairness evaluations should include both threshold-based and threshold-agnostic metrics,' but it does not empirically demonstrate the inadequacy of any single metric. The discussion in Section V describes individual metrics and their qualitative pros and cons, yet no comparison is made on a common dataset or operating point. Because this claim is central to the paper's message, I recommend either adding a small illustrative comparison of a few representative metrics (e.g., IR, FDR, GARBE, SFI/CFI/DFI, d-prime) on one publicly available dataset, or explicitly framing the claim as an open research question rather than a demonstrated conclusion.
minor comments (7)
  1. [Sec. V, Eqs. (2) and (3)] The symbol A(τ) is reused for the NIST geometric-mean ratio in Eq. (2) and for the maximum absolute FMR difference in Eq. (3); B(τ) is similarly overloaded. Please rename one set of definitions to avoid confusion.
  2. [Sec. V, Eq. (4)] There is a typo: the text says 'The formula for GABRE' but the metric is consistently called GARBE everywhere else, including Table III.
  3. [Sec. V, Eq. (8)] The formatting of the DFI expression is ambiguous: 'DFI = 1 − 1/N log2 N sum DKL' should use explicit parentheses, e.g., DFI = 1 − (1/(N log2 N)) Σ_d D_KL(P^(d) || P_ref).
  4. [Sec. V, Eq. (7)] The d-prime equation contains a stray 'q' symbol and the fraction under the square root is ambiguous: d′ = |μ_m − μ_nm| / sqrt( (σ_m^2 + σ_nm^2)/2 ) would be clearer.
  5. [Sec. IV, Table II and Fig. 3] Figure 3 compiles demographic distributions from 'original sources (wherever available) or from other works,' but no error bars or sample-size information are provided; please state in the caption that some distributions are approximate and refer readers to the original sources for exact numbers.
  6. [Sec. VI, Table IV] The method-type labels 'Data-Processing' and 'In-Processing' are inconsistent with the text, which uses 'Pre-Processing,' 'In-Processing,' and 'Post-Processing.' Please harmonize the terminology.
  7. [Sec. IV, MORPH entry] The MORPH dataset is cited via a cleaning report by Bingham et al. rather than the original MORPH source (Ricanek & Tesafaye). Please cite the original dataset publication alongside any curation report.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the review synthesizes independent studies and does not use the authors' own metrics or methods as load-bearing premises.

full rationale

I inspected the manuscript for self-definitional links, fitted-input predictions, and load-bearing self-citations. The paper is a literature review whose central claims—that demographic fairness is multifaceted, that no single dataset or metric captures it, and that some observed disparities may be driven by correlated non-demographic attributes—are supported by numerous independent external works (e.g., [26], [37], [65], [60], [64], [19]) rather than by the authors' own definitions or fitted quantities. The authors' own contributions appear only as surveyed content: Section V describes the SFI/CFI/DFI measures as 'introduced' in [98], and Section VI describes DeFT and regularized score calibration as methods proposed in [117] and [107], but the review does not use these to derive its conclusions. No equation in the paper reduces a predicted quantity to a fitted parameter, and no definition is stated in terms of the phenomenon it is meant to explain. The acknowledged issues of noisy demographic labels, heterogeneous thresholding protocols, and dataset skew are presented in Section VII as open challenges, not as hidden assumptions that make the synthesis circular. Self-citations occur, but none is load-bearing: the review's synthesized picture would stand without them. I therefore find no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new parameters or entities; it is a review. The main unstated premises are that the surveyed metrics and datasets meaningfully capture fairness, and that demographic labels are reliable enough for comparison.

assumptions (3)
  • domain assumption Demographic groups being compared are equivalent except for the demographic attribute of interest
    Explicitly stated in Section I as 'a commonly implicit but rarely emphasized assumption'; this equivalence is required for performance differences to be attributed to demographics, yet most datasets are not controlled for covariates.
  • domain assumption Demographic labels in the surveyed datasets are sufficiently accurate for group-wise comparisons
    Section VII (Noisy Labels) acknowledges that annotations come from classifiers or manual labeling and contain errors; the review still uses those labels to synthesize conclusions about bias.
  • domain assumption Fairness in face recognition is adequately captured by error rates and score distribution metrics
    The review restricts fairness to performance differentials (FMR, FNMR, and score distributions), excluding other fairness notions such as representational harm; this scope choice is stated in the Introduction and Section V.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Review of Demographic Fairness in Face Recognition." pith.science (2026). https://pith.science/paper/RAQNWKD2

@misc{pith2026250202309,
  author       = {Pith},
  title        = {Pith review of: Review of Demographic Fairness in Face Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RAQNWKD2}},
  note         = {Machine review of arXiv:2502.02309}
}
read the original abstract

Demographic fairness in face recognition (FR) has emerged as a critical area of research, given its impact on fairness, equity, and reliability across diverse applications. As FR technologies are increasingly deployed globally, disparities in performance across demographic groups -- such as race, ethnicity, and gender -- have garnered significant attention. These biases not only compromise the credibility of FR systems but also raise ethical concerns, especially when these technologies are employed in sensitive domains. This review consolidates extensive research efforts providing a comprehensive overview of the multifaceted aspects of demographic fairness in FR. We systematically examine the primary causes, datasets, assessment metrics, and mitigation approaches associated with demographic disparities in FR. By categorizing key contributions in these areas, this work provides a structured approach to understanding and addressing the complexity of this issue. Finally, we highlight current advancements and identify emerging challenges that need further investigation. This article aims to provide researchers with a unified perspective on the state-of-the-art while emphasizing the critical need for equitable and trustworthy FR systems.

Figures

Figures reproduced from arXiv: 2502.02309 by the authors.

Figure 1
Figure 1. Illustration of differences in performance across demo [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of false matches and false non-matches [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Distribution of images of commonly used FR datasets considering (a) [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Illustration of categories of bias mitigation methods in FR. [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigation of Accuracy and Bias in Face Recognition Trained with Synthetic Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Balanced synthetic face data from Stable Diffusion v3.5 reduces racial bias in face recognition models but does not yet match real-data accuracy on hard benchmarks.

Reference graph

Works this paper leans on

154 extracted references · 73 canonical work pages · cited by 1 Pith paper

  1. [26]

    Issues related to face recognition accuracy varying based on race and skin tone,

    K. Krishnapriya, V . Albiero, K. Vangara, M. C. King, and K. W. Bowyer, “Issues related to face recognition accuracy varying based on race and skin tone,” IEEE Transactions on Technology and Society, vol. 1, no. 1, pp. 8–20, 2020. 4, 6

  2. [41]

    An experimental evaluation of covariates effects on unconstrained face verification,

    B. Lu, J.-C. Chen, C. D. Castillo, and R. Chellappa, “An experimental evaluation of covariates effects on unconstrained face verification,” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 1, no. 1, pp. 42–55, 2019. 5, 6, 7, 14, 16, 18

  3. [65]

    On the “illusion

    P. J. Kurz, H. Wu, K. W. Bowyer, and P. Terh ¨orst, “On the “illusion” of gender bias in face recognition: Explaining the fairness issue through non-demographic attributes,” arXiv preprint arXiv:2501.12020 , 2025. 7, 8

  4. [117]

    Demographic fairness transformer for bias mitigation in face recognition,

    K. Kotwal and S. Marcel, “Demographic fairness transformer for bias mitigation in face recognition,” in 2024 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2024, pp. 1–10. 16, 18, 22

  5. [107]

    Mitigating demographic bias in face recognition via regularized score calibration,

    K. Kotwal and S. Marcel, “Mitigating demographic bias in face recognition via regularized score calibration,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, January 2024, pp. 1150–1159. 14, 17, 18

  6. [1]

    Demographic fairness in biometric systems: What do the experts say?

    C. Rathgeb, P. Drozdowski, D. Frings, N. Damer, and C. Busch, “Demographic fairness in biometric systems: What do the experts say?” IEEE Technology and Society Magazine , vol. 41, no. 4, pp. 71–82,

  7. [2]

    Fairface challenge at eccv 2020: Analyzing bias in face recognition,

    T. Sixta, J. C. Jacques Junior, P. Buch-Cardona, E. Vazquez, and S. Escalera, “Fairface challenge at eccv 2020: Analyzing bias in face recognition,” in Computer Vision–ECCV 2020 Workshops: Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16 . Springer, 2020, pp. 463–481. 1

  8. [3]

    Demographic bias in biometrics: A survey on an emerging challenge,

    P. Drozdowski, C. Rathgeb, A. Dantcheva, N. Damer, and C. Busch, “Demographic bias in biometrics: A survey on an emerging challenge,” IEEE Transactions on Technology and Society , vol. 1, no. 2, pp. 89– 103, 2020. 1, 2

Show all 154 references
  1. [4]

    Biometrics: Trust, but verify,

    A. Jain, D. Deb, and J. Engelsma, “Biometrics: Trust, but verify,” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 4, no. 3, pp. 303–323, 2021. 1, 2

  2. [5]

    Wrongfully accused by an algorithm,

    K. Hill, “Wrongfully accused by an algorithm,” The New York Times, 2020. [Online]. Available: https://www.nytimes.com/2020/06/ 24/technology/facial-recognition-arrest.html 1

  3. [6]

    Amazon’s face recognition falsely matched 28 members of congress with mugshots,

    J. Snow, “Amazon’s face recognition falsely matched 28 members of congress with mugshots,” ACLU Blog, 2018. [Online]. Available: https: //www.aclu.org/blog/privacy-technology/surveillance-technologies/ amazons-face-recognition-falsely-matched-28 1

  4. [7]

    Why new facial-recognition air- port screenings are raising concerns,

    L. Marshall, “Why new facial-recognition air- port screenings are raising concerns,” 2023. [On- line]. Available: https://www.colorado.edu/today/2023/07/11/ why-new-facial-recognition-airport-screenings-are-raising-concerns 1

  5. [8]

    The struggle to control facial recognition at airports,

    E. Falk, “The struggle to control facial recognition at airports,”

  6. [9]

    A survey on bias and fairness in machine learning,

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM computing surveys (CSUR), vol. 54, no. 6, pp. 1–35, 2021. 1, 2, 15

  7. [10]

    The effect of broad and specific demographic homogeneity on the imposter distributions and false match rates in face recognition algorithm performance,

    J. J. Howard, Y . B. Sirotin, and A. R. Vemury, “The effect of broad and specific demographic homogeneity on the imposter distributions and false match rates in face recognition algorithm performance,” in 2019 IEEE 10th International Conference on Biometrics Theory, Applicatio...

  8. [11]

    Bias in facial recognition technologies used by law enforcement: Understanding the causes and searching for a way out,

    A. Limant ˙e, “Bias in facial recognition technologies used by law enforcement: Understanding the causes and searching for a way out,” Nordic Journal of Human Rights , vol. 42, no. 2, pp. 115–134, 2024. 1

  9. [12]

    Law enforcement use of facial recognition: bias, disparate impacts on people of color, and the need for federal legislation,

    C. Jones, “Law enforcement use of facial recognition: bias, disparate impacts on people of color, and the need for federal legislation,” NCJL & Tech., vol. 22, p. 777, 2020. 1

  10. [13]

    Understanding bias in facial recognition technologies,

    D. Leslie, “Understanding bias in facial recognition technologies,” arXiv preprint arXiv:2010.07023 , 2020. 1

  11. [14]

    Watching the watchers: bias and vulnerability in remote proctoring software,

    B. Burgess, A. Ginsberg, E. W. Felten, and S. Cohney, “Watching the watchers: bias and vulnerability in remote proctoring software,” in 31st USENIX security symposium (USENIX security 22), 2022, pp. 571–588. 1

  12. [15]

    Saving face: Investigating the ethical concerns of facial recognition auditing,

    I. D. Raji, T. Gebru, M. Mitchell, J. Buolamwini, J. Lee, and E. Denton, “Saving face: Investigating the ethical concerns of facial recognition auditing,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2020, pp. 145–151. 1 IEEE TRANSACTIONS ON BIOMETRICS...

  13. [16]

    Bias in multimodal ai: Testbed for fair automatic recruitment,

    A. Pena, I. Serna, A. Morales, and J. Fierrez, “Bias in multimodal ai: Testbed for fair automatic recruitment,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 28–29. 1

  14. [17]

    Some research problems in biometrics: The future beckons,

    A. Ross, S. Banerjee, C. Chen, A. Chowdhury, V . Mirjalili, R. Sharma, T. Swearingen, and S. Yadav, “Some research problems in biometrics: The future beckons,” in 2019 International Conference on Biometrics (ICB). IEEE, 2019, pp. 1–8. 2

  15. [18]

    Challenges for automated face recognition systems,

    C. Busch, “Challenges for automated face recognition systems,” Nature Reviews Electrical Engineering , pp. 1–10, 2024. 2

  16. [19]

    Face recognition vendor test part 3: Demographic effects,

    P. Grother, M. Ngan, and K. Hanaoka, “Face recognition vendor test part 3: Demographic effects,” 12 2019. 2, 4, 6, 7, 8, 12, 15

  17. [20]

    Demographic differentials in face recognition algorithms,

    P. Grother, “Demographic differentials in face recognition algorithms,” EAB Virtual Event Series-Demographic Fairness in Biometric Systems,

  18. [21]

    Tbiom special issue on trustworthy biometrics-editorial,

    W. Deng, T. Hassner, X. Liu, and M. Pantic, “Tbiom special issue on trustworthy biometrics-editorial,” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 4, no. 3, pp. 301–302, 2022. 2

  19. [22]

    The hitchhiker’s guide to bias and fairness in facial affective signal processing: Overview and techniques,

    J. Cheong, S. Kalkan, and H. Gunes, “The hitchhiker’s guide to bias and fairness in facial affective signal processing: Overview and techniques,” IEEE Signal Processing Magazine , vol. 38, no. 6, pp. 39–49, 2021. 2

  20. [23]

    ISO/IEC DIS 19795-10. Information Technology – Biometric Performance Testing and Reporting – Part 10: Quantifying Biometric System Performance Variation Across Demo- graphic Group,

    ISO/IEC JTC1 SC37 Biometrics, “ISO/IEC DIS 19795-10. Information Technology – Biometric Performance Testing and Reporting – Part 10: Quantifying Biometric System Performance Variation Across Demo- graphic Group,” 2023. 2

  21. [24]

    A review on fairness in machine learning,

    D. Pessach and E. Shmueli, “A review on fairness in machine learning,” ACM Computing Surveys (CSUR) , vol. 55, no. 3, pp. 1–44, 2022. 2, 15

  22. [25]

    Racial bias within face recognition: A survey,

    S. Yucer, F. Tektas, N. Al Moubayed, and T. Breckon, “Racial bias within face recognition: A survey,” ACM Computing Surveys, 2023. 2

  23. [27]

    Face recognition performance: Role of demographic information,

    B. F. Klare, M. J. Burge, J. C. Klontz, R. W. V . Bruegge, and A. K. Jain, “Face recognition performance: Role of demographic information,” IEEE Transactions on Information Forensics and Security, vol. 7, no. 6, pp. 1789–1801, 2012. 4, 5, 6, 7, 16, 18

  24. [28]

    Ac- curacy comparison across face recognition algorithms: Where are we on measuring race bias?

    J. G. Cavazos, P. J. Phillips, C. D. Castillo, and A. J. O’Toole, “Ac- curacy comparison across face recognition algorithms: Where are we on measuring race bias?” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 3, no. 1, pp. 101–111, 2020. 4, 6, 7

  25. [29]

    Rethinking com- mon assumptions to mitigate racial bias in face recognition datasets,

    M. Gwilliam, S. Hegde, L. Tinubu, and A. Hanson, “Rethinking com- mon assumptions to mitigate racial bias in face recognition datasets,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 4123–4132. 4, 6

  26. [30]

    Does face recog- nition accuracy get better with age? deep face matchers say no,

    V . Albiero, K. Bowyer, K. Vangara, and M. King, “Does face recog- nition accuracy get better with age? deep face matchers say no,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020, pp. 261–269. 4

  27. [31]

    How does gender balance in training data affect face recognition accuracy?

    V . Albiero, K. Zhang, and K. W. Bowyer, “How does gender balance in training data affect face recognition accuracy?” in 2020 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2020, pp. 1–10. 4

  28. [32]

    What should be balanced in a “balanced

    H. Wu and K. W. Bowyer, “What should be balanced in a “balanced” face recognition dataset?” arXiv preprint arXiv:2304.09818 , 2023. 4, 6

  29. [33]

    Soft biometrics,

    R. Donida Labati, A. Ross, and A. Dantcheva, “Soft biometrics,” in Encyclopedia of Cryptography, Security and Privacy . Springer, 2022, pp. 1–4. 4

  30. [34]

    Racial faces in the wild: Reducing racial bias by information maximization adaptation network,

    M. Wang, W. Deng, J. Hu, X. Tao, and Y . Huang, “Racial faces in the wild: Reducing racial bias by information maximization adaptation network,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 692–702. 4, 6, 10

  31. [35]

    The impact of racial distribution in training data on face recognition bias: A closer look,

    M. Kolla and A. Savadamuthu, “The impact of racial distribution in training data on face recognition bias: A closer look,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 313–322. 4, 6

  32. [36]

    Understanding unequal gender classification accuracy from face images,

    V . Muthukumar, T. Pedapati, N. Ratha, P. Sattigeri, C.-W. Wu, B. Kingsbury, A. Kumar, S. Thomas, A. Mojsilovic, and K. R. Varshney, “Understanding unequal gender classification accuracy from face images,” arXiv preprint arXiv:1812.00099 , 2018. 4, 6

  33. [37]

    Demographic effects in facial recognition and their dependence on image acquisition: An evaluation of eleven commercial systems,

    C. M. Cook, J. J. Howard, Y . B. Sirotin, J. L. Tipton, and A. R. Vemury, “Demographic effects in facial recognition and their dependence on image acquisition: An evaluation of eleven commercial systems,” IEEE Transactions on Biometrics, Behavior, and Identity Science , vol. 1...

  34. [38]

    The validity and practicality of sun-reactive skin types i through vi,

    T. B. Fitzpatrick, “The validity and practicality of sun-reactive skin types i through vi,” Archives of dermatology, vol. 124, no. 6, pp. 869– 871, 1988. 4

  35. [39]

    The monk skin tone scale,

    E. Monk, “The monk skin tone scale,” OSF, 2023. 4

  36. [40]

    Gender shades: Intersectional accuracy disparities in commercial gender classification,

    J. Buolamwini and T. Gebru, “Gender shades: Intersectional accuracy disparities in commercial gender classification,” in Proceedings of the Conference on Fairness, Accountability and Transparency , ser. Proceedings of Machine Learning Research, S. Friedler and C. Wilson, Eds.,...

  37. [42]

    An other-race effect for face recognition algorithms,

    P. J. Phillips, F. Jiang, A. Narvekar, J. Ayyad, and A. J. O’Toole, “An other-race effect for face recognition algorithms,” ACM Transactions on Applied Perception (TAP) , vol. 8, no. 2, pp. 1–11, 2011. 5, 7

  38. [43]

    A review of face recognition against longitudinal child faces,

    K. Ricanek, S. Bhardwaj, and M. Sodomsky, “A review of face recognition against longitudinal child faces,” BIOSIG 2015 , pp. 15– 26, 2015. 5, 7, 8

  39. [44]

    Deep learning for face recognition: Pride or prejudiced?

    S. Nagpal, M. Singh, R. Singh, and M. Vatsa, “Deep learning for face recognition: Pride or prejudiced?” arXiv preprint arXiv:1904.01219 ,

  40. [45]

    Analysis of gender inequality in face recognition accuracy,

    V . Albiero, K. Krishnapriya, K. Vangara, K. Zhang, M. C. King, and K. W. Bowyer, “Analysis of gender inequality in face recognition accuracy,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision workshops , 2020, pp. 81–89. 5, 6, 7

  41. [46]

    Rethinking bias mitigation: Fairer architectures make for fairer face recognition,

    S. Dooley, R. Sukthanker, J. Dickerson, C. White, F. Hutter, and M. Goldblum, “Rethinking bias mitigation: Fairer architectures make for fairer face recognition,” Advances in Neural Information Processing Systems, vol. 36, pp. 74 366–74 393, 2023. 5, 7, 17, 18

  42. [47]

    Characterizing the variability in face recognition accuracy relative to race,

    K. Krishnapriya, K. Vangara, M. C. King, V . Albiero, K. Bowyer et al., “Characterizing the variability in face recognition accuracy relative to race,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019, pp. 0–0. 6, 7

  43. [48]

    Demo- graphic effects in facial recognition and their dependence on image acquisition: An evaluation of eleven commercial systems,

    C. Cook, J. Howard, Y . Sirotin, J. Tipton, and A. Vemury, “Demo- graphic effects in facial recognition and their dependence on image acquisition: An evaluation of eleven commercial systems,”IEEE Trans- actions on Biometrics, Behavior, and Identity Science, vol. 1, no. 1, pp. ...

  44. [49]

    Demographic effects across 158 facial recognition systems,

    C. M. Cook, J. J. Howard, Y . B. Sirotin, J. L. Tipton, and A. R. Vemury, “Demographic effects across 158 facial recognition systems,” Technical report, DHS, Tech. Rep., 2023. 6, 7, 8

  45. [50]

    Face recognition accuracy across demographics: Shining a light into the problem,

    H. Wu, V . Albiero, K. Krishnapriya, M. King, and K. Bowyer, “Face recognition accuracy across demographics: Shining a light into the problem,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 1041–1050. 6, 7, 14

  46. [51]

    Impact of blur and resolution on demographic disparities in 1-to-many facial identification,

    A. Bhatta, G. Pangelinan, M. C. King, and K. W. Bowyer, “Impact of blur and resolution on demographic disparities in 1-to-many facial identification,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 412–420. 6, 7, 14, 21

  47. [52]

    Lights, camera, matching: The role of image illumination in fair face recogni- tion,

    G. Pangelinan, G. Bezold, H. Wu, M. King, and K. Bowyer, “Lights, camera, matching: The role of image illumination in fair face recogni- tion,” arXiv preprint arXiv:2501.08910 , 2025. 6, 7, 22

  48. [53]

    Face verification subject to varying (age, ethnicity, and gender) demographics using deep learning,

    H. El Khiyari and H. Wechsler, “Face verification subject to varying (age, ethnicity, and gender) demographics using deep learning,”Journal of Biometrics and Biostatistics , vol. 7, no. 323, p. 11, 2016. 7, 8

  49. [54]

    Longitudinal study of automatic face recognition,

    L. Best-Rowden and A. K. Jain, “Longitudinal study of automatic face recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 1, pp. 148–162, 2017. 7, 8

  50. [55]

    Facegenderid: Exploiting gender information in dcnns face recognition systems,

    R. Vera-Rodriguez, M. Blazquez, A. Morales, E. Gonzalez-Sosa, J. C. Neves, and H. Proenc ¸a, “Facegenderid: Exploiting gender information in dcnns face recognition systems,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , 2019, ...

  51. [56]

    Does face recognition accuracy get better with age? deep face matchers say no,

    V . Albiero, K. Bowyer, K. Vangara, and M. C. King, “Does face recognition accuracy get better with age? deep face matchers say no,” in Proceedings of the IEEE Winter Conference on Applications of Computer Vision, vol. 1, 2020, pp. 250–258. 7, 8

  52. [57]

    Explaining bias in deep face recognition via image characteristics,

    A. Atzori, G. Fenu, and M. Marras, “Explaining bias in deep face recognition via image characteristics,” in 2022 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2022, pp. 1–10. 7, 8

  53. [58]

    Towards fair face verification: An in-depth analysis of demographic biases,

    I. Sarridis, C. Koutlis, S. Papadopoulos, and C. Diou, “Towards fair face verification: An in-depth analysis of demographic biases,” in Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 2023, pp. 194–208. 7, 8 IEEE TRANSACTIONS ON BI...

  54. [59]

    Face regions impact recognition accuracy differently across demographics,

    V . Albiero, K. W. Bowyer, and M. C. King, “Face regions impact recognition accuracy differently across demographics,” in 2022 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2022, pp. 1–9. 7, 9

  55. [60]

    The gender gap in face recognition accuracy is a hairy problem,

    A. Bhatta, V . Albiero, K. W. Bowyer, and M. C. King, “The gender gap in face recognition accuracy is a hairy problem,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 303–312. 7, 8, 14

  56. [61]

    Exploring causes of demographic variations in face recognition accuracy,

    G. Pangelinan, K. Krishnapriya, V . Albiero, G. Bezold, K. Zhang, K. Vangara, M. C. King, and K. W. Bowyer, “Exploring causes of demographic variations in face recognition accuracy,” in Computer Vision. Chapman and Hall/CRC, 2024, pp. 61–81. 7, 8

  57. [62]

    Can the accuracy bias by facial hairstyle be reduced through balancing the training data?

    K. Ozturk, H. Wu, and K. W. Bowyer, “Can the accuracy bias by facial hairstyle be reduced through balancing the training data?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 1519–1528. 7, 8

  58. [63]

    Facial hair area in face recognition across demographics: Small size, big effect,

    H. Wu, S. Tian, A. Bhatta, K. ¨Ozt¨urk, K. Ricanek, and K. W. Bowyer, “Facial hair area in face recognition across demographics: Small size, big effect,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 1131–1140. 7, 8

  59. [64]

    Fairness under cover: Evaluating the impact of occlusions on demographic bias in facial recognition,

    R. M. Mamede, P. C. Neto, and A. F. Sequeira, “Fairness under cover: Evaluating the impact of occlusions on demographic bias in facial recognition,” arXiv preprint arXiv:2408.10175 , 2024. 7, 9

  60. [66]

    Analyzing the impact of demographic and operational variables on 1- to-many face id search,

    G. Pangelinan, A. Bhatta, H. Wu, M. C. King, and K. W. Bowyer, “Analyzing the impact of demographic and operational variables on 1- to-many face id search,” IEEE Transactions on Technology and Society,

  61. [67]

    Can facial cosmetics affect the matching accuracy of face recognition systems?

    A. Dantcheva, C. Chen, and A. Ross, “Can facial cosmetics affect the matching accuracy of face recognition systems?” in 2012 IEEE Fifth International Conference on Biometrics: Theory, Applications and Systems (BTAS). IEEE, 2012, pp. 391–398. 8

  62. [68]

    Detection of age-induced makeup attacks on face recognition systems using multi-layer deep features,

    K. Kotwal, Z. Mostaani, and S. Marcel, “Detection of age-induced makeup attacks on face recognition systems using multi-layer deep features,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 2, no. 1, pp. 15–25, 2019. 8

  63. [69]

    Learning to learn across diverse data biases in deep face recognition,

    C. Liu, X. Yu, Y .-H. Tsai, M. Faraki, R. Moslemi, M. Chandraker, and Y . Fu, “Learning to learn across diverse data biases in deep face recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2022, pp. 4072–4082. 8, 22

  64. [70]

    A survey of face recognition techniques under occlusion,

    D. Zeng, R. Veldhuis, and L. Spreeuwers, “A survey of face recognition techniques under occlusion,” IET Biometrics, vol. 10, no. 6, pp. 581– 606, 2021. 8

  65. [71]

    Mfr 2021: Masked face recognition competition,

    F. Boutros, N. Damer, J. N. Kolf, K. Raja, F. Kirchbuchner, R. Ra- machandra, A. Kuijper, P. Fang, C. Zhang, F. Wang et al., “Mfr 2021: Masked face recognition competition,” in 2021 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2021, pp. 1–10. 8

  66. [72]

    Latent enhancing autoen- coder for occluded image classification,

    K. Kotwal, T. Deshmukh, and P. Gopal, “Latent enhancing autoen- coder for occluded image classification,” in 2024 IEEE International Conference on Image Processing (ICIP) . IEEE, 2024, pp. 894–900. 8

  67. [73]

    Com- positional convolutional neural networks: A robust and interpretable model for object recognition under occlusion,

    A. Kortylewski, Q. Liu, A. Wang, Y . Sun, and A. Yuille, “Com- positional convolutional neural networks: A robust and interpretable model for object recognition under occlusion,” International Journal of Computer Vision , vol. 129, pp. 736–760, 2021. 8

  68. [74]

    Gendered dif- ferences in face recognition accuracy explained by hairstyles, makeup, and facial morphology,

    V . Albiero, K. Zhang, M. C. King, and K. W. Bowyer, “Gendered dif- ferences in face recognition accuracy explained by hairstyles, makeup, and facial morphology,” IEEE Transactions on Information Forensics and Security, vol. 17, pp. 127–137, 2021. 8, 9, 10

  69. [75]

    MORPH-II: In- consistencies and Cleaning,

    G. Bingham, B. Yip, M. Ferguson, and C. Nansalo, “MORPH-II: In- consistencies and Cleaning,” University of North Carolina Wilmington NSF REU, 2017. 9, 10

  70. [76]

    Nist special database 32-multiple encounter dataset ii (meds-ii),

    A. P. Founds, N. Orlans, W. Genevieve, and C. I. Watson, “Nist special database 32-multiple encounter dataset ii (meds-ii),” 2011. 9, 10

  71. [77]

    Age and gender estimation of unfiltered faces,

    E. Eidinger, R. Enbar, and T. Hassner, “Age and gender estimation of unfiltered faces,” IEEE Transactions on information forensics and security, vol. 9, no. 12, pp. 2170–2179, 2014. 9, 10

  72. [78]

    A method for curation of web- scraped face image datasets,

    V . A. Kai Zhang and K. W. Bowyer, “A method for curation of web- scraped face image datasets,” in International Workshop on Biometrics and Forensics (IWBF), 2020. 9, 10

  73. [79]

    An asian face dataset and how race influences face recognition,

    Z. Xiong, Z. Wang, C. Du, R. Zhu, J. Xiao, and T. Lu, “An asian face dataset and how race influences face recognition,” in Pacific Rim Conference on Multimedia . Springer, 2018, pp. 372–383. 9, 10

  74. [80]

    Vggface2: A dataset for recognising faces across pose and age,

    Q. Cao, L. Shen, W. Xie, O. Parkhi, and A. Zisserman, “Vggface2: A dataset for recognising faces across pose and age,” in Proceedings of the IEEE International Conference on Automatic Face & Gesture Recognition. IEEE, 2018, pp. 67–74. 9, 10, 11

  75. [81]

    Demogpairs: Quantifying the impact of demographic imbalance in deep face recognition,

    I. Hupont and C. Fern ´andez, “Demogpairs: Quantifying the impact of demographic imbalance in deep face recognition,” in 2019 14th IEEE international conference on automatic face & gesture recognition (FG 2019). IEEE, 2019, pp. 1–7. 10

  76. [82]

    Investigating bias in deep face analysis: The kanface dataset and empirical study,

    M. Georgopoulos, Y . Panagakis, and M. Pantic, “Investigating bias in deep face analysis: The kanface dataset and empirical study,” Image and vision computing , vol. 102, p. 103954, 2020. 10, 17, 18

  77. [83]

    Mitigating bias in face recognition us- ing skewness-aware reinforcement learning,

    M. Wang and W. Deng, “Mitigating bias in face recognition us- ing skewness-aware reinforcement learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9322–9331. 10, 14, 16, 18

  78. [84]

    Sen- sitivenets: Learning agnostic representations with application to face images,

    A. Morales, J. Fierrez, R. Vera-Rodriguez, and R. Tolosana, “Sen- sitivenets: Learning agnostic representations with application to face images,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, vol. 43, no. 6, pp. 2158–2164, 2020. 10, 11, 18, 19

  79. [85]

    Balancing biases and preserving privacy on balanced faces in the wild,

    J. P. Robinson, C. Qin, Y . Henon, S. Timoner, and Y . Fu, “Balancing biases and preserving privacy on balanced faces in the wild,” IEEE Transactions on Image Processing , 2023. 10, 11, 19

  80. [86]

    Face recognition: too bias, or not too bias?

    J. P. Robinson, G. Livitz, Y . Henon, C. Qin, Y . Fu, and S. Timoner, “Face recognition: too bias, or not too bias?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 0–1. 10, 11, 14, 19

  81. [87]

    Casia-face- africa: A large-scale african face image database,

    J. Muhammad, Y . Wang, C. Wang, K. Zhang, and Z. Sun, “Casia-face- africa: A large-scale african face image database,” IEEE Transactions on Information Forensics and Security , vol. 16, pp. 3634–3646, 2021. 10, 11

  82. [88]

    Ms-celeb-1m: A dataset and benchmark for large-scale face recognition,

    Y . Guo, L. Zhang, Y . Hu, X. He, and J. Gao, “Ms-celeb-1m: A dataset and benchmark for large-scale face recognition,” in Proceedings of 14th European Conference on Computer Vision (ECCV) . Springer, 2016, pp. 87–102. 10

  83. [89]

    The megaface benchmark: 1 million faces for recognition at scale,

    I. Kemelmacher-Shlizerman, S. M. Seitz, D. Miller, and E. Brossard, “The megaface benchmark: 1 million faces for recognition at scale,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4873–4882. 11

  84. [90]

    Learning race from face: A survey,

    S. Fu, H. He, and Z.-G. Hou, “Learning race from face: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 36, no. 12, pp. 2483–2509, 2014. 11

  85. [91]

    Age and gender classification using convo- lutional neural networks,

    G. Levi and T. Hassner, “Age and gender classification using convo- lutional neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2015, pp. 34–42. 11

  86. [92]

    Demographic effects on estimates of automatic face recognition performance,

    A. J. O’Toole, P. J. Phillips, X. An, and J. Dunlop, “Demographic effects on estimates of automatic face recognition performance,” Image and Vision Computing , vol. 30, no. 3, pp. 169–176, 2012. 12

  87. [93]

    Face recognition vendor test (frvt) part 8: Summarizing demographic differentials,

    P. Grother, “Face recognition vendor test (frvt) part 8: Summarizing demographic differentials,” National Institute of Standards and Tech- nology (NIST), vol. 8429, 2022. 12

  88. [94]

    Assessing uncertainty in similarity scoring: Performance & fairness in face recognition,

    J.-R. Conti and S. Cl ´emenc ¸on, “Assessing uncertainty in similarity scoring: Performance & fairness in face recognition,” in The Twelfth International Conference on Learning Representations , 2024. 12, 13

  89. [95]

    Mitigating gender bias in face recognition using the von mises-fisher mixture model,

    J.-R. Conti, N. Noiry, S. Clemencon, V . Despiegel, and S. Gentric, “Mitigating gender bias in face recognition using the von mises-fisher mixture model,” in International Conference on Machine Learning . PMLR, 2022, pp. 4344–4369. 12, 15, 19

  90. [96]

    Fairness in biometrics: a figure of merit to assess biometric verification systems,

    T. de Freitas Pereira and S. Marcel, “Fairness in biometrics: a figure of merit to assess biometric verification systems,” IEEE Transactions on Biometrics, Behavior, and Identity Science, vol. 4, no. 1, pp. 19–29,

  91. [97]

    Statistical methods for assessing differences in false non-match rates across demographic groups,

    M. Schuckers, S. Purnapatra, K. Fatima, D. Hou, and S. Schuckers, “Statistical methods for assessing differences in false non-match rates across demographic groups,” in International Conference on Pattern Recognition. Springer, 2022, pp. 570–581. 12, 13

  92. [98]

    Fairness Index Measures to Evaluate Bias in Biometric Recognition,

    K. Kotwal and S. Marcel, “Fairness Index Measures to Evaluate Bias in Biometric Recognition,” inProceedings of the International Conference on Pattern Recognition Workshops. Springer, 2022, pp. 479–493. 13, 14

  93. [99]

    Evaluating proposed fairness models for face recogni- tion algorithms,

    J. J. Howard, E. J. Laird, R. E. Rubin, Y . B. Sirotin, J. L. Tipton, and A. R. Vemury, “Evaluating proposed fairness models for face recogni- tion algorithms,” in International Conference on Pattern Recognition Workshops. Springer, 2022, pp. 431–447. 13

  94. [100]

    Fair face verification by using non-sensitive soft-biometric attributes,

    E. Villalobos, D. Mery, and K. Bowyer, “Fair face verification by using non-sensitive soft-biometric attributes,” IEEE Access , vol. 10, pp. 30 168–30 179, 2022. 13, 14 IEEE TRANSACTIONS ON BIOMETRICS, BEHA VIOR AND IDENTITY SCIENCE 26

  95. [101]

    Sum of group error differences: A critical exami- nation of bias evaluation in biometric verification and a dual-metric measure,

    A. Elobaid, N. Ramoly, L. Younes, S. Papadopoulos, E. Ntoutsi, and I. Kompatsiaris, “Sum of group error differences: A critical exami- nation of bias evaluation in biometric verification and a dual-metric measure,” in 2024 IEEE 18th International Conference on Automatic Face a...

  96. [102]

    Comprehensive equity index (cei): Defi- nition and application to bias evaluation in biometrics,

    I. Solano, A. Pe ˜na, A. Morales, J. Fierrez, R. Tolosana, F. Zamora- Martinez, and J. S. Agustin, “Comprehensive equity index (cei): Defi- nition and application to bias evaluation in biometrics,” in International Conference on Pattern Recognition. Springer, 2024, pp. 110–126. 13, 14

  97. [103]

    Faircal: Fairness calibration for face verification,

    T. Salvador, S. Cairns, V . V oleti, N. Marshall, and A. Oberman, “Faircal: Fairness calibration for face verification,” in International Conference on Learning Representations , 2022. 14, 18, 19

  98. [104]

    A comprehensive study on face recognition biases beyond demographics,

    P. Terh ¨orst, J. N. Kolf, M. Huber, F. Kirchbuchner, N. Damer, A. M. Moreno, J. Fierrez, and A. Kuijper, “A comprehensive study on face recognition biases beyond demographics,” IEEE Transactions on Technology and Society, vol. 3, no. 1, pp. 16–30, 2021. 14

  99. [105]

    Jointly de-biasing face recognition and demographic attribute estimation,

    S. Gong, X. Liu, and A. Jain, “Jointly de-biasing face recognition and demographic attribute estimation,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16. Springer, 2020, pp. 330–347. 14, 16, 18

  100. [106]

    Comparison-level mitigation of ethnic bias in face recognition,

    P. Terh ¨orst, M. L. Tran, N. Damer, F. Kirchbuchner, and A. Kuijper, “Comparison-level mitigation of ethnic bias in face recognition,” in Proceedings of the International Workshop on Biometrics and Foren- sics. IEEE, 2020, pp. 1–6. 14, 19

  101. [108]

    FRCSyn-onGoing: Benchmarking and comprehensive evaluation of real and synthetic data to improve face recognition systems,

    P. Melzi, R. Tolosana, R. Vera-Rodriguez, M. Kim, C. Rathgeb, X. Liu, I. DeAndres-Tame, A. Morales, J. Fierrez, J. Ortega-Garcia et al. , “FRCSyn-onGoing: Benchmarking and comprehensive evaluation of real and synthetic data to improve face recognition systems,” Informa- tion F...

  102. [109]

    FRCSyn challenge at W ACV 2024: Face recognition challenge in the era of synthetic data,

    ——, “FRCSyn challenge at W ACV 2024: Face recognition challenge in the era of synthetic data,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 892–901. 14, 21

  103. [110]

    Frcsyn challenge at cvpr 2024: Face recognition challenge in the era of synthetic data,

    I. DeAndres-Tame, R. Tolosana, P. Melzi, R. Vera-Rodriguez, M. Kim, C. Rathgeb, X. Liu, A. Morales, J. Fierrez, J. Ortega-Garcia et al. , “Frcsyn challenge at cvpr 2024: Face recognition challenge in the era of synthetic data,” in Proceedings of the IEEE/CVF Conference on Comp...

  104. [111]

    Anatomizing bias in facial analysis,

    R. Singh, P. Majumdar, S. Mittal, and M. Vatsa, “Anatomizing bias in facial analysis,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, 2022, pp. 12 351–12 358. 15

  105. [112]

    Bias mitigation for machine learning classifiers: A comprehensive survey,

    M. Hort, Z. Chen, J. M. Zhang, M. Harman, and F. Sarro, “Bias mitigation for machine learning classifiers: A comprehensive survey,” ACM Journal on Responsible Computing, vol. 1, no. 2, pp. 1–52, 2024. 15

  106. [113]

    Longitudinal study of child face recognition,

    D. Deb, N. Nain, and A. K. Jain, “Longitudinal study of child face recognition,” in 2018 International Conference on Biometrics (ICB) . IEEE, 2018, pp. 225–232. 16, 18

  107. [114]

    Facenet: A unified embedding for face recognition and clustering,

    F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 815–823. 16

  108. [115]

    Analyzing and reducing the damage of dataset bias to face recognition with synthetic data,

    A. Kortylewski, B. Egger, A. Schneider, T. Gerig, A. Morel-Forster, and T. Vetter, “Analyzing and reducing the damage of dataset bias to face recognition with synthetic data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , 2019...

  109. [116]

    Exploring racial bias within face recognition via per-subject adversarially-enabled data augmentation,

    S. Yucer, S. Akc ¸ay, N. Al-Moubayed, and T. P. Breckon, “Exploring racial bias within face recognition via per-subject adversarially-enabled data augmentation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 18–19. 16, 18

  110. [118]

    Uncovering and mitigating algorithmic bias through learned latent structure,

    A. Amini, A. P. Soleimany, W. Schwarting, S. N. Bhatia, and D. Rus, “Uncovering and mitigating algorithmic bias through learned latent structure,” in Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society , 2019, pp. 289–295. 16, 18

  111. [119]

    Deep class- skewed learning for face recognition,

    P. Wang, F. Su, Z. Zhao, Y . Guo, Y . Zhao, and B. Zhuang, “Deep class- skewed learning for face recognition,” Neurocomputing, vol. 363, pp. 35–45, 2019. 16, 18

  112. [120]

    Feature transfer learning for face recognition with under-represented data,

    X. Yin, X. Yu, K. Sohn, X. Liu, and M. Chandraker, “Feature transfer learning for face recognition with under-represented data,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 5704–5713. 16, 18

  113. [121]

    Mitigating face recognition bias via group adaptive classifier,

    S. Gong, X. Liu, and A. Jain, “Mitigating face recognition bias via group adaptive classifier,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3414–3424. 16, 18

  114. [122]

    Gradient attention balance network: Mitigating face recognition racial bias via gradient attention,

    L. Huang, M. Wang, J. Liang, W. Deng, H. Shi, D. Wen, Y . Zhang, and J. Zhao, “Gradient attention balance network: Mitigating face recognition racial bias via gradient attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. ...

  115. [123]

    Invariant feature regularization for fair face recognition,

    J. Ma, Z. Yue, K. Tomoyuki, S. Tomoki, K. Jayashree, S. Pranata, and H. Zhang, “Invariant feature regularization for fair face recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2023, pp. 20 861–20 870. 16, 18

  116. [124]

    Toward fairness in face matching algorithms,

    J. Alasadi, A. Al Hilli, and V . K. Singh, “Toward fairness in face matching algorithms,” in Proceedings of the 1st International Workshop on Fairness, Accountability, and Transparency in MultiMedia , 2019, pp. 19–25. 17, 18

  117. [125]

    Additive adversarial learning for unbiased authentication,

    J. Liang, Y . Cao, C. Zhang, S. Chang, K. Bai, and Z. Xu, “Additive adversarial learning for unbiased authentication,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 11 428–11 437. 17, 18

  118. [126]

    Learning fair face representation with progressive cross transformer,

    Y . Li, Y . Sun, Z. Cui, S. Shan, and J. Yang, “Learning fair face representation with progressive cross transformer,” arXiv preprint arXiv:2108.04983, 2021. 17, 18

  119. [127]

    Consistent instance false positive improves fairness in face recognition,

    X. Xu, Y . Huang, P. Shen, S. Li, J. Li, F. Huang, Y . Li, and Z. Cui, “Consistent instance false positive improves fairness in face recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 578–586. 17, 18

  120. [128]

    Meta balanced network for fair face recognition,

    M. Wang, Y . Zhang, and W. Deng, “Meta balanced network for fair face recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 8433–8448, 2021. 17, 18

  121. [129]

    Fair con- trastive learning for facial attribute classification,

    S. Park, J. Lee, P. Lee, S. Hwang, D. Kim, and H. Byun, “Fair con- trastive learning for facial attribute classification,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 10 389–10 398. 17, 18

  122. [130]

    Sensitive loss: Improving accuracy and fairness of face representations with discrimination-aware deep learning,

    I. Serna, A. Morales, J. Fierrez, and N. Obradovich, “Sensitive loss: Improving accuracy and fairness of face representations with discrimination-aware deep learning,” Artificial Intelligence , vol. 305, p. 103682, 2022. 17, 18

  123. [131]

    Fairness- aware contrastive learning with partially annotated sensitive attributes,

    F. Zhang, K. Kuang, L. Chen, Y . Liu, C. Wu, and J. Xiao, “Fairness- aware contrastive learning with partially annotated sensitive attributes,” in Proceedings of the International Conference on Learning Represen- tations, 2023. 17, 18

  124. [132]

    Mixfairface: Towards ultimate fairness via mixfair adapter in face recognition,

    F.-E. Wang, C.-Y . Wang, M. Sun, and S.-H. Lai, “Mixfairface: Towards ultimate fairness via mixfair adapter in face recognition,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 12, 2023, pp. 14 531–14 538. 17, 18

  125. [133]

    Instance-consistent fair face recognition,

    Y . Li, Y . Sun, Z. Cui, P. Shen, and S. Shan, “Instance-consistent fair face recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 17, 18

  126. [134]

    The impact of age and threshold variation on facial recognition algorithm performance using images of children,

    D. Michalski, S. Y . Yiu, and C. Malec, “The impact of age and threshold variation on facial recognition algorithm performance using images of children,” in 2018 International Conference on Biometrics (ICB). IEEE, 2018, pp. 217–224. 19

  127. [135]

    Face recognition algorithm bias: Performance differences on images of children and adults,

    N. Srinivas, K. Ricanek, D. Michalski, D. S. Bolme, and M. King, “Face recognition algorithm bias: Performance differences on images of children and adults,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops , 2019, pp. 0–0. 19

  128. [136]

    Post-comparison mitigation of demographic bias in face recognition using fair score normalization,

    P. Terh ¨orst, J. N. Kolf, N. Damer, F. Kirchbuchner, and A. Kuijper, “Post-comparison mitigation of demographic bias in face recognition using fair score normalization,” Pattern Recognition Letters, vol. 140, pp. 332–338, 2020. 19

  129. [137]

    Pass: protected attribute suppression system for mitigating bias in face recognition,

    P. Dhar, J. Gleason, A. Roy, C. D. Castillo, and R. Chellappa, “Pass: protected attribute suppression system for mitigating bias in face recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 15 087–15 096. 19

  130. [138]

    Oneface: one threshold for all,

    J. Liu, Z. Yu, H. Qin, Y . Wu, D. Liang, G. Zhao, and K. Xu, “Oneface: one threshold for all,” in European Conference on Computer Vision . Springer, 2022, pp. 545–561. 19, 20 IEEE TRANSACTIONS ON BIOMETRICS, BEHA VIOR AND IDENTITY SCIENCE 27

  131. [139]

    Score normalization for demographic fairness in face recognition,

    Y . Linghu, T. de Freitas Pereira, C. Ecabert, S. Marcel, and M. G¨unther, “Score normalization for demographic fairness in face recognition,” in 2024 IEEE International Joint Conference on Biometrics (IJCB), 2024, pp. 1–11. 19, 20

  132. [140]

    Mitigating bias in facial recognition systems: Centroid fairness loss optimization,

    J.-R. Conti and S. Cl ´emenc ¸on, “Mitigating bias in facial recognition systems: Centroid fairness loss optimization,” in International Confer- ence on Pattern Recognition . Springer, 2024, pp. 371–385. 19

  133. [141]

    Towards gender-neutral face descriptors for mitigating bias in face recognition,

    P. Dhar, J. Gleason, H. Souri, C. D. Castillo, and R. Chellappa, “Towards gender-neutral face descriptors for mitigating bias in face recognition,” arXiv preprint arXiv:2006.07845 , 2020. 19

  134. [142]

    Rectifying the data bias in knowledge distillation,

    B. Liu, S. Zhang, G. Song, H. You, and Y . Liu, “Rectifying the data bias in knowledge distillation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 1477–1486. 20

  135. [143]

    Prune responsibly,

    M. Paganini, “Prune responsibly,” arXiv preprint arXiv:2009.09936 ,

  136. [144]

    Bias in pruned vision mod- els: In-depth analysis and countermeasures,

    E. Iofinova, A. Peste, and D. Alistarh, “Bias in pruned vision mod- els: In-depth analysis and countermeasures,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 24 364–24 373. 20

  137. [145]

    Fairgrape: Fairness-aware gradient pruning method for face attribute classification,

    X. Lin, S. Kim, and J. Joo, “Fairgrape: Fairness-aware gradient pruning method for face attribute classification,” in European Conference on Computer Vision. Springer, 2022, pp. 414–432. 20

  138. [146]

    Mst- kd: Multiple specialized teachers knowledge distillation for fair face recognition,

    E. Caldeira, J. S. Cardoso, A. F. Sequeira, and P. C. Neto, “Mst- kd: Multiple specialized teachers knowledge distillation for fair face recognition,” arXiv preprint arXiv:2408.16563 , 2024. 20

  139. [147]

    Quantization and training of neural networks for efficient integer-arithmetic-only inference,

    B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. ...

  140. [148]

    Does lossy image compression affect racial bias within face recognition?

    S. Yucer, M. Poyser, N. Al Moubayed, and T. P. Breckon, “Does lossy image compression affect racial bias within face recognition?” in 2022 IEEE International Joint Conference on Biometrics (IJCB) . IEEE, 2022, pp. 1–10. 21

  141. [149]

    Gone with the bits: Benchmarking bias in facial phenotype degradation under low-rate neural compression,

    T. Qiu, A. Nichani, R. Tadayon, and H. Jeong, “Gone with the bits: Benchmarking bias in facial phenotype degradation under low-rate neural compression,” in ICML 2024 Next Generation of AI Safety Workshop, 2024. 21

  142. [150]

    Demographic bias in low- resolution deep face recognition in the wild,

    A. Atzori, G. Fenu, and M. Marras, “Demographic bias in low- resolution deep face recognition in the wild,” IEEE Journal of Selected Topics in Signal Processing , vol. 17, no. 3, pp. 599–611, 2023. 21

  143. [151]

    Digi2real: Bridging the realism gap in syn- thetic data face recognition via foundation models,

    A. George and S. Marcel, “Digi2real: Bridging the realism gap in syn- thetic data face recognition via foundation models,” in Proceedings of the Winter Conference on Applications of Computer Vision Workshops, 2025, pp. 1469–1478. 22

  144. [152]

    Bias and diversity in synthetic-based face recognition,

    M. Huber, A. T. Luu, F. Boutros, A. Kuijper, and N. Damer, “Bias and diversity in synthetic-based face recognition,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 6215–6226. 22

  145. [153]

    A large-scale study of performance and equity of commercial remote identity verification technologies across demographics,

    K. Fatima, M. Schuckers, G. Cruz-Ortiz, D. Hou, S. Purnapatra, T. An- drews, A. Neupane, B. Marshall, and S. Schuckers, “A large-scale study of performance and equity of commercial remote identity verification technologies across demographics,” in 2024 IEEE International Joint...

  146. [2024]

    Available: https://journals.library.columbia.edu/index

    [Online]. Available: https://journals.library.columbia.edu/index. php/stlr/blog/view/607 1

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.