Pith. sign in

REVIEW 5 cited by

Understanding Unequal Gender Classification Accuracy from Face Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.00099 v1 pith:3UAH4R7K submitted 2018-11-30 cs.CV cs.CYstat.ML

classification cs.CVcs.CYstat.ML
keywords facegenderclassificationskintypeaccuracyacrossdifferences
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Accuracy on dark-skinned females is significantly worse than on any other group. In this paper, we conduct several analyses to try to uncover the reason for this gap. The main finding, perhaps surprisingly, is that skin type is not the driver. This conclusion is reached via stability experiments that vary an image's skin type via color-theoretic methods, namely luminance mode-shift and optimal transport. A second suspect, hair length, is also shown not to be the driver via experiments on face images cropped to exclude the hair. Finally, using contrastive post-hoc explanation techniques for neural networks, we bring forth evidence suggesting that differences in lip, eye and cheek structure across ethnicity lead to the differences. Further, lip and eye makeup are seen as strong predictors for a female face, which is a troubling propagation of a gender stereotype.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 45 citations worldwide. Full citation record

  1. Toward Calibrated, Fair, and accurate Deepfake Detection

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    Face-Feature Tuning is a label-free logit remapping method that reduces FPR/TPR gaps across groups in deepfake detection while preserving overall accuracy.

  2. When Algorithms Play Favorites: Lookism in the Generation and Perception of Faces

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Text-to-image models link facial attractiveness to unrelated positive traits, and gender classifiers misclassify faces generated with negative trait labels more often, with the largest effects for non-White women.

  3. Facial Analysis Systems and Down Syndrome

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Commercial facial analysis tools are less accurate on faces of people with Down syndrome, especially for gender classification of males and age classification of adults.

  4. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  5. The Folly of AI for Age Verification

    cs.CY 2025-05 conditional novelty 3.0 of 10

    Deploying AI for age verification is likely to be ineffective and inequitable, the paper argues by analogy to facial recognition and remote proctoring systems.

Pith tools