Pith. sign in

REVIEW 5 cited by

FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.04913 v1 pith:PNJIXIGN submitted 2019-08-14 cs.CV cs.LG

classification cs.CVcs.LG
keywords racedatasetdatasetsfacegroupsgendernovelaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing public face datasets are strongly biased toward Caucasian faces, and other races (e.g., Latino) are significantly underrepresented. This can lead to inconsistent model accuracy, limit the applicability of face analytic systems to non-White race groups, and adversely affect research findings based on such skewed data. To mitigate the race bias in these datasets, we construct a novel face image dataset, containing 108,501 images, with an emphasis of balanced race composition in the dataset. We define 7 race groups: White, Black, Indian, East Asian, Southeast Asian, Middle East, and Latino. Images were collected from the YFCC-100M Flickr dataset and labeled with race, gender, and age groups. Evaluations were performed on existing face attribute datasets as well as novel image datasets to measure generalization performance. We find that the model trained from our dataset is substantially more accurate on novel datasets and the accuracy is consistent between race and gender groups.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 143 citations worldwide. Full citation record

  1. On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FairCLIP's claimed fairness and performance gains over CLIP do not reproduce on two datasets, and its official implementation diverges from the paper's own formulation.

  2. Analyzing Character Representation in Media Content using Multimodal Foundation Model: Effectiveness and Trust

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A 30-participant user study shows that viewers partially understand AI-generated gender and age representation charts for films, find them moderately useful, and trust the gender model more than the age model.

  3. A Responsible Face Recognition Approach for Small and Mid-Scale Systems Through Personalized Neural Networks

    cs.CV 2025-05 conditional novelty 6.0 of 10

    MOTE replaces fixed face embeddings with per-identity binary classifiers trained using KDE-generated synthetic samples, improving gender fairness and privacy at the cost of storage and enrollment time.

  4. Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Using wrong/correct hard-sample head comparisons, LTC finds spurious CLIP attention heads and corrects them to raise worst-group accuracy on biased benchmarks.

  5. Vision-Language Models display a strong gender bias

    cs.CV 2025-08 reject novelty 3.0 of 10

    Using cosine similarity in CLIP embedding space, the paper finds that male and female face sets are differentially associated with occupation and activity statements across all four tested models.

Pith tools