Pith. sign in

REVIEW 1 cited by

Massively Annotated Datasets for Assessment of Synthetic and Real Data in Face Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.15234 v1 pith:T4BBHJJ3 submitted 2024-04-23 cs.CV

classification cs.CV
keywords datasetsrealsyntheticmodelsattributefacerecognitionannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Face recognition applications have grown in parallel with the size of datasets, complexity of deep learning models and computational power. However, while deep learning models evolve to become more capable and computational power keeps increasing, the datasets available are being retracted and removed from public access. Privacy and ethical concerns are relevant topics within these domains. Through generative artificial intelligence, researchers have put efforts into the development of completely synthetic datasets that can be used to train face recognition systems. Nonetheless, the recent advances have not been sufficient to achieve performance comparable to the state-of-the-art models trained on real data. To study the drift between the performance of models trained on real and synthetic datasets, we leverage a massive attribute classifier (MAC) to create annotations for four datasets: two real and two synthetic. From these annotations, we conduct studies on the distribution of each attribute within all four datasets. Additionally, we further inspect the differences between real and synthetic datasets on the attribute set. When comparing through the Kullback-Leibler divergence we have found differences between real and synthetic samples. Interestingly enough, we have verified that while real samples suffice to explain the synthetic distribution, the opposite could not be further from being true.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Autonomous AI Surveillance: Multimodal Deep Learning for Cognitive and Behavioral Monitoring

    cs.CV 2025-07 reject novelty 2.0 of 10

    A classroom monitoring system combining YOLOv8, MTCNN, and LResNet reports high detection accuracies, but the paper lacks reproducible artifacts and contains inconsistent results.

Pith tools