Pith. sign in

REVIEW 3 cited by

Data Representativity for Machine Learning and AI Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.04706 v2 pith:2MJ4IXMM submitted 2022-03-09 stat.ML cs.AIcs.LG

classification stat.MLcs.AIcs.LG
keywords datarepresentativityrepresentativesamplesystemsconceptsevaluatefocus
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data representativity is crucial when drawing inference from data through machine learning models. Scholars have increased focus on unraveling the bias and fairness in models, also in relation to inherent biases in the input data. However, limited work exists on the representativity of samples (datasets) for appropriate inference in AI systems. This paper reviews definitions and notions of a representative sample and surveys their use in scientific AI literature. We introduce three measurable concepts to help focus the notions and evaluate different data samples. Furthermore, we demonstrate that the contrast between a representative sample in the sense of coverage of the input space, versus a representative sample mimicking the distribution of the target population is of particular relevance when building AI systems. Through empirical demonstrations on US Census data, we evaluate the opposing inherent qualities of these concepts. Finally, we propose a framework of questions for creating and documenting data with data representativity in mind, as an addition to existing dataset documentation templates.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Source-dataset selection for medical transfer learning is driven by community practice and perceived similarity, and 'more similar is better' does not consistently hold.

  2. Robustness of transferability estimation metrics for medical imaging

    eess.IV 2026-08 conditional novelty 5.0 of 10

    Transferability estimation metric rankings in medical imaging are unstable to target resampling and to the evaluation metric used for the reference ranking.

  3. Predictive Representativity: Uncovering Racial Bias in AI-based Skin Cancer Detection

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A new fairness audit metric, Predictive Representativity, reveals that five common skin-cancer classifiers trained on HAM10000 have much lower precision for darker Fitzpatrick skin types when tested on an independent ...

Pith tools