Pith. sign in

REVIEW 2 cited by

Data Representativity for Machine Learning and AI Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.04706 v2 pith:2MJ4IXMM submitted 2022-03-09 stat.ML cs.AIcs.LG

Data Representativity for Machine Learning and AI Systems

classification stat.ML cs.AIcs.LG
keywords datarepresentativityrepresentativesamplesystemsconceptsevaluatefocus
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Data representativity is crucial when drawing inference from data through machine learning models. Scholars have increased focus on unraveling the bias and fairness in models, also in relation to inherent biases in the input data. However, limited work exists on the representativity of samples (datasets) for appropriate inference in AI systems. This paper reviews definitions and notions of a representative sample and surveys their use in scientific AI literature. We introduce three measurable concepts to help focus the notions and evaluate different data samples. Furthermore, we demonstrate that the contrast between a representative sample in the sense of coverage of the input space, versus a representative sample mimicking the distribution of the target population is of particular relevance when building AI systems. Through empirical demonstrations on US Census data, we evaluate the opposing inherent qualities of these concepts. Finally, we propose a framework of questions for creating and documenting data with data representativity in mind, as an addition to existing dataset documentation templates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Prototypical Signature Approach for Writer-Independent Offline Signature Verification

    cs.CV 2026-06 unverdicted novelty 6.0

    Prototypical signatures enable generation of diverse negative samples for writer-independent offline signature verification, improving skilled forgery detection and allowing scalable linear SVM alternatives to RBF models.

  2. Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

    cs.CV 2025-10 conditional novelty 6.0

    Source-dataset selection for medical transfer learning is driven by community practice and perceived similarity, and 'more similar is better' does not consistently hold.