Pith. sign in

REVIEW 1 cited by

Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00252 v1 pith:NN5I62AA submitted 2024-08-30 cs.LG

classification cs.LG
keywords datasetcreatorsresponsiblecreationdatasetscrucialcurrentlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing demand for high-quality datasets in machine learning has raised concerns about the ethical and responsible creation of these datasets. Dataset creators play a crucial role in developing responsible practices, yet their perspectives and expertise have not yet been highlighted in the current literature. In this paper, we bridge this gap by presenting insights from a qualitative study that included interviewing 18 leading dataset creators about the current state of the field. We shed light on the challenges and considerations faced by dataset creators, and our findings underscore the potential for deeper collaboration, knowledge sharing, and collective development. Through a close analysis of their perspectives, we share seven central recommendations for improving responsible dataset creation, including issues such as data quality, documentation, privacy and consent, and how to mitigate potential harms from unintended use cases. By fostering critical reflection and sharing the experiences of dataset creators, we aim to promote responsible dataset creation practices and develop a nuanced understanding of this crucial but often undervalued aspect of machine learning research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new dataset annotates conspiracy texts with six cognitive traits, and experiments show LLMs reproduce conspiracy reasoning more readily than they deflect it.

Pith tools