A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.
The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence
3 Pith papers cite this work, alongside 28 external citations. Polarity classification is still indexing.
abstract
As the production of and reliance on datasets to produce automated decision-making systems (ADS) increases, so does the need for processes for evaluating and interrogating the underlying data. After launching the Dataset Nutrition Label in 2018, the Data Nutrition Project has made significant updates to the design and purpose of the Label, and is launching an updated Label in late 2020, which is previewed in this paper. The new Label includes context-specific Use Cases &Alerts presented through an updated design and user interface targeted towards the data scientist profile. This paper discusses the harm and bias from underlying training data that the Label is intended to mitigate, the current state of the work including new datasets being labeled, new and existing challenges, and further directions of the work, as well as Figures previewing the new label.
citation-role summary
citation-polarity summary
fields
cs.CY 3roles
background 2representative citing papers
In interviews with 11 Portuguese-language model developers, four AI ethics tools guided general ethical reflection but failed to surface Portuguese-specific harms like cultural misrepresentation and low language performance.
Structured dataset documentation shows little engagement with major reflexivity themes from FAccT literature, leading to a new codebook and extended datasheet questions.
citing papers explorer
-
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.
-
Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Study
In interviews with 11 Portuguese-language model developers, four AI ethics tools guided general ethical reflection but failed to surface Portuguese-specific harms like cultural misrepresentation and low language performance.