Pith. sign in

REVIEW 2 major objections 4 minor 85 references

A Critical Field Guide for Working with Machine Learning Datasets

T0 review · 2 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper argues that conscientious dataset stewardship—working through lifecycle questions on origins, usage, and stewardship—can help practitioners avoid technical, legal, and ethical harms and build more reliable machine learning…

desk verdict A well-made teaching synthesis, not a research contribution; the efficacy promise is untested and should be read as aspiration. read the letter →

arxiv 2501.15491 v1 pith:RUPIZUPQ submitted 2025-01-26 cs.CY

classification cs.CY
keywords machinelearningdatasetsdatasetlifecycledatastewardshipbiasconsentdeprecationcriticalstudiesdatasheetsfor
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that machine learning datasets are powerful but unwieldy resources whose problems—technical, legal, and ethical—can be managed through conscientious stewardship. It offers a practical field guide giving questions, suggestions, and strategies for every phase of a dataset's life. The central promise is that practitioners who ask the guide's lifecycle questions will be more capable of avoiding dataset-specific harms and constructing more reliable systems. A sympathetic reader would care because the guide translates critical AI scholarship into accessible, usable practices for students, journalists, artists, researchers, and developers.

What carries the argument

The central mechanism is the Dataset Lifecycle framework, a set of critical questions organized into three stages: Origins (what is the dataset's story, who created and consented to it), Usage (what story will you tell with it), and Stewardship (what story will it keep telling after you). Each stage prompts reflection on provenance, consent, annotation, missing data, transformation, licensing, harm mitigation, documentation, and deprecation. The framework carries the argument by turning the abstract claim that 'datasets are not neutral' into a repeatable questioning practice.

What would settle it

A controlled study in which two groups of practitioners—one using the lifecycle questions, one not—build or select datasets for comparable tasks, with independent audit of resulting harms, errors, and legal or ethical issues, would settle the claim; if the question-using group shows no measurable improvement, the guide's central value proposition fails.

Watch

Extended reading notes

Core claim

The guide's central claim is that no dataset is neutral or ready to use off the shelf: datasets are contingent on how they are made, who made them, and the settings in which they circulate, and they remain tied to the people they represent and affect. Working through the lifecycle framework—Origins, Usage, and Stewardship—makes these entanglements visible and gives practitioners concrete questions to ask before, during, and after a project. The paper asserts that such critical care yields more robust datasets, more reliable results, greater protection from legal and ethical liability, and more conscientious outcomes for those impacted.

Load-bearing premise

The guide assumes that reflecting on these questions will actually change practitioners' decisions and reduce harm, but it provides no evidence or user study that the questions alter behavior or outcomes.

Editorial extensions

If this is right

  • Practitioners who work through the Origins questions are more likely to detect biased collection methods, missing consent, and licensing restrictions before a project starts.
  • Practitioners who apply the Usage questions are more likely to catch transformation choices that erase context or skew results, such as dropping missing data or binning continuous values.
  • Practitioners who follow the Stewardship questions are more likely to document derivatives, monitor for dataset deprecation, and plan for ethical archiving.
  • Widespread adoption of the lifecycle questions would push datasheets and deprecation frameworks toward field-standard practice, reducing the continued use of 'zombie datasets.'

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The lifecycle questions could be turned into a routinized checklist embedded in dataset repositories or model registries, making the critical reflection the guide advocates a structural part of dataset access.
  • A testable extension would be to measure whether datasets accompanied by completed datasheets and lifecycle documentation are reused less frequently in inappropriate contexts, or attract more caution from downstream users.
  • The framework's logic generalizes beyond existing datasets to generated or synthetic data, where provenance and consent questions become murkier and the guide's emphasis on documenting transformations is even more salient.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This manuscript is a field guide to working critically with machine learning datasets. It introduces key concepts (data, datasets, parts of datasets, types, transformations) and offers a lifecycle framework with questions about origins, usage, and stewardship. It draws on critical data studies, documents common pitfalls, and provides references to tools and practices such as datasheets, deprecation frameworks, and FAIR/CARE principles. The guide is written for a broad audience including researchers, journalists, artists, and developers.

Significance. If adopted, the guide could serve as a useful pedagogical and reference resource that bridges critical AI scholarship and practical data work. Its strengths are the breadth of synthesized literature, the concrete questions and checklists, the inclusion of case studies, and the clear distinction between technical and sociotechnical dimensions of dataset care. However, the paper makes strong efficacy claims about improved outcomes without empirical support, and some benefit claims (e.g., reduced legal liability) are overstated. With appropriate qualifications, the guide would be a valuable contribution to the critical data studies literature.

major comments (2)
  1. [Abstract; Sections 1.1, 2, 6; Section 8] The guide repeatedly claims that working through its recommendations will make practitioners 'more capable of avoiding the problems unique to datasets' and 'construct more reliable, robust solutions.' The only mechanism proposed is reflection on a set of questions; no user study, behavioral outcome measure, or comparison against alternative interventions is presented. The conclusion (Section 8) hedges by calling the guide 'a starting point,' but the earlier statements are unqualified. This is a load-bearing issue because the guide's value proposition rests on the assumption that reflection changes dataset practice. Please either soften these claims to indicate potential benefit, or add an explicit discussion of the evidence status and limitations, noting that the efficacy of such checklists remains an open empirical question.
  2. [Section 2] The claim that proactive attention to legal and ethical concerns yields 'INCREASED PROTECTION FROM LIABILITY' is stated as a benefit. Although the text disclaims that it is not legal advice, this is a factual/legal assertion that is not substantiated with legal analysis or citations to specific remedies. At a minimum, phrase this as 'may reduce legal and ethical risk' and advise readers to consult counsel, rather than implying guaranteed protection.
minor comments (4)
  1. [Section 1.1] The overview states that readers will find TYPES of datasets in Section 5 and TRANSFORM in Section 4, but the Table of Contents and section headers place Types in Section 4 and Transforming in Section 5. Please correct the cross-references.
  2. [Section 3] The invented term 'Data subjectees' is defined, but it could be confused with 'data subjects' throughout the guide. Since the authors already cite the direct/indirect stakeholder distinction from Friedman and Hendry, consider using that established vocabulary to improve clarity.
  3. [References] References [57] and [53] are duplicates of the same CARE Principles citation; please consolidate.
  4. [Section 7.3] Long sequences of block characters appear to be intentional design elements from the spreadsheet origin of the guide. In a text-only version these render as garbled characters and may be inaccessible; consider replacing them with descriptive text or an accessible version.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the guide contains no derivations or fitted predictions; its claims are practical recommendations grounded in external scholarship.

full rationale

This document is a field guide rather than a technical derivation, so the circularity axis is essentially vacuous. It contains no equations, no fitted parameters, no statistical predictions, and no uniqueness theorem. Its central claims — that asking critical questions about dataset origins, usage, and stewardship can help practitioners avoid harms — are normative exhortations with practical checklists, not results derived from those checklists. The abstract and Section 2 make an efficacy promise, and the reader's take correctly notes that the promise is unsupported by a user study or outcome evaluation; however, that is an evidentiary weakness, not circularity, because the advice does not reduce to its own conclusion by construction. The paper's self-references to Knowing Machines outputs (e.g., the deprecation framework in [50] and the reading list in [22]) are citations to prior work used as recommended resources, not load-bearing derivations: the guide does not claim to prove the framework, it points to it as a tool. The conclusion explicitly hedges, calling the guide 'not as a definitive source' but 'a starting point,' which further undercuts any reading of the text as claiming a forced or self-validating result. Because there is no specific equation, fitted parameter, or self-citation chain that makes a prediction equivalent to its input, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

The guide makes no quantitative claims, so the free-parameter ledger is empty. Its load-bearing premises are philosophical and normative assumptions imported from the critical data studies literature (the relational nature of data, the impossibility of anonymization, the value of participatory design), adopted rather than argued for in the text. The single invented term is 'data subjectees' (Section 3), a definitional neologism with no empirical handle. The guide's own statements at Section 8 ('not a definitive source... a starting point') and the disclaimers in Sections 6.1 and 7.3 partially acknowledge the limits of these premises.

assumptions (3)
  • domain assumption Data are a relational category: what counts as data depends on who uses them, how, and for which purposes (Leonelli).
    Adopted in Section 1.2 as the philosophical foundation for the whole guide. It is a stance from the philosophy of science, asserted here as the correct lens rather than demonstrated.
  • domain assumption There is no such thing as anonymizing data; re-identification is always possible.
    Section 6.2 states this categorically ('There is no such thing as anonymizing identifying data') and the consent and access guidance in Sections 6.1 and 6.3 depends on it. It is a widely held privacy-research position, presented without supporting evidence in the guide.
  • domain assumption Participatory, community-led design produces better and less harmful outcomes ('Build with, not for').
    Invoked in Sections 6.1 and 6.2 via the Design Justice Network. The guide's repeated recommendation to consult data subjects and subjectees assumes this participatory premise; no evidence for its efficacy is given.
invented entities (1)
  • Data subjectees
    purpose: A new term for people affected by a machine learning system's outputs even when their data is not in the dataset, distinct from data subjects whose data is contained in it.
    Introduced in Section 3. It is a definitional neologism, useful for drawing the distinction between direct and indirect stakeholders (following Friedman and Hendry), but it has no operationalization or empirical handle beyond the guide's usage of the word.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Critical Field Guide for Working with Machine Learning Datasets." pith.science (2026). https://pith.science/paper/RUPIZUPQ

@misc{pith2026250115491,
  author       = {Pith},
  title        = {Pith review of: A Critical Field Guide for Working with Machine Learning Datasets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RUPIZUPQ}},
  note         = {Machine review of arXiv:2501.15491}
}
read the original abstract

Machine learning datasets are powerful but unwieldy. Despite the fact that large datasets commonly contain problematic material--whether from a technical, legal, or ethical perspective--datasets are valuable resources when handled carefully and critically. A Critical Field Guide for Working with Machine Learning Datasets suggests practical guidance for conscientious dataset stewardship. It offers questions, suggestions, strategies, and resources for working with existing machine learning datasets at every phase of their lifecycle. It combines critical AI theories and applied data science concepts, explained in accessible language. Equipped with this understanding, students, journalists, artists, researchers, and developers can be more capable of avoiding the problems unique to datasets. They can also construct more reliable, robust solutions, or even explore new ways of thinking with machine learning datasets that are more critical and conscientious.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 62 canonical work pages

  1. [4]

    Browne, Dark matters: on the surveillance of blackness

    S. Browne, Dark matters: on the surveillance of blackness. Durham, [North Carolina] ; Duke University Press, 2015

  2. [5]

    Making data colonialism liveable: how might data’s social order be regulated?,

    N. Couldry and U. A. Mejias, “Making data colonialism liveable: how might data’s social order be regulated?,” Internet Policy Rev., vol. 8, no. 2, Jun. 2019, Accessed: Mar. 21, 2021. https://policyreview.info/articles/analysis/making-data-colonialism-liveable-how-might-datas-social-order-be-regulated

  3. [6]

    Machine Bias,

    J. Angwin, J. Larson, S. Mattu, and L. Kirchner, “Machine Bias,” ProPublica, May 2016, Accessed: Apr. 27, 2019. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing

  4. [7]

    17. Classification — Computational and Inferential Thinking,

    D. Wagner, “17. Classification — Computational and Inferential Thinking,” in Computational and Inferential Thinking: The Foundations of Data Science, 2nd ed., A. Adhikari, J. DeNero, and D. Wagner, Eds. Accessed: Nov. 25, 2022. https://computerscience.chemeketa.edu/datasci-text/chapters/17/Classification.html

  5. [8]

    Leonelli, Data-Centric Biology: A Philosophical Study

    S. Leonelli, Data-Centric Biology: A Philosophical Study. University of Chicago Press, 2016. doi: 10.7208/chicago/9780226416502.001.0001

  6. [9]

    The Relevance of Algorithms,

    T. Gillespie, “The Relevance of Algorithms,” T. Gillespie, P. J. Boczkowski, and K. A. Foot, Eds. Cambridge, MA: MIT Press, 2014

  7. [10]

    Raw data

    L. Gitelman, Ed., “Raw data” is an oxymoron. Cambridge, Massachusetts ; London, England: The MIT Press, 2013

  8. [11]

    Koopman, How We Became Our Data: A Genealogy of the Informational Person

    C. Koopman, How We Became Our Data: A Genealogy of the Informational Person. University of Chicago Press, 2019. doi: 10.7208/9780226626611

Show all 85 references
  1. [12]

    C. L. Borgman, Big Data, Little Data, No Data: Scholarship in the Networked World. 2015. doi: 10.7551/mitpress/9963.001.0001

  2. [13]

    Kitchin, Data Lives

    R. Kitchin, Data Lives. Policy Press, 2021

  3. [14]

    Gleick, The information: a history, a theory, a flood, 1st Vintage Books ed., 2012

    J. Gleick, The information: a history, a theory, a flood, 1st Vintage Books ed., 2012. New York: Vintage Books, 2011

  4. [15]

    N. B. Thylstrup, D. Agostinho, A. Ring, C. D’Ignazio, and K. Veel, Eds., Uncertain Archives: Critical Keywords for Big Data. 2021. Accessed: Mar. 29, 2021. [Online]. Available: https://doi.org/10.7551/mitpress/12236.001.0001

  5. [16]

    Algorithmic culture,

    T. Striphas, “Algorithmic culture,” Eur. J. Cult. Stud., vol. 18, no. 4–5, pp. 395–412, Aug. 2015, doi: 10.1177/1367549415577392

  6. [17]

    Datson, Rules

    L. Datson, Rules. Princeton University Press, 2022. https://press.princeton.edu/books/hardcover/9780691156989/rules

  7. [18]

    Chollet, Deep Learning with Python, Second Edition

    F. Chollet, Deep Learning with Python, Second Edition. New York: Manning Publications Co. LLC, 2021

  8. [19]

    The ontology explorer: A method to make visible data infrastructures for population management,

    W. Van Rossem and A. Pelizza, “The ontology explorer: A method to make visible data infrastructures for population management,” Big Data Soc., vol. 9, no. 1, p. 20539517221104090, Jan. 2022, doi: 10.1177/20539517221104087

  9. [20]

    Lawsuits allege Microsoft, Amazon and Google violated Illinois facial recognition privacy law,

    “Lawsuits allege Microsoft, Amazon and Google violated Illinois facial recognition privacy law,” TechCrunch. https://social.techcrunch.com/2020/07/15/facial-recognition-lawsuit-vance-janecyk-bipa/

  10. [21]

    Facial recognition’s ‘dirty little secret’: Social media photos used without consent,

    “Facial recognition’s ‘dirty little secret’: Social media photos used without consent,” NBC News. https://www.nbcnews.com/tech/internet/facial-recognition-s-dirty-little-secret-millions-online-photos-scraped-n981921

  11. [22]

    Critical Dataset Studies Reading List,

    F. Corry, E. B. Kang, H. Sridharan, S. Luccioni, M. Ananny, and K. Crawford, “Critical Dataset Studies Reading List,” Knowing Machines. https://knowingmachines.org/reading-list

  12. [23]

    Yolanda Gil: Teaching Data Science to Non-Programmers,

    Y. Gil, “Yolanda Gil: Teaching Data Science to Non-Programmers,” Jan. 03, 2020. https://www.isi.edu/~gil/teaching/TeachingDataScienceToNonProgrammers.html

  13. [24]

    Datasheets for Datasets,

    T. Gebru et al., “Datasheets for Datasets,” ArXiv180309010 Cs, Mar. 2020, Accessed: Apr. 01, 2021. http://arxiv.org/abs/1803.09010

  14. [25]

    Crawford, Atlas of AI: power, politics, and the planetary costs of artificial intelligence

    K. Crawford, Atlas of AI: power, politics, and the planetary costs of artificial intelligence. New Haven: Yale University Press, 2021

  15. [26]

    Cady, The data science handbook

    F. Cady, The data science handbook. Hoboken, NJ: John Wiley & Sons, Inc., 2017

  16. [27]

    Friedman and D

    B. Friedman and D. G. Hendry, Value Sensitive Design: Shaping Technology with Moral Imagination. 2019. doi: 10.7551/mitpress/7585.001.0001

  17. [28]

    Using artificial intelligence, geo-journalism and data journalism, journalists dodge some of the dangers of covering the Amazon,

    “Using artificial intelligence, geo-journalism and data journalism, journalists dodge some of the dangers of covering the Amazon,” LatAm Journalism Review by the Knight Center, Jul. 19, 2022. https://latamjournalismreview.org/articles/using-artificial-intelligence-geo-journali...

  18. [30]

    Ithaca | Restoring and attributing ancient texts using deep neural networks,

    “Ithaca | Restoring and attributing ancient texts using deep neural networks,” Ithaca. https://ithaca.deepmind.com

  19. [31]

    On the Endless Infrastructural Reach of a Phoneme,

    P. Oliveira, “On the Endless Infrastructural Reach of a Phoneme,” Transmedialeart Digit. Cult., no. 3, https://archive.transmediale.de/content/on-the-endless-infrastructural-reach-of-a-phoneme

  20. [32]

    CALLFRIEND Egyptian Arabic

    Canavan, Alexandra and Zipperlen, George, “CALLFRIEND Egyptian Arabic.” Linguistic Data Consortium, p. 1401448 KB, 1996. doi: 10.35111/NNM5-KP69

  21. [33]

    CALLHOME Egyptian Arabic Speech

    Canavan, Alexandra, Zipperlen, George, and Graff, David, “CALLHOME Egyptian Arabic Speech.” Linguistic Data Consortium, p. 1807744 KB, 1997. doi: 10.35111/D8YB-9M13

  22. [34]

    Eine Software des BAMF bringt Menschen in Gefahr,

    A. Biselli, “Eine Software des BAMF bringt Menschen in Gefahr,” Vice, Aug. 20, 2018. https://www.vice.com/de/article/a3q8wj/fluechtlinge-bamf-sprachanalyse-software-entscheidet-asyl

  23. [35]

    Personal conversation,

    P. Oliveira, “Personal conversation,” Aug. 16, 2022

  24. [36]

    On̸ ̸The App̸arentl̸y ̸Me̸aningl̸e̸ss Texture of Noise̸

    P. Oliveira, “On̸ ̸The App̸arentl̸y ̸Me̸aningl̸e̸ss Texture of Noise̸.” http://meaninglesstexture.schloss-post.com/

  25. [37]

    Researchers train AI on ‘synthetic data’ to uncover Syrian war crimes,

    M. Murgia, “Researchers train AI on ‘synthetic data’ to uncover Syrian war crimes,” FT.com, Dec. 2021, Ahttps://www.proquest.com/docview/2617702310/citation/42933059DCF5492DPQ/1

  26. [38]

    AI Emerges as Crucial Tool for Groups Seeking Justice for Syria War Crimes,

    “AI Emerges as Crucial Tool for Groups Seeking Justice for Syria War Crimes,” Dow Jones Institutional News, Dow Jones & Company Inc, New York, United States, Feb. 13, 2021. http://www.proquest.com/docview/2489065006/citation/C2B3DB568C8D4995PQ/1

  27. [39]

    Data Science 4 All – a user friendly data science learning site

    “Data Science 4 All – a user friendly data science learning site.” https://datascience4all.org/

  28. [40]

    McKinney, Python for Data Analysis, 3E, Open 3rd Edition

    W. McKinney, Python for Data Analysis, 3E, Open 3rd Edition. O’Reilly. https://wesmckinney.com/book/

  29. [41]

    Accessed: Jul

    Cleaning Data for Effective Data Science. Accessed: Jul. 31, 2022. https://learning.oreilly.com/library/view/cleaning-data-for/9781801071291/

  30. [42]

    Consider Data Cleaning v1.1,

    K. Kuksenok, “Consider Data Cleaning v1.1,” presented at the Resistance AI Workshop at NeurIPS2020, Nov. 2020. https://ksen0.github.io/code-data-work/

  31. [43]

    McPherson, Feminist in a Software Lab: Difference + Design

    T. McPherson, Feminist in a Software Lab: Difference + Design. Cambridge, Massachusetts ; London, England: Harvard University Press, 2018

  32. [44]

    The Library of Missing Datasets — MIMI ỌNỤỌHA,

    M. Onuoha, “The Library of Missing Datasets — MIMI ỌNỤỌHA,” MIMI ỌNỤỌHA. https://mimionuoha.com/the-library-of-missing-datasets

  33. [45]

    Demarginalizing the Intersection of Race and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory and Antiracist Politics,

    K. Crenshaw, “Demarginalizing the Intersection of Race and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory and Antiracist Politics,” Univ. Chic. Leg. Forum, vol. 1989, pp. 139–168, 1989

  34. [46]

    The Point of Collection,

    M. Onuoha, “The Point of Collection,” Medium, Oct. 31, 2016. https://points.datasociety.net/the-point-of-collection-8ee44ad7c2fa

  35. [47]

    Responsible Data Handbook | Getting Data

    “Responsible Data Handbook | Getting Data.” https://the-engine-room.github.io/responsible-data-handbook/chapters/chapter-02a-getting-data.html

  36. [48]

    Responsible Data Handbook

    The Engine Room, “Responsible Data Handbook.” https://the-engine-room.github.io/responsible-data-handbook/

  37. [49]

    Whose Ground Truth? Accounting for Individual and Collective Identities Underlying Dataset Annotation,

    E. Denton, M. Díaz, I. Kivlichan, V. Prabhakaran, and R. Rosen, “Whose Ground Truth? Accounting for Individual and Collective Identities Underlying Dataset Annotation,” arXiv, arXiv:2112.04554, Dec. 2021. doi: 10.48550/arXiv.2112.04554

  38. [50]

    A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication,

    A. S. Luccioni, F. Corry, H. Sridharan, M. Ananny, J. Schultz, and K. Crawford, “A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication,” in 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun....

  39. [51]

    Design Justice for Action,

    Design Justice Network, “Design Justice for Action,” Des. Justice Zines, no. #3

  40. [52]

    Ten simple rules for responsible big data research,

    M. Zook et al., “Ten simple rules for responsible big data research,” PLOS Comput. Biol., vol. 13, no. 3, p. e1005399, Mar. 2017, doi: 10.1371/journal.pcbi.1005399

  41. [54]

    The Data Ethics Canvas

    Open Data Institute, “The Data Ethics Canvas.” https://theodi.org/article/the-data-ethics-canvas-2021/

  42. [55]

    Local Contexts – Grounding Indigenous Rights

    “Local Contexts – Grounding Indigenous Rights.” https://localcontexts.org/

  43. [56]

    FAIR Principles,

    “FAIR Principles,” GO FAIR. https://www.go-fair.org/fair-principles/

  44. [57]

    CARE Principles of Indigenous Data Governance,

    “CARE Principles of Indigenous Data Governance,” Global Indigenous Data Alliance. https://www.gida-global.org/care

  45. [58]

    The Problem With Bias: Allocative Versus Representational Harms in Machine Learning.,

    S. Barocas, K. Crawford, A. Shapiro, and H. Wallach, “The Problem With Bias: Allocative Versus Representational Harms in Machine Learning.,” in Proceedings of SIGCIS, Philadelphia, PA, 2017

  46. [59]

    Benjamin, Race After Technology: Abolitionist Tools for the New Jim Code, 1 edition

    R. Benjamin, Race After Technology: Abolitionist Tools for the New Jim Code, 1 edition. Medford, MA: Polity, 2019

  47. [60]

    Data and its (dis)contents: A survey of dataset development and use in machine learning research,

    A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis)contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11, p. 100336, Nov. 2021, doi: 10.1016/j.patter.2021.100336

  48. [61]

    A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle,

    H. Suresh and J. V. Guttag, “A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle,” in Equity and Access in Algorithms, Mechanisms, and Optimization, Oct. 2021, pp. 1–9. doi: 10.1145/3465416.3483305

  49. [62]

    The Data-Production Dispositif

    M. Miceli and J. Posada, “The Data-Production Dispositif.” arXiv, May 24, 2022. doi: 10.48550/arXiv.2205.11963

  50. [64]

    Apprich, W

    C. Apprich, W. H. K. Chun, F. Cramer, and H. Steyerl, Pattern Discrimination. Minneapolis: University of Minnesota Press, 2018

  51. [65]

    Feeling fixes: Mess and emotion in algorithmic audits,

    O. Keyes and J. Austin, “Feeling fixes: Mess and emotion in algorithmic audits,” Big Data Soc., vol. 9, no. 2, p. 20539517221113772, Jul. 2022, doi: 10.1177/20539517221113772

  52. [66]

    Language (Technology) is Power: A Critical Survey of ‘Bias’ in NLP

    S. L. Blodgett, S. Barocas, H. Daumé III, and H. Wallach, “Language (Technology) is Power: A Critical Survey of ‘Bias’ in NLP.” arXiv, May 29, 2020. doi: 10.48550/arXiv.2005.14050

  53. [67]

    G. C. Bowker and S. L. Star, Sorting things out: classification and its consequences. Cambridge, Mass: MIT Press, 1999

  54. [68]

    Large image datasets: A pyrrhic win for computer vision?

    V. U. Prabhu and A. Birhane, “Large image datasets: A pyrrhic win for computer vision?” arXiv, Jul. 23, 2020. doi: 10.48550/arXiv.2006.16923

  55. [69]

    Blair, P

    A. Blair, P. Duguid, A.-S. Goeing, and A. Grafton, Information: A Historical Companion. Princeton University Press, 2021. doi: 10.1515/9780691209746

  56. [70]

    The Power to Name: Representation in Library Catalogs,

    H. A. Olson, “The Power to Name: Representation in Library Catalogs,” SignsJournal Women Cult. Soc., vol. 26, no. 3, 2001, doi: 10.1086/495624

  57. [71]

    Critical Algorithm Studies: a Reading List,

    T. Gillespie and N. Seaver, “Critical Algorithm Studies: a Reading List,” Social Media Collective, Nov. 05, 2015. https://socialmediacollective.org/reading-lists/critical-algorithm-studies/

  58. [72]

    50 Years of Test (Un)fairness: Lessons for Machine Learning,

    B. Hutchinson and M. Mitchell, “50 Years of Test (Un)fairness: Lessons for Machine Learning,” in Proceedings of the Conference on Fairness, Accountability, and Transparency, Atlanta GA USA, Jan. 2019, pp. 49–58. doi: 10.1145/3287560.3287600

  59. [73]

    Adversarially Constructed Evaluation Sets Are More Challenging, but May Not Be Fair

    J. Phang, A. Chen, W. Huang, and S. R. Bowman, “Adversarially Constructed Evaluation Sets Are More Challenging, but May Not Be Fair.” arXiv, Nov. 15, 2021.http://arxiv.org/abs/2111.08181

  60. [74]

    A Survey on Bias and Fairness in Machine Learning

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A Survey on Bias and Fairness in Machine Learning.” arXiv, Jan. 25, 2022. http://arxiv.org/abs/1908.09635

  61. [75]

    State-of-the-art generalisation research in NLP: a taxonomy and review

    D. Hupkes et al., “State-of-the-art generalisation research in NLP: a taxonomy and review.” arXiv, Oct. 10, 2022. http://arxiv.org/abs/2210.03050

  62. [76]

    Explainable deep learning for efficient and robust pattern recognition: A survey of recent developments,

    X. Bai et al., “Explainable deep learning for efficient and robust pattern recognition: A survey of recent developments,” Pattern Recognit., vol. 120, p. 108102, Dec. 2021, doi: 10.1016/j.patcog.2021.108102

  63. [77]

    Demarginalizing the Intersection of Race and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory and Antiracist Politics,

    K. Crenshaw, What Does Intersectionality Mean? : 1A. 2021. https://www.npr.org/2021/03/29/982357959/what-does-intersectionality-mean [78]K. Crenshaw, “Demarginalizing the Intersection of Race and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory an...

  64. [79]

    Intersectionality,

    B. Cooper, “Intersectionality,” in The Oxford Handbook of Feminist Theory, vol. 1, L. Disch and M. Hawkesworth, Eds. Oxford University Press, 2016. doi: 10.1093/oxfordhb/9780199328581.013.20

  65. [80]

    Intersectionality,

    B. Gipson, F. Corry, and S. U. Noble, “Intersectionality,” in Uncertain Archives: Critical Keywords for Big Data, 2021. https://doi.org/10.7551/mitpress/12236.003.0027

  66. [81]

    S. U. Noble and B. M. Tynes, Eds., The intersectional Internet : race, sex, class and culture online. New York: Peter Lang Publishing, Inc, 2016

  67. [82]

    Intersectional AI Toolkit,

    S. Ciston, “Intersectional AI Toolkit,” Intersectional AI Toolkit. https://intersectionalai.com/

  68. [83]

    Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation,

    A. Wang, V. V. Ramaswamy, and O. Russakovsky, “Towards Intersectionality in Machine Learning: Including More Identities, Handling Underrepresentation, and Performing Evaluation,” in 2022 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA, Jun. 2022...

  69. [84]

    Don’t ask if artificial intelligence is good or fair, ask how it shifts power,

    P. Kalluri, “Don’t ask if artificial intelligence is good or fair, ask how it shifts power,” Nature, vol. 583, no. 7815, Art. no. 7815, Jul. 2020, doi: 10.1038/d41586-020-02003-2

  70. [85]

    Data Violence and How Bad Engineering Choices Can Damage Society,

    A. L. Hoffmann, “Data Violence and How Bad Engineering Choices Can Damage Society,” Medium, Apr. 30, 2018. https://medium.com/s/story/data-violence-and-how-bad-engineering-choices-can-damage-society-39e44150e1d4

  71. [86]

    Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?,

    M. Miceli, J. Posada, and T. Yang, “Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?,” Proc. ACM Hum.-Comput. Interact., vol. 6, no. GROUP, p. 34:1-34:14, Jan. 2022, doi: 10.1145/3492853

  72. [87]

    Managing Bias When Library Collections Become Data,

    C. N. Coleman, “Managing Bias When Library Collections Become Data,” Int. J. Librariansh., vol. 5, no. 1, Art. no. 1, Jul. 2020, doi: 10.23974/ijol.2020.vol5.1.162

  73. [88]

    Sinclair and J

    K. Sinclair and J. Clark, Making a New Reality. 2020. https://makinganewreality.org/making-a-new-reality-a-toolkit-for-inclusive-media-futures-a3bdc0e68f20

  74. [89]

    Hugging Face

    “Hugging Face.” https://huggingface.co/datasets

  75. [90]

    “Kaggle.” https://www.kaggle.com/datasets

  76. [91]

    Papers With Code

    “Papers With Code.” https://paperswithcode.com/datasets

  77. [92]

    re3data.org

    “re3data.org.” https://www.re3data.org/ . [93]“Zenodo.” https://zenodo.org/ . [94]M. Khan and A. Hanna, “The Subjects and Stages of AI Dataset Development: A Framework for Dataset Accountability.” Rochester, NY, Sep. 13, 2022. Accessed: Sep. 19, 2022. https://papers.ssrn.com/a...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.