Pith. sign in

REVIEW 2 major objections 2 minor 160 references

Evaluating Structured Documentation as a Tool for Reflexivity in Dataset Development

T0 review · 2 major / 2 minor · reviewed 2026-05-13 · grok-4.3

Pith's one-line read Structured documentation frameworks like datasheets engage little with major reflexivity themes from the literature.

desk verdict The paper documents limited reflexivity in dataset docs with new codebook and questions, but the general claim needs stronger sampling evidence. read the letter →

arxiv 2605.11345 v1 submitted 2026-05-11 cs.CY

classification cs.CY
keywords reflexivitydatasetdocumentationdatasheetsdatastatementsmachinelearningFAccTthematicanalysisdevelopment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper investigates whether tools such as datasheets, data statements, and nutrition labels actually help dataset developers practice reflexivity about the value judgments in their work. It applies thematic analysis to reflexivity literature and discourse analysis to the frameworks plus published examples of their use. The analysis reveals little incorporation of core reflexivity ideas such as positionality, power relations, or contextual values. This finding matters because dataset creation shapes downstream machine learning systems, yet current documentation may not prompt the reflection needed to surface those influences. The authors supply a codebook of reflexivity topics, practical strategies, and a set of extended questions for datasheets to close the gap.

What carries the argument

A codebook of reflexivity topics derived from the literature, used to evaluate incorporation in documentation frameworks and their published applications via thematic and discourse analysis.

What would settle it

A broader survey that identifies frequent, detailed engagement with multiple reflexivity themes such as positionality and value conflicts across a larger set of published dataset documentations would undermine the general-lack finding.

Watch

Extended reading notes

Core claim

Through mixed-method thematic analysis of reflexivity literature and corpus-assisted discourse analysis of frameworks and applications, the paper establishes a general lack of engagement with major reflexivity themes in both the design of structured dataset documentation and in how those frameworks are applied in published work.

Load-bearing premise

The chosen reflexivity themes from the literature and the selected sample of frameworks plus published applications are representative enough to support a claim of general lack beyond the cases examined.

Editorial extensions

If this is right

  • Framework creators should revise datasheets and similar tools to include questions that explicitly prompt reflexivity on positionality and power.
  • Dataset developers can apply the provided codebook and extended questions to surface value-laden choices during documentation.
  • Published applications of documentation frameworks can serve as better models if they address reflexivity themes more directly.
  • The FAccT community can use the gap identified to prioritize reflexivity in future dataset work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Adding reflexivity prompts might change how developers document datasets in practice, which could be tested by comparing before-and-after documentation quality.
  • The finding points to a possible mismatch between stated goals of documentation tools and their actual effects on critical reflection.
  • Extending the analysis to documentation in non-academic settings such as industry datasets could reveal whether the lack is specific to research publications.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper claims that structured dataset documentation frameworks (e.g., datasheets, data statements, dataset nutrition labels), despite their stated goal of facilitating reflexivity in ML dataset development, show a general lack of engagement with major reflexivity themes drawn from FAccT and related literature. This holds both for the frameworks themselves and for published applications of the frameworks. The authors use mixed-method thematic analysis and corpus-assisted discourse analysis to derive a codebook of reflexivity topics, empirically demonstrate the gap, and propose actionable strategies plus a set of extended datasheet questions to address it.

Significance. If the sampling and analysis support the generalization to a 'general lack,' the work would be significant for the FAccT and responsible AI communities by providing empirical evidence of a disconnect between the reflexive intent of documentation frameworks and their actual content and usage. The codebook, recommendations, and concrete extended questions offer practical value for improving future frameworks and dataset practices. The mixed-methods design and focus on actionable outputs are strengths that could help translate the findings into impact.

major comments (2)
  1. [Methods] Methods section: The paper does not provide sufficient detail on the inclusion criteria, search strategy, time periods, venues, or keywords used to select the dataset documentation frameworks and the corpus of published applications. Given that the central empirical claim is a 'general lack' across the field (rather than within a convenience sample), the representativeness of the chosen frameworks and applications must be explicitly justified and documented to support the generalization.
  2. [Analysis and Results] Analysis and Results: The thematic analysis would be strengthened by including inter-coder reliability metrics, the full codebook with definitions and examples of application to framework text and published uses, and a clear mapping from the reflexivity literature themes to the coded categories. Without these, it is difficult to assess whether the evidence robustly supports the 'general lack' finding or whether alternative theme selections could alter the conclusion.
minor comments (2)
  1. [Abstract] Abstract: Consider adding a brief statement of the number of frameworks examined and the size of the application corpus to immediately convey the empirical scope.
  2. [Discussion] Discussion: A dedicated limitations subsection would help by addressing potential selection biases in the reflexivity themes, frameworks, and corpus, as well as the generalizability of the proposed extended questions.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their constructive and detailed feedback. The comments highlight important areas for improving transparency and rigor, and we will incorporate revisions to address them fully.

read point-by-point responses
  1. Referee: [Methods] Methods section: The paper does not provide sufficient detail on the inclusion criteria, search strategy, time periods, venues, or keywords used to select the dataset documentation frameworks and the corpus of published applications. Given that the central empirical claim is a 'general lack' across the field (rather than within a convenience sample), the representativeness of the chosen frameworks and applications must be explicitly justified and documented to support the generalization.

    Authors: We agree that more explicit documentation of our sampling process is required to support the generalization. In the revised manuscript, we will expand the Methods section to detail the inclusion criteria, search strategy, time periods, venues, and keywords used to identify the documentation frameworks and the corpus of published applications. We will also add a justification of representativeness, explaining how the selected frameworks represent the major approaches in the literature and how the applications corpus provides broad coverage, while noting the boundaries of the sample. revision: yes

  2. Referee: [Analysis and Results] Analysis and Results: The thematic analysis would be strengthened by including inter-coder reliability metrics, the full codebook with definitions and examples of application to framework text and published uses, and a clear mapping from the reflexivity literature themes to the coded categories. Without these, it is difficult to assess whether the evidence robustly supports the 'general lack' finding or whether alternative theme selections could alter the conclusion.

    Authors: We agree these additions will strengthen the presentation of the analysis. In the revision, we will report inter-coder reliability metrics (including Cohen's kappa) from the thematic coding process. The complete codebook with definitions and examples drawn from both framework texts and published applications will be provided in an appendix. We will also add a mapping table that explicitly connects the reflexivity themes identified in the FAccT and related literature to our coded categories. These changes will allow readers to evaluate the robustness of the 'general lack' finding more directly. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical thematic analysis of external frameworks and literature

full rationale

The paper conducts a mixed-method thematic analysis and corpus-assisted discourse analysis on selected dataset documentation frameworks and published applications, drawing reflexivity themes from FAccT and related literature. The central empirical claim of general lack of engagement is presented as an observation from this external evaluation rather than a derivation that reduces to self-defined inputs, fitted parameters, or self-citation chains. No equations, predictions, or uniqueness theorems are invoked; the methodology relies on codebook development and analysis of independently sourced materials. The representativeness concern raised in the skeptic note is a question of sampling validity, not a circular reduction of the result to its own construction.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Based on abstract only: the analysis assumes reflexivity literature provides a stable set of major themes that can be reliably applied to documentation frameworks via thematic analysis.

assumptions (1)
  • domain assumption Reflexivity concepts from FAccT literature can be operationalized into discrete themes suitable for thematic analysis of documentation frameworks.
    The paper relies on this to create its codebook and evaluate engagement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Evaluating Structured Documentation as a Tool for Reflexivity in Dataset Development." pith.science (2026). https://pith.science/paper/2605.11345

@misc{pith2026260511345,
  author       = {Pith},
  title        = {Pith review of: Evaluating Structured Documentation as a Tool for Reflexivity in Dataset Development},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2605.11345}},
  note         = {Machine review of arXiv:2605.11345}
}
read the original abstract

It is prominently recognized that dataset development in machine learning is a value-laden process from problem formulation to data processing, use, and reuse. Structured documentation frameworks such as datasheets, data statements, and dataset nutrition labels have been created to aid developers in documenting how their datasets were produced and, according to the creators of the frameworks, to facilitate reflexivity in dataset development. While reflexivity is a stated goal, it is unclear whether and to what extent these structured dataset documentation frameworks incorporate concepts from reflexivity literature (at FAccT and elsewhere) and whether the use of the frameworks demonstrates reflexivity. Here, we adopt mixed-method thematic analysis and corpus-assisted discourse analysis to explore how reflexivity is incorporated in structured documentation frameworks and their responses. We demonstrate empirically that there is a general lack of engagement with major themes of reflexivity in both dataset documentation frameworks and published applications of these frameworks. We present a codebook of major reflexivity topics, recommend actionable strategies, and propose a set of extended datasheet questions to more effectively incorporate these topics into structured documentation frameworks and in the FAccT literature.

Figures

Figures reproduced from arXiv: 2605.11345 by the authors.

Figure 1
Figure 1. Our multi-stage approach to conceptualizing reflexivity and analyzing it within structured documentation frameworks [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Distribution of topics across individual papers. [PITH_FULL_IMAGE:figures/full_fig_p038_2.png] view at source ↗
Figure 3
Figure 3. Distribution of topics across individual questions. [PITH_FULL_IMAGE:figures/full_fig_p039_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

160 extracted references · 160 canonical work pages

  1. [1]

    [n. d.]. https://alliedmedia.org/projects/our-data-bodies-odb

  2. [2]

    Call For Datasets and Benchmarks - NeurIPS

    2021. Call For Datasets and Benchmarks - NeurIPS. https://web.archive.org/web/20210407213644/https://neurips.cc/Conferences/2021/ CallForDatasetsBenchmarks

  3. [3]

    Call For Datasets and Benchmarks - NeurIPS

    2022. Call For Datasets and Benchmarks - NeurIPS. https://web.archive.org/web/20220521130452/https://neurips.cc/Conferences/2022/ CallForDatasetsBenchmarks

  4. [4]

    Call For Datasets and Benchmarks - NeurIPS

    2023. Call For Datasets and Benchmarks - NeurIPS. https://web.archive.org/web/20230325021706/https://neurips.cc/Conferences/2023/ CallForDatasetsBenchmarks

  5. [5]

    Call For Datasets and Benchmarks - NeurIPS

    2024. Call For Datasets and Benchmarks - NeurIPS. https://neurips.cc/Conferences/2024/CallForDatasetsBenchmarks

  6. [6]

    Gustaf Ahdritz, Nazim Bouatta, Sachin Kadyan, Lukas Jarosch, Dan Berenberg, Ian Fisk, Andrew Martin Watkins, Stephen Ra, Richard Bonneau, and Mohammed AlQuraishi. 2023. OpenProteinSet: Training data for structural biology at scale. Advances in Neural Information Processing Systems

  7. [7]

    Danielle Allard and Tami Oliphant. 2024. With a Little Help from Our Friends: Applying a Critical Friends Orientation to Critical Literature Reviews.Proceedings of the Association for Information Science and Technology61, 1 (2024), 13–24. doi:10.1002/pra2.1004

  8. [8]

    Doris Allhutter and Bettina Berendt. 2020. Deconstructing FAT: using memories to collectively explore implicit assumptions, values and context in practices of debiasing and discrimination-awareness. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* ’20). Association for Computing Machinery, New York, NY, USA, 687. do...

Show all 160 references
  1. [9]

    Clyde Ancarno. 2020. Corpus-Assisted Discourse Studies. InThe Cambridge Handbook of Discourse Studies, Alexandra Georgakopoulou and Anna De Fina (Eds.). Cambridge University Press, Cambridge, 165–185

  2. [10]

    Gabriele Bammer. 2017. Toolkits for transdisciplinary research. https://i2insights.org/2017/07/25/toolkits-for-transdisciplinarity/

  3. [11]

    Shaowen Bardzell and Jeffrey Bardzell. 2011. Towards a feminist HCI methodology: social science, feminism, and HCI. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’11). Association for Computing Machinery, New York, NY, USA, 675–684. doi:10.1...

  4. [12]

    Björn Barz and Joachim Denzler. 2021. WikiChurches: A Fine-Grained Dataset of Architectural Styles with Real-World Challenges. Advances in Neural Information Processing Systems

  5. [13]

    2023.Insolvent: How to Reorient Computing for Just Sustainability

    Christoph Becker. 2023.Insolvent: How to Reorient Computing for Just Sustainability. MIT Press

  6. [14]

    Bilel Benbouzid. 2023. Fairness in machine learning from the perspective of sociology of statistics: How machine learning is becoming scientific by turning its back on metrological realism. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency ...

  7. [15]

    Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? . InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21). Association for ...

  8. [16]

    Eshta Bhardwaj, Harshit Gujral, Siyi Wu, Ciara Zogheib, Tegan Maharaj, and Christoph Becker. 2024. Machine Learning Data Practices through a Data Curation Lens: An Evaluation Framework. In2024 ACM Conference on Fairness, Accountability, and Transparency. Association for Comput...

  9. [17]

    Eshta Bhardwaj, Harshit Gujral, Siyi Wu, Ciara Zogheib, Tegan Maharaj, and Christoph Becker. 2024. The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track.Advances in Neural Information Processing Systems37 (20...

  10. [18]

    2009.Natural language processing with Python

    Steven Bird, Ewan Klein, and Edward Loper. 2009.Natural language processing with Python. O’Reilly, Cambridge. FAccT ’26, June 25–28, 2026, Montreal, QC, Canada Bhardwaj and Zogheib, et al

  11. [19]

    Florian Bordes, Shashank Shekhar, Mark Ibrahim, Diane Bouchacourt, Pascal Vincent, and Ari S. Morcos. 2023. PUG: Photorealistic and Semantically Controllable Synthetic Data for Representation Learning. InThirty-seventh Conference on Neural Information Processing Systems Datase...

  12. [20]

    2000.Pascalian meditations

    Pierre Bourdieu. 2000.Pascalian meditations. Stanford University Press, Stanford, Calif

  13. [21]

    2004.Science of Science and Reflexivity

    Pierre Bourdieu. 2004.Science of Science and Reflexivity. University of Chicago Press, Chicago, IL

  14. [22]

    Alicia E Boyd. 2023. QUINTA: Reflexive Sensibility For Responsible AI Research and Data-Driven Processes. doi:abs/2509.16347v1

  15. [23]

    Karen L. Boyd. 2021. Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training Data.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–27. doi:10.1145/3479582

  16. [24]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis.Qualitative Research in Sport, Exercise and Health 11, 4 (2019), 589–597. doi:10.1080/2159676X.2019.1628806

  17. [25]

    Virginia Braun and Victoria Clarke. 2021. Can I use TA? Should I use TA? Should I not use TA? Comparing reflexive thematic analysis and other pattern-based qualitative analytic approaches.Counselling and Psychotherapy Research21, 1 (2021), 37–47. doi:10.1002/capr.12360

  18. [26]

    Bruton, Alicia L

    Samuel V. Bruton, Alicia L. Macchione, Mitch Brown, and Mohammad Hosseini. 2025. Citation Ethics: An Exploratory Survey of Norms and Behaviors.Journal of academic ethics23, 2 (2025), 329–346. doi:10.1007/s10805-024-09539-2

  19. [27]

    David Byrne. 2022. A worked example of Braun and Clarke’s approach to reflexive thematic analysis.Quality & Quantity56, 3 (2022), 1391–1412. doi:10.1007/s11135-021-01182-y

  20. [28]

    Scott Allen Cambo and Darren Gergle. 2022. Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data Science. InCHI Conference on Human Factors in Computing Systems. ACM, New Orleans LA USA, 1–19. doi:10.1145/3491102.3501998

  21. [29]

    Tricco, Zachary Munn, Danielle Pollock, Ashrita Saran, Anthea Sutton, Howard White, and Hanan Khalil

    Fiona Campbell, Andrea C. Tricco, Zachary Munn, Danielle Pollock, Ashrita Saran, Anthea Sutton, Howard White, and Hanan Khalil

  22. [30]

    Big Picture

    Mapping reviews, scoping reviews, and evidence and gap maps (EGMs): the same but different— the “Big Picture” review family. Systematic Reviews12, 1 (2023), 45. doi:10.1186/s13643-023-02178-5

  23. [31]

    Chmielinski, Sarah Newman, Matt Taylor, Josh Joseph, Kemi Thomas, Jessica Yurkofsky, and Yue Chelsea Qiu

    Kasia S. Chmielinski, Sarah Newman, Matt Taylor, Josh Joseph, Kemi Thomas, Jessica Yurkofsky, and Yue Chelsea Qiu. 2022. The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence. doi:10.48550/arXiv.2201.03954

  24. [32]

    Connolly, Daniel M

    Charlotte J. Connolly, Daniel M. Hueholt, and Melissa A. Burt. 2025. Datasheets for Earth Science Datasets.Bulletin of the American Meteorological Society106, 4 (April 2025), E642–E648. doi:10.1175/BAMS-D-24-0203.1

  25. [33]

    Payton Croskey, Fabian Offert, Jennifer Jacobs, and Kai M. Thaler. 2025. Liberatory Collections and Ethical AI: Reimagining AI Development from Black Community Archives and Datasets. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ...

  26. [34]

    Melissa Curran and Ashley K. Randall. 2021. Positionality Statements. https://onlinelibrary.wiley.com/pb-assets/assets/14756811/ Positionality-Statements.pdf

  27. [35]

    Davies and Constantin Holmer

    Sarah R. Davies and Constantin Holmer. 2024. Care, collaboration, and service in academic data work: biocuration as ‘academia otherwise’.Information, Communication & Society27, 4 (2024), 683–701. doi:10.1080/1369118X.2024.2315285

  28. [36]

    Suzanne Day. 2012. A Reflexive Lens: Exploring Dilemmas of Qualitative Methodology Through the Concept of Reflexivity.Qualitative Sociology Review8, 1 (2012), 60–85. doi:10.18778/1733-8077.8.1.04

  29. [37]

    Siddharth Peter De Souza and Linnet Taylor. 2025. Rebooting the global consensus: Norm entrepreneurship, data governance and the inalienability of digital bodies.Big Data & Society12, 2 (2025), 20539517251330191. doi:10.1177/20539517251330191

  30. [38]

    Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. 2023. Ob...

  31. [39]

    Melissa Dell, Jacob Carlson, Tom Bryan, Emily Silcock, Abhishek Arora, Zejiang Shen, Luca D’Amico-Wong, Quan Le, Pablo Querubin, and Leander Heldring. 2023. American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers. Advances in Neural Information Pr...

  32. [40]

    Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. 2021. On the genealogy of machine learning datasets: A critical history of ImageNet.Big Data & Society8, 2 (2021), 20539517211035955. doi:10.1177/20539517211035955

  33. [41]

    Nadine Desrochers, Adèle Paul-Hus, and Jen Pecoskie. 2017. Five decades of gratitude: A meta-synthesis of acknowledgments research. Journal of the Association for Information Science and Technology68, 12 (2017), 2821–2833. doi:10.1002/asi.23903

  34. [42]

    Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruksachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Trans...

  35. [43]

    Nicholas Diakopoulos. 2016. Accountability in algorithmic decision making.Commun. ACM59, 2 (2016), 56–62. doi:10.1145/2844110

  36. [44]

    Catherine D’Ignazio and Lauren Klein. 2020. 7. Show Your Work. InData Feminism. https://data-feminism.mitpress.mit.edu/pub/ 0vgzaln4/release/3

  37. [45]

    Mary Dixon-Woods, Debbie Cavers, Shona Agarwal, Ellen Annandale, Antony Arthur, Janet Harvey, Ron Hsu, Savita Katbamna, Richard Olsen, Lucy Smith, Richard Riley, and Alex J. Sutton. 2006. Conducting a critical interpretive synthesis of the literature on Evaluating Structured D...

  38. [46]

    2020.Data Feminism

    Catherine D’Ignazio and Lauren Klein. 2020.Data Feminism. MIT Press

  39. [47]

    Catherine D’Ignazio and Lauren Klein. 2023. Introducing Data Feminism. InWomen’s Empowerment and Its Limits: Interdisciplinary and Transnational Perspectives Toward Sustainable Progress, Elisa Fornalé and Federica Cristani (Eds.). Springer International Publishing, 139–151. ht...

  40. [48]

    Mustafa Emirbayer and Matthew Desmond. 2012. Race and reflexivity.Ethnic and Racial Studies35, 4 (2012), 574–599. doi:10.1080/ 01419870.2011.606910

  41. [49]

    Kim V. L. England. 1994. Getting Personal: Reflexivity, Positionality, and Feminist Research.The Professional Geographer46, 1 (1994), 80–89. doi:10.1111/j.0033-0124.1994.00080.x

  42. [50]

    Linda Finlay. 2002. Negotiating the swamp: the opportunity and challenge of reflexivity in research practice.Qualitative Research2, 2 (2002), 209–230. doi:10.1177/146879410200200205

  43. [51]

    Kate Flemming. 2010. Synthesis of quantitative and qualitative research: an example using Critical Interpretive Synthesis.Journal of Advanced Nursing66, 1 (2010), 201–217. doi:10.1111/j.1365-2648.2009.05173.x

  44. [52]

    Louise Folkes. 2023. Moving beyond ‘shopping list’ positionality: Using kitchen table reflexivity and in/visible tools to develop reflexive qualitative research.Qualitative Research23, 5 (2023), 1301–1318. doi:10.1177/14687941221098922

  45. [53]

    Simon Frieder, Luca Pinchetti, Alexis Chevalier, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Christian Petersen, and Julius Berner. 2023. Mathematical Capabilities of ChatGPT. Advances in Neural Information Processing Systems

  46. [54]

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Muss...

  47. [55]

    Jingxian Gan and Yong Qi. 2021. Selection of the Optimal Number of Topics for LDA Topic Model—Taking Patent Policy Analysis as an Example.Entropy23, 10 (2021), 1301. doi:10.3390/e23101301

  48. [56]

    Anatol Garioud, Nicolas Gonthier, Loic Landrieu, Apolline De Wit, Marion Valette, Marc Poupée, Sebastien Giordano, and Boris Wattrelos. 2023. FLAIR : a Country-Scale Land Cover Semantic Segmentation Dataset From Multi-Source Optical Imagery. In Thirty-seventh Conference on Neu...

  49. [57]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford

  50. [58]

    arXiv:1803.09010 (2018)

    Datasheets for Datasets. arXiv:1803.09010 (2018). http://arxiv.org/abs/1803.09010 arXiv:1803.09010 [cs]

  51. [59]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford

  52. [60]

    ACM64, 12 (2021), 86–92

    Datasheets for datasets.Commun. ACM64, 12 (2021), 86–92. doi:10.1145/3458723

  53. [61]

    Mathew Gillings and Andrew Hardie. 2023. The interpretation of topic models for scholarly analysis: An evaluation and critique of current practice.Digital Scholarship in the Humanities38, 2 (2023), 530–543

  54. [62]

    Gonzalez Zelay and Carlos Vladimiro. 2019. Towards Explaining the Effects of Data Preprocessing on Machine Learning. In2019 IEEE 35th International Conference on Data Engineering (ICDE). 2086–2090. doi:10.1109/ICDE.2019.00245

  55. [63]

    David S. A. Guttormsen and Fiona Moore. 2023. ‘Thinking About How We Think’: Using Bourdieu’s Epistemic Reflexivity to Reduce Bias in International Business Research.Management International Review63, 4 (2023), 531–559. doi:10.1007/s11575-023-00507-3

  56. [64]

    Donna Haraway. 1988. Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective.Feminist Studies14, 3 (1988), 575. doi:10.2307/3178066

  57. [65]

    Strong Objectivity

    Sandra Harding. 1992. After the Neutrality Ideal: Science, Politics, and "Strong Objectivity".Social Research59, 3 (1992), 567–587. https://www.jstor.org/stable/40970706

  58. [66]

    Strong Objectivity?

    Sandra Harding. 1992. Rethinking Standpoint Epistemology: What Is "Strong Objectivity?".The Centennial Review36, 3 (1992), 437–470. https://www.jstor.org/stable/23739232

  59. [67]

    Sheikh Md Shakeel Hassan, Arthur Feeney, Akash Dhruv, Jihoon Kim, Youngjoon Suh, Jaiyoung Ryu, Yoonjin Won, and Aparna Chandramowlishwaran. 2023. BubbleML: A Multiphase Multiphysics Dataset and Benchmarks for Machine Learning. Advances in Neural Information Processing Systems

  60. [68]

    Heger, Liz B

    Amy K. Heger, Liz B. Marquis, Mihaela Vorvoreanu, Hanna Wallach, and Jennifer Wortman Vaughan. 2022. Understanding Machine Learning Practitioners’ Data Documentation Perceptions, Needs, Challenges, and Desiderata.Proceedings of the ACM on Human- Computer Interaction6, CSCW2 (2...

  61. [69]

    Simpkins, Fritz Gerald P

    Anna Heinke, LingLing Huang, Kyongmi U. Simpkins, Fritz Gerald P. Kalaw, Apoorva Karsolia, Kiratjit Singh, Sanjay Soundarajan, Camille Nebeker, Sally L. Baxter, Cecilia S. Lee, Aaron Y. Lee, Bhavesh Patel, and the AI-READI Consortium. 2025. Dataset Documentation for Responsibl...

  62. [70]

    David J. Hess. 2013. Neoliberalism and the History of STS Theory: Toward a Reflexive Sociology.Social Epistemology27, 2 (2013), 177–193. doi:10.1080/02691728.2013.793754 FAccT ’26, June 25–28, 2026, Montreal, QC, Canada Bhardwaj and Zogheib, et al

  63. [71]

    Sharlene Nagy Hesse-Bibber and Deborah Piatelli. 2012. The Feminist Practice of Holistic Reflexivity. InHandbook of Feminist Research: Theory and Praxis. SAGE Publications, Inc., 557–582. https://doi.org/10.4135/9781483384740.n27

  64. [72]

    Simon David Hirsbrunner, Michael Tebbe, and Claudia Müller-Birn. 2024. From critical technical practice to reflexive data science. Convergence: The International Journal of Research into New Media Technologies30, 1 (2024), 190–215. doi:10.1177/13548565221132243

  65. [73]

    Rink Hoekstra and Simine Vazire. 2021. Aspiring to greater intellectual humility in science.Nature Human Behaviour5, 12 (2021), 1602–1607. doi:10.1038/s41562-021-01203-8

  66. [74]

    Thibaut Horel, Lorenzo Masoero, Raj Agrawal, Daria Roithmayr, and Trevor Campbell. 2021. The CPD Data Set: Personnel, Use of Force, and Complaints in the Chicago Police Department. Advances in Neural Information Processing Systems

  67. [75]

    Rodrigo Hormazabal, Changyoung Park, Soonyoung Lee, Sehui Han, Yeonsik Jo, Jaewan Lee, Ahra Jo, Seung Hwan Kim, Jaegul Choo, Moontae Lee, and Honglak Lee. 2022. CEDe: A collection of expert-curated datasets with atom-level entity annotations for Optical Chemical Structure Reco...

  68. [76]

    Zhe Huang, Liang Wang, Giles Blaney, Christopher Slaughter, Devon McKeon, Ziyu Zhou, Robert Jacob, and Michael C. Hughes

  69. [77]

    Advances in Neural Information Processing Systems

    The Tufts fNIRS Mental Workload Dataset & Benchmark for Brain-Computer Interfaces that Generalize. Advances in Neural Information Processing Systems

  70. [78]

    Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell

  71. [79]

    InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21)

    Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21). Association for Computing Machinery, New York, NY, USA, 560–575. do...

  72. [80]

    Othering & Belonging Institute. 2025. Power Analysis. https://belonging.berkeley.edu/transformative-research-toolkit/power-analysis

  73. [81]

    Green, and Tariq Iqbal

    Md Mofijul Islam, Reza Manuel Mirzaiee, Alexi Gladstone, Haley N. Green, and Tariq Iqbal. 2022. CAESAR: An Embodied Simulator for Generating Multimodal Referring Expression Datasets. Advances in Neural Information Processing Systems

  74. [82]

    Danielle Jacobson and Nida Mustafa. 2019. Social Identity Map: A Reflexivity Tool for Practicing Explicit Positionality in Critical Qualitative Research.International Journal of Qualitative Methods18 (2019), 1609406919870075. doi:10.1177/1609406919870075

  75. [83]

    Jamieson, Gisela H

    Michelle K. Jamieson, Gisela H. Govaart, and Madeleine Pownall. 2023. Reflexivity in quantitative research: A rationale and beginner’s guide.Social and Personality Psychology Compass17, 4 (2023), e12735. doi:10.1111/spc3.12735

  76. [84]

    Eun Seo Jo and Timnit Gebru. 2020. Lessons from archives: strategies for collecting sociocultural data in machine learning. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency. ACM, Barcelona Spain, 306–316. doi:10.1145/3351095.3372829

  77. [85]

    Fanny Jourdan, Yannick Chevalier, and Cécile Favre. 2025. FairTranslate: an English-French Dataset for Gender Bias Evaluation in Machine Translation by Overcoming Gender Binarity. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25...

  78. [86]

    Gaoussou Youssouf Kebe, Padraig Higgins, Patrick Jenkins, Kasra Darvish, Rishabh Sachdeva, Ryan Barron, John Winder, Donald Engel, Edward Raff, Francis Ferraro, and Cynthia Matuszek. 2021. A Spoken Language Dataset of Descriptions for Speech-Based Grounded Language Learning. A...

  79. [87]

    Jane Kenway and Julie McLeod. 2004. Bourdieu’s reflexive sociology and ‘spaces of points of view’: whose reflexivity, which perspective? British Journal of Sociology of Education25, 4 (2004), 525–544. doi:10.1080/0142569042000236998 Publisher: Routledge

  80. [88]

    David Kerr and David C. Klonoff. 2019. Digital Diabetes Data and Artificial Intelligence: A Time for Humility Not Hubris.Journal of Diabetes Science and Technology13, 1 (2019), 123–127. doi:10.1177/1932296818796508

  81. [89]

    Mehtab Khan and Alex Hanna. 2022. The Subjects and Stages of AI Dataset Development: A Framework for Dataset Accountability. (2022). doi:10.2139/ssrn.4217148

  82. [90]

    2021.Beyond Interdisciplinarity: Boundary Work, Communication, and Collaboration

    Julie Thompson Klein. 2021.Beyond Interdisciplinarity: Boundary Work, Communication, and Collaboration. Oxford University Press

  83. [91]

    Angelie Kraft and Eloïse Soulier. 2024. Knowledge-Enhanced Language Models Are Not Bias-Proof: Situated Knowledge and Epistemic Injustice in AI. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’24). Association for Computing Machin...

  84. [92]

    Sneha Kudugunta, Isaac Rayburn Caswell, Biao Zhang, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. MADLAD-400: A Multilingual And Document-Level Large Audited Dataset. Advances in Neural Information Processing Systems

  85. [93]

    Jiyoung Lee, Seungho Kim, Seunghyun Won, Joonseok Lee, Marzyeh Ghassemi, James Thorne, Jaeseok Choi, O.-Kil Kwon, and Edward Choi. 2023. VisAlign: Dataset for Measuring the Alignment between AI and Humans in Visual Perception. Advances in Neural Information Processing Systems

  86. [94]

    Jingjin Li, Qisheng Li, Rong Gong, Lezhi Wang, and Shaomei Wu. 2025. Our Collective Voices: The Social and Technical Values of a Grassroots Chinese Stuttered Speech Dataset. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Ass...

  87. [95]

    Calvin Liang. 2021. Reflexivity, positionality, and disclosure in HCI. https://medium.com/@caliang/reflexivity-positionality-and- disclosure-in-hci-3d95007e9916 Evaluating Structured Documentation as a Tool for Reflexivity FAccT ’26, June 25–28, 2026, Montreal, QC, Canada

  88. [96]

    Zhaojiang Lin, Andrea Madotto, Genta Indra Winata, Peng Xu, Feijun Jiang, Yuxiang Hu, Chen Shi, and Pascale Fung. 2021. BiToD: A Bilingual Multi-Domain Dataset For Task-Oriented Dialogue Modeling. Advances in Neural Information Processing Systems

  89. [97]

    Jasmine R Linabary, Danielle J Corple, and Cheryl Cooky. 2021. Of wine and whiteboards: Enacting feminist reflexivity in collaborative research.Qualitative Research21, 5 (2021), 719–735. doi:10.1177/1468794120946988

  90. [98]

    Climate change

    Ming Liu and Jingyi Huang. 2022. “Climate change” vs. “global warming”: A corpus-assisted discourse analysis of two popular terms in The New York Times.Journal of World Languages8, 1 (2022), 34–55. doi:10.1515/jwl-2022-0004 Publisher: De Gruyter Mouton

  91. [99]

    Utkarsh Mall, Bharath Hariharan, and Kavita Bala. 2022. Change Event Dataset for Discovery from Spatio-temporal Remote Sensing Imagery. Advances in Neural Information Processing Systems

  92. [100]

    Jiayuan Mao, Xuelin Yang, Xikun Zhang, Noah Goodman, and Jiajun Wu. 2022. CLEVRER-Humans: Describing Physical and Causal Events the Human Way. Advances in Neural Information Processing Systems

  93. [101]

    Karl Maton. 2003. Reflexivity, Relationism, & Research: Pierre Bourdieu and the Epistemic Conditions of Social Scientific Knowledge. Space and Culture6, 1 (2003), 52–65. doi:10.1177/1206331202238962

  94. [102]

    William Mattingly. 2021. Introduction to Topic Modeling and Text Classification. https://topic-modeling.pythonhumanities.com/intro. html

  95. [103]

    McCorkel and Kristen Myers

    Jill A. McCorkel and Kristen Myers. 2003. What Difference Does Difference Make? Position and Privilege in the Field.Qualitative Sociology26, 2 (2003), 199–231. doi:10.1023/A:1022967012774

  96. [104]

    Hill, Javier Hernandez, Jonathan Lester, and Tadas Baltrusaitis

    Daniel McDuff, Miah Wander, Xin Liu, Brian L. Hill, Javier Hernandez, Jonathan Lester, and Tadas Baltrusaitis. 2022. SCAMPS: Synthetics for Camera Measurement of Physiological Signals. Advances in Neural Information Processing Systems

  97. [105]

    Katrina Skewes McFerran, Sandra Garrido, and Suvi Saarikallio. 2016. A Critical Interpretive Synthesis of the Literature Linking Music and Adolescent Mental Health.Youth & Society48, 4 (2016), 521–538. doi:10.1177/0044118X13501343

  98. [106]

    Bender, and Batya Friedman

    Angelina McMillan-Major, Emily M. Bender, and Batya Friedman. 2024. Data Statements: From Technical Concept to Community Practice.ACM J. Responsib. Comput.1, 1 (2024), 1:1–1:17. doi:10.1145/3594737

  99. [107]

    Milagros Miceli, Tianling Yang, Laurens Naudts, Martin Schuessler, Diana Serbanescu, and Alex Hanna. 2021. Documenting Computer Vision Datasets: An Invitation to Reflexive Data Practices. InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (F...

  100. [108]

    Nafise Sadat Moosavi, Andreas Rücklé, Dan Roth, and Iryna Gurevych. 2021. SciGen: a Dataset for Reasoning-Aware Text Generation from Scientific Tables. Advances in Neural Information Processing Systems

  101. [109]

    Vera Liao, Casey Dugan, and Thomas Erickson

    Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q. Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How Data Science Workers Work with Data: Discovery, Capture, Curation, Design, Creation. InProceedings of the 2019 CHI Conference on Human Factors in C...

  102. [110]

    Michael Muller and Angelika Strohmayer. 2022. Forgetting Practices in the Data Sciences. InCHI Conference on Human Factors in Computing Systems. ACM, New Orleans LA USA, 1–19. doi:10.1145/3491102.3517644

  103. [111]

    Gokul Nc, Manideep Ladi, Sumit Negi, Prem Selvaraj, Pratyush Kumar, and Mitesh M. Khapra. 2022. Addressing Resource Scarcity across Sign Languages with Multilingual Pretraining and Unified-Vocabulary Datasets. Advances in Neural Information Processing Systems

  104. [112]

    Noblit and R

    George W. Noblit and R. Dwight Hare. 1988.Meta-Ethnography: Synthesizing Qualitative Studies. SAGE Publications

  105. [113]

    Yasumasa Onoe, Michael JQ Zhang, Eunsol Choi, and Greg Durrett. 2021. CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge. Advances in Neural Information Processing Systems

  106. [114]

    Cliodhna O’Connor and Helene Joffe. 2020. Intercoder reliability in qualitative research: Debates and practical guidelines.International journal of qualitative methods19 (2020), 1609406919899220

  107. [115]

    Dongwei Pan, Long Zhuo, Jingtan Piao, Huiwen Luo, Wei Cheng, Yuxin Wang, Siming Fan, Shengqi Liu, Lei Yang, Bo Dai, Ziwei Liu, Chen Change Loy, Chen Qian, Wayne Wu, Dahua Lin, and Kwan-Yee Lin. 2023. RenderMe-360: A Large Digital Asset Library and Benchmarks Towards High-fidel...

  108. [116]

    Orestis Papakyriakopoulos, Anna Seo Gyeong Choi, William Thong, Dora Zhao, Jerone Andrews, Rebecca Bourke, Alice Xiang, and Allison Koenecke. 2023. Augmented Datasheets for Speech Datasets and Ethical Decision-Making. InProceedings of the 2023 ACM Conference on Fairness, Accou...

  109. [117]

    Samir Passi and Solon Barocas. 2019. Problem Formulation and Fairness. InProceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). Association for Computing Machinery, New York, NY, USA, 39–48. doi:10.1145/3287560.3287567

  110. [118]

    Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vanderplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesna...

  111. [119]

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web FAccT ’26, June 25–28, 2026...

  112. [120]

    Wanda Pillow. 2003. Confession, catharsis, or cure? Rethinking the uses of reflexivity as methodological power in qualitative research. International Journal of Qualitative Studies in Education16, 2 (2003), 175–196. doi:10.1080/0951839032000060635

  113. [121]

    Lindsay Poirier. 2022. Accountable Data: The Politics and Pragmatics of Disclosure Datasets. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machinery, New York, NY, USA, 1446–1456. doi:10.1145/35311...

  114. [122]

    Lindsay Poirier, Juniper Huang, and Casey MacGibbon. 2025. What Remains Opaque in Transparency Initiatives: Visualizing Phantom Reductions through Devious Data Analysis. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Associa...

  115. [123]

    Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. 2022. Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machin...

  116. [124]

    White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes

    Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI accountability gap: defining an end-to-end framework for internal algorithmic auditing. InProceedin...

  117. [125]

    Chirag Raman, Jose Vargas Quiros, Stephanie Tan, Ashraful Islam, Ekin Gedik, and Hayley Hung. 2022. ConfLab: A Data Collection Concept, Dataset, and Benchmark for Machine Analysis of Free-Standing Social Interactions in the Wild. Advances in Neural Information Processing Systems

  118. [126]

    Fernández-Giménez, and Elisa Oteros-Rozas

    Federica Ravera, Maria E. Fernández-Giménez, and Elisa Oteros-Rozas. 2023. Reflexivity, embodiment, and ethics of care in rangeland political ecology: reflections of three feminist researchers on the experience of transdisciplinary knowledge co-production.Frontiers in Human Dy...

  119. [127]

    Fábio Ribeiro and Juliana Miraldi. 2022. Bourdieu, Reflexivity, and Scientific Practice.Configurações. Revista Ciências Sociais29 (2022), 111–130. doi:10.4000/configuracoes.15157

  120. [128]

    Mohammad Rashidujjaman Rifat, Abdullah Hasan Safir, Sourav Saha, Jahedul Alam Junaed, Maryam Saleki, Mohammad Ruhul Amin, and Syed Ishtiaque Ahmed. 2024. Data, Annotation, and Meaning-Making: The Politics of Categorization in Annotating a Dataset of Faith-based Communal Violen...

  121. [129]

    Negar Rostamzadeh, Diana Mincu, Subhrajit Roy, Andrew Smart, Lauren Wilcox, Mahima Pushkarna, Jessica Schrouff, Razvan Amironesei, Nyalleng Moorosi, and Katherine Heller. 2022. Healthsheet: Development of a Transparency Artifact for Health Datasets. InProceedings of the 2022 A...

  122. [130]

    Everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. InProceedings of the 2021 CHI Conference on Human Factors in Computing System...

  123. [131]

    Kate Sanders, Reno Kriz, Anqi Liu, and Benjamin Van Durme. 2022. Ambiguous Images With Human Judgments for Robust Visual Event Classification. Advances in Neural Information Processing Systems

  124. [132]

    Sanja Scepanovic, Ivica Obadic, Sagar Joglekar, Laura Giustarini, Cristiano Nattero, Daniele Quercia, and Xiao Xiang Zhu. 2023. MedSat: A Public Health Dataset for England Featuring Medical Prescriptions and Satellite Imagery. InThirty-seventh Conference on Neural Information ...

  125. [133]

    Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021. Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development.Proceedings of the ACM on Human-Computer Interaction5, CSCW2 (2021), 1–37. doi:10.1145/3476058

  126. [134]

    Benjamin M Schmidt. 2012. Words alone: Dismantling topic models in the humanities.Journal of Digital Humanities2, 1 (2012), 49–65

  127. [135]

    Patrick Schramowski, Christopher Tauchmann, and Kristian Kersting. 2022. Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’2...

  128. [136]

    Hope Schroeder, Akshansh Pareek, and Solon Barocas. 2025. Disclosure without Engagement: An Empirical Review of Positionality Statements at FAccT. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Mach...

  129. [137]

    Tal Schuster, Ashwin Kalyan, Alex Polozov, and Adam Tauman Kalai. 2021. Programming Puzzles. Advances in Neural Information Processing Systems

  130. [138]

    Nick Seaver. 2019. Knowing Algorithms. InDigitalSTS: a field guide for science & technology studies, Janet Vertesi, David Ribes, Carl DiSalvo, Laura Forlano, Steven J. Jackson, Yanni Alexander Loukissas, Daniela K. Rosner, and Hanna Rose Shell (Eds.). Princeton University Pres...

  131. [139]

    Masters, Matilde L

    Stephen Secules, Cassandra McCall, Joel Alejandro Mejia, Chanel Beebe, Adam S. Masters, Matilde L. Sánchez-Peña, and Martina Svyantek. 2021. Positionality practices and dimensions of impact on equity research: A collaborative inquiry and call to the community. Journal of Engin...

  132. [140]

    Raesetje Sefala, Timnit Gebru, Luzango Mfupe, Nyalleng Moorosi, and Richard Klein. 2021. Constructing a Visual Dataset to Study the Effects of Spatial Apartheid in South Africa. Advances in Neural Information Processing Systems

  133. [141]

    Omar Shouman, Wassim Gabriel, Victor-George Giurcoiu, Vitor Sternlicht, and Mathias Wilhelm. 2022. PROSPECT: Labeled Tandem Mass Spectrometry Dataset for Machine Learning in Proteomics. Advances in Neural Information Processing Systems

  134. [142]

    Jan Simson, Alessandro Fabris, and Christoph Kern. 2024. Lazy Data Practices Harm Fairness Research. InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’24). Association for Computing Machinery, New York, NY, USA, 642–659. doi:10.1145...

  135. [143]

    Ought-Is

    Bryan A. Sisk, Jessica Mozersky, Alison L. Antes, and James M. DuBois. 2020. The “Ought-Is” Problem: An Implementation Science Framework for Translating Ethical Norms Into Practice.The American Journal of Bioethics20, 4 (2020), 62–70. doi:10.1080/15265161. 2020.1730483

  136. [144]

    Jessica Soedirgo and Aarie Glas. 2020. Toward Active Reflexivity: Positionality and Practice in the Production of Knowledge.PS: Political Science & Politics53, 3 (2020), 527–531. doi:10.1017/S1049096519002233

  137. [145]

    Moein Sorkhei, Yue Liu, Hossein Azizpour, Edward Azavedo, Karin Dembrower, Dimitra Ntoula, Athanasios Zouzos, Fredrik Strand, and Kevin Smith. 2021. CSAW-M: An Ordinal Classification Dataset for Benchmarking Mammographic Masking of Cancer. Advances in Neural Information Proces...

  138. [146]

    Lucy Suchman. 2002. Located accountabilities in technology production.Scandinavian Journal of Information Systems14, 2 (2002), 7

  139. [147]

    Harini Suresh and John V. Guttag. 2021. A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. InEquity and Access in Algorithms, Mechanisms, and Optimization. 1–9. doi:10.1145/3465416.3483305

  140. [148]

    Zhiyuan Tang, Dong Wang, Yanguang Xu, Jianwei Sun, Xiaoning Lei, Shuaijiang Zhao, Cheng Wen, Xingjun Tan, Chuandong Xie, Shuran Zhou, Rui Yan, Chenjia Lv, Yang Han, Wei Zou, and Xiangang Li. 2021. KeSpeech: An Open Source Speech Dataset of Mandarin and Its Eight Subdialects. I...

  141. [149]

    Anissa Tanweer, Emily Kalah Gade, P. M. Krafft, and Sarah Dreier. 2021. Why the Data Revolution Needs Qualitative Thinking.Harvard Data Science Review3, 3 (2021). doi:10.1162/99608f92.eee0b0da

  142. [150]

    Thomer, Dharma Akmon, Jeremy J

    Andrea K. Thomer, Dharma Akmon, Jeremy J. York, Allison R. B. Tyler, Faye Polasek, Sara Lafia, Libby Hemphill, and Elizabeth Yakel

  143. [151]

    doi:10.1145/3555139

    The Craft and Coordination of Data Curation: Complicating Workflow Views of Data Science.Proceedings of the ACM on Human-Computer Interaction6, CSCW2 (2022), 414:1–414:29. doi:10.1145/3555139

  144. [152]

    Nanna Bonde Thylstrup. 2022. The ethics and politics of data sets in the age of machine learning: deleting traces and encountering remains.Media, Culture & Society44, 4 (2022), 655–671. doi:10.1177/01634437211060226

  145. [153]

    Michael Veale, Max Van Kleek, and Reuben Binns. 2018. Fairness and Accountability Design Needs for Algorithmic Support in High-Stakes Public Sector Decision-Making. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM, Montreal QC Canada, 1–14. d...

  146. [154]

    Junjue Wang, Zhuo Zheng, Ailong Ma, Xiaoyan Lu, and Yanfei Zhong. 2021. LoveDA: A Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation. Advances in Neural Information Processing Systems

  147. [155]

    Tenenbaum, and Chuang Gan

    Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B. Tenenbaum, and Chuang Gan. 2021. STAR: A Benchmark for Situated Reasoning in Real-World Videos. Advances in Neural Information Processing Systems

  148. [156]

    Ch, Andreas Grammenos, Jing Han, Apinan Hasthanasombat, Erika Bondareva, Ting Dang, Andres Floto, Pietro Cicuta, and Cecilia Mascolo

    Tong Xia, Dimitris Spathis, Chlo{\"e} Brown, J. Ch, Andreas Grammenos, Jing Han, Apinan Hasthanasombat, Erika Bondareva, Ting Dang, Andres Floto, Pietro Cicuta, and Cecilia Mascolo. 2021. COVID-19 Sounds: A Large-Scale Audio Dataset for Digital Respiratory Screening. Advances ...

  149. [157]

    Juan Xie, Qing Ke, Ying Cheng, and Nancy Everhart. 2020. Meta-synthesis in Library & Information Science Research.The Journal of Academic Librarianship46, 5 (2020), 102217. doi:10.1016/j.acalib.2020.102217

  150. [158]

    Xinyu Yang, Weixin Liang, and James Zou. 2023. Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFace. https://openreview.net/forum?id=xC8xh2RSs2

  151. [159]

    2008.Five Faces of Oppression

    Iris Marion Young. 2008.Five Faces of Oppression. Routledge, 55–71

  152. [160]

    Hard-to-Reach

    Radim Řehůřek and Petr Sojka. 2010.Software Framework for Topic Modelling with Large Corpora. University of Malta. https: //repozitar.cz/publication/15725/cs/Software-Framework-for-Topic-Modelling-with-Large-Corpora/Rehurek-Sojka FAccT ’26, June 25–28, 2026, Montreal, QC, Cana...

Pith tools

Reviewed May 13, 2026 · model on record in the stance chip above.