REVIEW 2 major objections 4 minor 2 cited by
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Public instruments for measuring LLM representational harms often go unused by practitioners.
desk verdict A genuinely useful qualitative study of why practitioners don't use public representational-harm instruments, but the abstract overstates prevalence in a way the authors' own limitations section disavows. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study is carried by semi-structured interviews with 12 practitioners, scaffolded by a set of seven instrument desiderata: validity, reliability, specificity, extensibility, scalability, interpretability, and actionability. The interview guide prompts participants to discuss each desideratum and to describe any additional challenges. The analysis then maps reported challenges onto the "not useful" versus "not used" distinction, and the authors use measurement theory and pragmatic measurement as the interpretive lens for their recommendations, particularly the ideas that instruments should document their systematized concept and that extensible, modular instruments can be adapted by practitioners.
What would settle it
A random-sample survey of practitioners with a high response rate, or an audit of actual instrument usage logs across organizations, that found most practitioners routinely use public instruments as-is and rate them valid, specific, and actionable, would contradict the paper's central claim.
Extended reading notes
Core claim
The paper's central claim is that the public measurement instruments for representational harms are frequently unusable in practice, and that the reasons fall into two classes. In one class, instruments are not useful: practitioners reported that concepts are often undefined or detached from theory, that datasets contain mislabeled examples, that benchmarks may be contaminated by training data, that instruments are not specific to a system's context, and that outputs cannot be interpreted or acted upon. In the other class, instruments are not used even when potentially useful, because organizations impose security, data-licensing, time, or incentive constraints. The authors report that every participant discussed challenges that prevented use, that validity and specificity were the primary considerations, and that practitioners often responded by building their own bespoke instruments.
Load-bearing premise
The load-bearing premise is that the 12 recruited practitioners, most reached through the authors' networks and snowball sampling, are telling the truth about their measurement practices and are not just echoing the challenges the interview guide put in front of them.
Editorial extensions
If this is right
- Instrument designers should document the systematized concept behind each instrument, separating definition from operationalization, so practitioners can judge what is being measured.
- Public instruments should ship with interpretability aids, such as distributions of scores on known datasets and guidance on what a given score means.
- Designing instruments to be open, modular, and extensible can help practitioners adapt them to local contexts while preserving validity and reliability.
- Organizations and regulators can remove practical and institutional barriers, such as security constraints and missing incentives, rather than leaving uptake to designers alone.
- If instruments remain unused, practitioners will continue to build bespoke instruments, making measurements harder to compare across organizations.
Reading between the lines
- A natural extension of the paper's logic is that public evaluation artifacts in NLP more broadly may need to be assessed for uptake, not just for technical performance.
- The paper's interviews were not designed to estimate prevalence, so a large survey of practitioners would be a reasonable next step to test how widespread the reported barriers are.
- Because the authors note that participants saw their difficulties as extending to other abstract or contested concepts, the "not useful, not used" split may generalize well beyond representational harms.
- Practitioners measuring harms in low-resource languages likely face even steeper shortages of usable instruments, a gap the paper only touches on.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a semi-structured interview study with 12 practitioners who evaluate LLM-based systems, focusing on their experiences using publicly available instruments for measuring representational harms. The authors find that practitioners in their sample faced two broad types of challenges: instruments are sometimes not useful because they lack validity, specificity, interpretability, or actionability for the practitioner's context, and even useful instruments are sometimes not used because of practical and institutional barriers such as security requirements, data licensing, and organizational culture. Drawing on measurement theory and pragmatic measurement, the paper offers recommendations for instrument designers, including systematic concept systematization, providing interpretability resources, and designing extensible instruments. The paper includes a detailed interview guide, a systematic literature review of desiderata, a positionality statement, and an explicit limitations section.
Significance. If the findings are interpreted as a qualitative characterization of challenges rather than a prevalence estimate, the paper makes a useful contribution to the responsible AI and NLP evaluation literature. Its distinction between instruments that are not useful and instruments that are useful but not used is a valuable framing that connects practitioner experience to measurement theory. The paper also gives concrete, actionable recommendations grounded in established measurement frameworks, and it is unusually transparent about recruitment difficulties, coding procedures, and study limitations. The interview guide and PRISMA-based literature review are strengths that will facilitate replication and extension. However, the headline claim in the abstract and conclusion goes beyond what the sampling design can support, and the interview protocol's prompting of the seven a priori desiderata weakens the claim that those desiderata emerged from practitioners' own accounts.
major comments (2)
- [Abstract; §1; §6; Limitations] The abstract, introduction, and conclusion claim that practitioners are 'often' unable to use publicly available instruments for measuring representational harms. This is a prevalence claim, but the Limitations section explicitly states that the participant pool is 'likely skewed toward practitioners who faced challenges' and that the findings 'do not enable us to answer questions about the prevalence of the challenges.' A self-selected, snowball-recruited sample of 12 practitioners cannot support the quantifier 'often.' I recommend rewording the central claim to an existence claim about the challenges practitioners can face (for example, 'practitioners in our sample reported being unable to use...' or 'practitioners can face challenges that leave them unable to use...'), and removing or qualifying 'often' in the abstract, §1, and §6.
- [§3 Interview scaffolding; §4 Desiderata alignment; Appendix B Q20] The claim in §4 that 'the desiderata identified in Table 3 aligned closely with considerations participants described' and that 'participants did not report considering any additional desiderata' is partly an artifact of the interview protocol. Appendix B Q20 explicitly asks participants to confirm or deny each of the seven a priori desiderata, so the observed alignment is partly by construction. The paper should either analyze unprompted mentions separately from prompted confirmations, or substantially soften the claim to say that participants affirmed these desiderata when asked. As written, the inductive–deductive coding description (§3) does not fully address the circularity introduced by the scaffolding.
minor comments (4)
- [Table 2; §3] Table 2 uses participant IDs P01–P12 while the text (e.g., §3 and §4) refers to P1–P12; please unify the notation.
- [§4.3] The sentence beginning 'Finally, we note that multiple participants lived in countries where English is not the primary language...' introduces a topic that is not strictly a challenge 'exacerbated by, but extend[ing] beyond, representational harms'; consider moving it to the Limitations or presenting it as a separate observation.
- [Appendix A.1.2] The annotation strategy states that the desiderata list was developed before mapping passages and that no additional desiderata emerged; this is a limitation of the systematic review that should be acknowledged in the main text alongside the interview scaffolding limitation.
- [Table 1] The row labeled 'Other' contains the phrase 'Matched guise probing'; for consistency with the other entries, consider using a full sentence or consistent formatting.
Circularity Check
The desiderata-alignment finding is partly built into the interview protocol; the central two-type taxonomy retains independent evidence, so the paper is only partially circular.
-
self definitional
[§3 Methods (Interviews and Thematic analysis), §4 opening, Appendix B Q20]
"To scaffold the interviews, we identified a set of desiderata ... validity, reliability, specificity, extensibility, scalability, interpretability, and actionability. ... For each of the challenges defined below, say either: 'It sounds like you mentioned an issue to do with [challenge]. Is that correct?', or 'I don't think you mentioned [challenge]. Did you experience any issues with this?' ... We found that the desiderata identified in Table 3 aligned closely with considerations participants described as being central to their decisions about the usefulness of measurement instruments."
The desiderata were fixed before data collection, and Q20 asked participants to confirm or deny experiencing each one; the thematic analysis then started from the same desiderata as the initial code set. The reported alignment between the a priori desiderata and practitioners' considerations is therefore substantially an artifact of the elicitation and coding protocol, not an independent discovery. The open-ended questions did yield one genuinely emergent category (practical and institutional barriers), and participants could decline prompts, so the central two-type taxonomy is not wholly circular; but the seven-desiderata alignment and the claim that all participants considered validity and specificity are to a large degree constructed by the protocol.
full rationale
The paper's core finding—that practitioners face both usefulness and uptake challenges—does not reduce entirely to its inputs: participants' concrete accounts (e.g., data contamination, licensing/security, organizational disincentives) provide independent content, and the limitations section honestly disclaims prevalence claims. However, the specific claim that the authors' seven pre-defined desiderata 'aligned closely' with practitioner considerations is circular in part, because the interview guide explicitly prompted participants about each desideratum and the initial coding scheme used those same desiderata. I do not count the 'often' prevalence wording as circularity: it is an overgeneralization/sampling-validity concern flagged by the authors' own Limitations section. No load-bearing self-citation chain or imported uniqueness theorem is present, so this is a partial, protocol-driven circularity rather than a fully constructed one.
Assumptions & free parameters
assumptions (4)
- domain assumption Saturation after multiple consecutive interviews justifies treating the identified challenge set as complete for thematic purposes.
- domain assumption Participants' self-reports during one-hour interviews accurately capture their measurement practices and constraints.
- domain assumption Prompting participants with the seven a priori desiderata does not materially induce the challenges they report.
- domain assumption The seven desiderata identified from the systematic review and author experience are complete.
Cite this review
Pith. "Pith review of Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems." pith.science (2026). https://pith.science/paper/O4ICOBPR
@misc{pith2026250604482,
author = {Pith},
title = {Pith review of: Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/O4ICOBPR}},
note = {Machine review of arXiv:2506.04482}
}
read the original abstract
The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instruments meet the needs of practitioners tasked with evaluating LLM-based systems. Via semi-structured interviews with 12 such practitioners, we find that practitioners are often unable to use publicly available instruments for measuring representational harms. We identify two types of challenges. In some cases, instruments are not useful because they do not meaningfully measure what practitioners seek to measure or are otherwise misaligned with practitioner needs. In other cases, instruments - even useful instruments - are not used by practitioners due to practical and institutional barriers impeding their uptake. Drawing on measurement theory and pragmatic measurement, we provide recommendations for addressing these challenges to better meet practitioner needs.
Figures
Forward citations
Cited by 2 Pith papers
-
Grounded Chess Reasoning in Language Models via Master Distillation
Master Distillation turns opaque chess-engine search into natural-language CoT, lifting a 4B model (C1) from near-zero to 48.1% accuracy that beats its teacher and most larger LMs.
-
Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictabi...
Reference graph
Works this paper leans on
-
[1]
Robert Adcock and David Collier. 2001. https://doi.org/10.1017/S0003055401003100 Measurement Validity : A Shared Standard for Qualitative and Quantitative Research . American Political Science Review, 95(3):529--546
-
[2]
Maria Antoniak and David Mimno. 2021. https://doi.org/10.18653/v1/2021.acl-long.148 Bad seeds: Evaluating lexical methods for bias measurement . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1889--1904, Onl...
-
[3]
Fairness Toolkits , A Checkbox Culture ?
Agathe Balayn, Mireia Yurrita, Jie Yang, and Ujwal Gadiraju. 2023. https://doi.org/10.1145/3600211.3604674 “ Fairness Toolkits , A Checkbox Culture ?” On the Factors that Fragment Developer Practices in Handling Algorithmic Harms . In Proceedings of the 2023 AAAI / ACM Conference on AI , Ethics , and Society , AIES , pages 482--495, Montreal QC Canada. ACM
arXiv 2023
-
[4]
Carliss Y. Baldwin and Kim B. Clark. 2000. https://doi.org/10.7551/mitpress/2366.001.0001 Design Rules, Volume 1: The Power of Modularity . The MIT Press
-
[5]
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: Allocative versus representational harms in machine learning. In Proceedings of SIGCIS, Philadelphia, PA
work page 2017
-
[6]
Emily Bender. 2019. https://thegradient.pub/the-benderrule-on-naming-the-languages-we-study-and-why-it-matters/ The \#benderrule: On naming the languages we study and why it matters . The Gradient
work page 2019
-
[7]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of bias in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454--5476, Online. Association for Computational Linguistics
-
[8]
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. https://doi.org/10.18653/v1/2021.acl-long.81 Stereotyping N orwegian salmon: An inventory of pitfalls in fairness benchmark datasets . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferenc...
Show all 76 references
-
[9]
Rishi Bommasani and Percy Liang. 2024. Trustworthy social bias measurement. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 210--224
2024
-
[10]
Virginia Braun and Victoria Clarke. 2006. https://doi.org/10.1191/1478088706qp063oa Using thematic analysis in psychology . Qualitative Research in Psychology, 3(2):77--101
2006 doi
-
[11]
Virginia Braun and Victoria Clarke. 2019. https://doi.org/10.1080/2159676X.2019.1628806 Reflecting on reflexive thematic analysis . Qualitative Research in Sport, Exercise and Health, 11(4):589--597
2019
-
[12]
Kelly Caine. 2016. https://doi.org/10.1145/2858036.2858498 Local Standards for Sample Size at CHI . In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems , pages 981--992, San Jose California USA. ACM
2016
-
[13]
Bryson, and Arvind Narayanan
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. https://doi.org/10.1126/science.aal4230 Semantics derived automatically from language corpora contain human-like biases . Science, 356(6334):183--186
2017 doi
-
[14]
Yang Trista Cao, Yada Pruksachatkun, Kai-Wei Chang, Rahul Gupta, Varun Kumar, Jwala Dhamala, and Aram Galstyan. 2022. https://doi.org/10.18653/v1/2022.acl-short.62 On the intrinsic and extrinsic fairness evaluation metrics for contextualized language representations . In Proce...
2022 doi
-
[15]
Tommaso Caselli, Valerio Basile, Jelena Mitrovi \'c , and Michael Granitzer. 2021. https://doi.org/10.18653/v1/2021.woah-1.3 H ate BERT : Retraining BERT for abusive language detection in E nglish . In Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021), page...
2021 doi
-
[16]
Kate Crawford. 2017. The trouble with bias. In Conference on Neural Information Processing Systems (invited speaker)
2017
-
[17]
Pieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett, and Zeerak Talat. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.1207 Metrics for what, metrics for whom: Assessing actionability of bias evaluation metrics in NLP . In Proceedings of the 2024 Conference o...
2024 doi
-
[18]
Pieter Delobelle, Ewoenam Tokpo, Toon Calders, and Bettina Berendt. 2022. https://doi.org/10.18653/v1/2022.naacl-main.122 Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models . In Proceedings of the 2022 Conference of the N...
2022 doi
-
[19]
Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. https://doi.org/10.1145/3531146.3533113 Exploring How Machine Learning Practitioners ( Try To ) Use Fairness Toolkits . In 2022 ACM Conference o...
2022
-
[20]
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.150 Harms of gender exclusivity and challenges in non-binary representation in language technologies . In Proceedings of the ...
2021 doi
-
[21]
Sunipa Dev, Emily Sheng, Jieyu Zhao, Aubrie Amstutz, Jiao Sun, Yu Hou, Mattie Sanseverino, Jiin Kim, Akihiro Nishi, Nanyun Peng, and Kai-Wei Chang. 2022. https://doi.org/10.18653/v1/2022.findings-aacl.24 On measures of biases and harms in NLP . In Findings of the Association f...
2022 doi
-
[23]
Yupei Du, Qixiang Fang, and Dong Nguyen. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.785 Assessing the reliability of word embedding gender bias measures . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10012--10034, Onli...
2021 doi
-
[24]
Estabrooks, Maureen Boyle, Karen M
Paul A. Estabrooks, Maureen Boyle, Karen M. Emmons, Russell E. Glasgow, Bradford W. Hesse, Robert M. Kaplan, Alexander H. Krist, Richard P. Moser, and Martina V. Taylor. 2012. https://doi.org/10.1136/amiajnl-2011-000576 Harmonized patient-reported data elements in the electron...
2012 doi
-
[25]
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daum \'e III, Alexandra Olteanu, Emily Sheng, Dan Vann, and Hanna Wallach. 2023. https://doi.org/10.18653/v1/2023.acl-long.343 F air P rism: Evaluating fairness-related harms in text generation . In Proceedings of ...
2023 doi
-
[26]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. https://doi.org/10.1162/coli_a_00524 Bias and fairness in large language models: A survey . Computational Linguistics, 50(3):1097--1179
2024 doi
-
[27]
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2023. https://doi.org/10.1613/jair.1.13715 Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text . Journal of Artificial Intelligence Research, 77:103--166
2023 doi
-
[28]
Glasgow and William T
Russell E. Glasgow and William T. Riley. 2013. https://doi.org/10.1016/j.amepre.2013.03.010 Pragmatic Measures . American Journal of Preventive Medicine, 45(2):237--243
2013 doi
-
[29]
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Mu \ n oz S \'a nchez, Mugdha Pandya, and Adam Lopez. 2021. https://doi.org/10.18653/v1/2021.acl-long.150 Intrinsic bias metrics do not correlate with application bias . In Proceedings of the 59th Annual Meeting of the Asso...
2021 doi
-
[30]
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. https://doi.org/10.18653/v1/2023.findings-acl.139 This prompt is measuring < mask > : evaluating bias evaluation in language models . In Findings of the Association for Computational Linguistics...
2023 doi
-
[31]
Rishav Hada, Agrima Seth, Harshita Diddee, and Kalika Bali. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.115 fifty shades of bias : Normative ratings of gender bias in GPT generated E nglish text . In Proceedings of the 2023 Conference on Empirical Methods in Natural Lang...
2023 doi
-
[32]
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. https://doi.org/10.18653/v1/2022.acl-long.234 T oxi G en: A large-scale machine-generated dataset for adversarial and implicit hate speech detection . In Proceedings of the 60th A...
2022 doi
-
[33]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://openreview.net/forum?id=XPZIaotutsD \ DEBERTA \ : \ DECODING \ - \ enhanced \ \ bert \ \ with \ \ disentangled \ \ attention \ . In International Conference on Learning Representations
2021
-
[34]
Monique Hennink and Bonnie N. Kaiser. 2022. https://doi.org/10.1016/j.socscimed.2021.114523 Sample sizes for saturation in qualitative research: A systematic review of empirical tests . Social Science & Medicine, 292:114523
2022
-
[35]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. https://doi.org/10.1038/s41586-024-07856-5 Ai generates covertly racist decisions about people based on their dialect . Nature, 633(8028):147--154
2024 doi
-
[36]
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé, Miro Dudik, and Hanna Wallach. 2019. https://doi.org/10.1145/3290605.3300830 Improving Fairness in Machine Learning Systems : What Do Industry Practitioners Need ? In Proceedings of the 2019 CHI Conference on Human Factors...
2019
-
[37]
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. https://arxiv.org/abs/2312.06674 Llama guard: Llm-based input-output safeguard for human-ai conversations ...
2023 arXiv
-
[38]
Jacobs and Hanna Wallach
Abigail Z. Jacobs and Hanna Wallach. 2021. https://doi.org/10.1145/3442188.3445901 Measurement and Fairness . In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , FAccT , pages 375--385, New York, NY, USA. Association for Computing Machinery
2021
-
[39]
Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. 2023. https://doi.org/10.1609/aaai.v37i12.26670 Taxonomizing and measuring representational harms: a look at image tagging . In Proceedings of the Thirty-Seventh ...
2023 doi
-
[40]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in large language models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery
2023
-
[41]
Michelle Seng Ah Lee and Jat Singh. 2021. https://doi.org/10.1145/3411764.3445261 The Landscape and Gaps in Open Source Fairness Toolkits . In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , CHI , pages 1--13, Yokohama Japan. ACM
2021
-
[42]
Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman
Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. https://doi.org/10.1145/3534678.3539147 A new generation of perspective api: Efficient multilingual character-level transformers . In Proceedings of the 28th ACM SIGKDD Co...
2022
-
[43]
Lewis and Caitlin Dorsey
Cara C. Lewis and Caitlin Dorsey. 2020. https://doi.org/10.1007/978-3-030-03874-8_9 Advancing Implementation Science Measurement . In Bianca Albers, Aron Shlonsky, and Robyn Mildon, editors, Implementation Science 3.0 , pages 227--251. Springer International Publishing, Cham
2020 doi
-
[44]
Cara C Lewis, Kayne D Mettert, Cameo F Stanick, Heather M Halko, Elspeth A Nolen, Byron J Powell, and Bryan J Weiner. 2021. https://doi.org/10.1177/26334895211037391 The psychometric and pragmatic evidence rating scale ( PAPERS ) for measure development and evaluation . Implem...
2021 doi
-
[45]
Ahmed Magooda, Alec Helyar, Kyle Jackson, David Sullivan, Chad Atalla, Emily Sheng, Dan Vann, Richard Edgar, Hamid Palangi, Roman Lutz, Hongliang Kong, Vincent Yun, Eslam Kamal, Federico Zarfati, Hanna Wallach, Sarah Bird, and Mei Chen. 2023. http://arxiv.org/abs/2310.17750 A ...
2023 arXiv
-
[46]
Gaurav Maheshwari, Aur \'e lien Bellet, Pascal Denis, and Mikaela Keller. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.558 Fair without leveling down: A new intersectional fairness definition . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language...
2023 doi
-
[47]
Martinez, Cara C
Ruben G. Martinez, Cara C. Lewis, and Bryan J. Weiner. 2014. https://doi.org/10.1186/s13012-014-0118-8 Instrumentation issues in implementation science . Implementation science: IS, 9:118
2014 doi
-
[48]
Bowman, and Rachel Rudinger
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. https://doi.org/10.18653/v1/N19-1063 On measuring social biases in sentence encoders . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational...
2019 doi
-
[49]
David L Morgan. 2008. Snowball sampling. The SAGE encyclopedia of qualitative research methods, 2:815--16
2008
-
[50]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...
2021 doi
-
[51]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[52]
Jekaterina Novikova, Ond r ej Du s ek, Amanda Cercas Curry, and Verena Rieser. 2017. https://doi.org/10.18653/v1/D17-1238 Why we need new evaluation metrics for NLG . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2241--2252, C...
2017 doi
-
[53]
Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. 2024. http://arxiv.org/abs/2402.17861 Towards AI Accountability Infrastructure : Gaps and Opportunities in AI Audit Tooling . arXiv preprint
2024 arXiv
-
[54]
Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, Roger Chou, Julie Glanville, Jeremy M Grimshaw, Asbj rn Hr \'o bjartsson, Manoj M Lalu, Tianjing Li, El...
2021 doi
-
[55]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. https://doi.org/10.18653/v1/2022.findings-acl.165 BBQ : A hand-built bias benchmark for question answering . In Findings of the Association for ...
2022 doi
-
[56]
Ian Porada, Alexandra Olteanu, Kaheer Suleman, Adam Trischler, and Jackie Cheung. 2024. https://doi.org/10.18653/v1/2024.findings-acl.909 Challenges to evaluating the generalization of coreference resolution models: A measurement modeling perspective . In Findings of the Assoc...
2024 doi
-
[57]
Byron J Powell, Thomas J Waltz, Matthew J Chinman, Laura J Damschroder, Jeffrey L Smith, Monica M Matthieu, Enola K Proctor, and JoAnn E Kirchner. 2015. https://doi.org/10.1186/s13012-015-0209-1 A refined compilation of implementation strategies: results from the Expert Recomm...
2015 doi
-
[58]
Ehud Reiter. 2018. https://doi.org/10.1162/coli_a_00322 A structured review of the validity of BLEU . Computational Linguistics, 44(3):393--401
2018 doi
-
[59]
Way, Jennifer Thom, and Henriette Cramer
Brianna Richardson, Jean Garcia-Gathright, Samuel F. Way, Jennifer Thom, and Henriette Cramer. 2021. https://doi.org/10.1145/3411764.3445604 Towards Fairness in Practice : A Practitioner - Oriented Rubric for Evaluating Fair ML Toolkits . In Proceedings of the 2021 CHI Confere...
2021
-
[60]
Morgan Klaus Scheuerman. 2024. https://doi.org/10.1145/3630106.3658918 In the Walled Garden : Challenges and Opportunities for Research on the Practices of the AI Tech Industry . In The 2024 ACM Conference on Fairness , Accountability , and Transparency , pages 456--466, Rio d...
2024
-
[61]
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2023. Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. arXiv preprint arXiv:2310.11324
2023 arXiv
-
[62]
Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022. https://openreview.net/forum?id=rIhzjia7SLa Quantifying social biases using templates is unreliable . In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022
2022
-
[63]
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. https://doi.org/10.18653/v1/2021.acl-long.330 Societal biases in language generation: Progress and challenges . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the...
2021 doi
-
[64]
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. https://doi.org/10.18653/v1/D19-1339 The woman worked as a babysitter: On biases in language generation . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9...
2019 doi
-
[65]
Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, and David Jurgens. 2024. https://doi.org/10.18653/v1/2024.naacl-long.295 You don`t need a personality test to know these models are unreliable: Assessing the reliability of...
2024 doi
-
[66]
Mario Luis Small. 2009. https://doi.org/10.1177/1466138108099586 `how many cases do i need?': On science and the logic of case selection in field-based research . Ethnography, 10(1):5--38
2009 doi
-
[67]
Stanick, Heather M
Cameo F. Stanick, Heather M. Halko, Elspeth A. Nolen, Byron J. Powell, Caitlin N. Dorsey, Kayne D. Mettert, Bryan J. Weiner, Melanie Barwick, Luke Wolfenden, Laura J. Damschroder, and Cara C. Lewis. 2021. https://doi.org/10.1093/tbm/ibz164 Pragmatic measures for implementation...
2021 doi
-
[68]
Kaiser Sun, Adina Williams, and Dieuwke Hupkes. 2023. https://doi.org/10.18653/v1/2023.conll-1.19 The validity of evaluation results: Assessing concurrence across compositionality benchmarks . In Proceedings of the 27th Conference on Computational Natural Language Learning (Co...
2023 doi
-
[69]
Oskar Van Der Wal, Dominik Bachmann, Alina Leidinger, Leendert Van Maanen, Willem Zuidema, and Katrin Schulz. 2024. https://doi.org/10.1613/jair.1.15195 Undesirable Biases in NLP : Addressing Challenges of Measurement . Journal of Artificial Intelligence Research, 79:1--40
2024 doi
-
[70]
Aniket Vashishtha, S Sai Prasad, Payal Bajaj, Vishrav Chaudhary, Kate Cook, Sandipan Dandapat, Sunayana Sitaram, and Monojit Choudhury. 2023. https://doi.org/10.18653/v1/2023.findings-eacl.167 Performance and risk trade-offs for multi-word text prediction at scale . In Finding...
2023 doi
-
[71]
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson. 2022. https://aclanthology.org/2022.coling-1.113/ A study of implicit bias in pretrained language models against people with disabilities . In Proceedings of the 29th International Conference on Computational Linguist...
2022
-
[72]
Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P
Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaugha...
2025 arXiv
-
[73]
Waltz, Byron J
Thomas J. Waltz, Byron J. Powell, Monica M. Matthieu, Laura J. Damschroder, Matthew J. Chinman, Jeffrey L. Smith, Enola K. Proctor, and JoAnn E. Kirchner. 2015. https://doi.org/10.1186/s13012-015-0295-0 Use of concept mapping to characterize relationships among implementation ...
2015 doi
-
[74]
Angelina Wang, Xuechunzi Bai, Solon Barocas, and Su Lin Blodgett. 2023. Measuring stereotype harm from machine learning errors requires understanding who is being harmed by which errors in what ways. In ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimiz...
2023
-
[75]
Vera Liao
Ziang Xiao, Susu Zhang, Vivian Lai, and Q. Vera Liao. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.676 Evaluating evaluation metrics: A framework for analyzing NLG evaluation metrics using measurement theory . In Proceedings of the 2023 Conference on Empirical Methods in ...
2023 doi
- [76]
-
[77]
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daum \'e III, Kaheer Suleman, and Alexandra Olteanu. 2022. https://doi.org/10.18653/v1/2022.naacl-main.24 Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications . In Proceedings of the 2022 Co...
2022 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.