REVIEW 3 major objections 4 minor 102 references
The Only Way is Ethics: A Guide to Ethical Research with Large Language Models
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The authors claim no integrated, practical LLM ethics guide exists, and present one: a living whitepaper that turns a literature review into stage-by-stage Do's and Don'ts for NLP practitioners.
desk verdict A genuinely useful lifecycle-organized ethics guide for LLM practitioners, with a methodology that needs one round of tightening to match its 'comprehensive' billing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying device is the simplified project lifecycle, a six-stage diagram that structures both the whitepaper and this overview. It turns 'ethics' from a diffuse topic into a sequence of concrete decisions: what to do at project start, how to compile and prepare data, how to choose and develop a model, how to evaluate it, and how to deploy it. For each stage the paper pairs a set of Do's and Don'ts with a small set of tools — ethics sheets and internal audits at the start, datasheets and scraping guidance during data work, model cards and debiasing cautions during development, harm-evaluation taxonomies and red-teaming at evaluation, and release strategies and monitoring at deployment. The machinery works by making each recommendation point to a named, usable resource.
What would settle it
A traceability audit would settle it: compare every Do and Don't against the full set of papers returned by the two literature searches and the tutorial references. If a substantial share of recommendations cannot be traced to a cited source, or if major harm categories documented in the cited taxonomies (for example, non-English or non-text harms) are missing from the guide, the distillation claim fails.
Extended reading notes
Core claim
The paper claims that existing ethical guidance is split between broad regulatory frameworks, which are hard to act on, and submission checklists, which are too thin, leaving LLM practitioners without an integrated, practical reference. It attempts to close that gap with the LLM Ethics Whitepaper, whose distilled content appears here as Do's and Don'ts organized by a simplified project lifecycle. Each lifecycle stage — project start, data compilation, data preparation, model development and selection, evaluation, deployment — comes with concrete recommendations and pointers to toolkits such as ethics sheets, datasheets, model cards, bias-test suites, and staged-release strategies. The authors state that the recommendations were drawn from a literature review of the ACL Anthology and Semantic Scholar, supplemented by the authors' expertise, and that the whitepaper cites the hundreds of underlying papers. The present paper is offered as the condensed pocket version.
Load-bearing premise
The load-bearing premise is that the authors' literature selection and expert synthesis give a representative picture of LLM ethics, a premise the paper itself narrows by conceding its resources are English-centric, mostly Western, and text-only.
Editorial extensions
If this is right
- A researcher can use the paper as a reference during a project rather than as a post-hoc checklist, because each section maps to a distinct stage of work.
- Following the Do's and Don'ts yields concrete artifacts: ethics sheets before starting, datasheets for new datasets, model cards for released models, and staged-release plans before deployment.
- The guide is meant to be more actionable than broad AI frameworks and more detailed than association submission checklists, filling the middle space for LLM-specific work.
- Because the resource is hosted as a living document on GitHub, the recommendations can be revised in response to practitioner feedback and future editions can extend beyond the current English, Western, text-only scope.
Reading between the lines
- If adopted widely, the stage-by-stage format could push ethics from a review-time formality toward a set of project artifacts that reviewers and auditors can check for.
- The acknowledged English-centric and Western skew implies a testable prediction: practitioners working on non-English or multimodal systems will find fewer tools and recommendations directly applicable to their setting.
- The paper's Do's and Don'ts could be operationalized further by linking each one to a verifiable output, effectively turning the guide into a lightweight audit checklist.
- Because the companion whitepaper and this overview are versioned separately, the 'living' claim means the arXiv paper may not remain identical to the GitHub resource; readers should treat the online version as the current authority.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces a practical guide to ethical research with large language models, structured around a simplified project lifecycle (project start, data compilation, data preparation, model development, evaluation, deployment). The authors report a literature review of the ACL Anthology and Semantic Scholar, supplemented by expert-selected resources, and distill the findings into concrete Do's and Don'ts plus a list of useful toolkits for each lifecycle stage. The paper also points to a longer companion whitepaper (Ungless et al., 2024) that contains the full citations and detailed discussion. The stated contribution is a practitioner-facing resource that fills a gap between broad AI frameworks (NIST, EU AI Act) and association-level checklists (NeurIPS, ACL).
Significance. If the guide is adopted, it would provide NLP researchers with an actionable, lifecycle-structured reference that translates a large body of ethics literature into specific recommendations. The open and living format on GitHub, the explicit separation of Do's and Don'ts, and the curation of toolkits are genuine practical strengths. The paper is also honest about its limitations, acknowledging English-centricity, a largely Western perspective, and a text-only focus. However, the significance is contingent on the representativeness of the literature selection: the authors' claim of a 'comprehensive directory' is load-bearing because the authority of the Do's and Don'ts rests on the quality and coverage of the underlying review. As reported, the methodology does not allow a reader to verify that claim.
major comments (3)
- [Section 2] The protocol for the literature review is described with search strings but without the information needed to audit the resulting 'comprehensive directory': the number of hits retrieved, the screening criteria, the inclusion/exclusion counts, and the inter-reviewer agreement are all omitted. The subsequent admission that resources were added 'ad hoc' and removed for subjective reasons (e.g., 'limited relevance', 'too narrow') means the final set is an unmeasured mixture of systematic search and expert curation. Because the Do's and Don'ts are explicitly 'drawn from' this review, the representativeness of the selection is load-bearing for the guide's authority. I recommend either reporting a full PRISMA-style flow and screening protocol or reframing the deliverable as a curated, non-exhaustive resource rather than a 'comprehensive directory'.
- [Abstract and Section 1] The claim that 'as yet there is no work that integrates these resources into a single practical guide' is a strong negative existence claim, yet the evidence reported is confined to two convenience sources (ACL Anthology and Semantic Scholar) and the authors' own familiarity. No recall check or comparison with an independent seed set of LLM-ethics resources is provided, so a reader cannot verify that the gap is not an artifact of the search strategy. Please either soften the claim to 'we are not aware of' or provide a more systematic basis for the gap claim.
- [Section 9 versus Section 3.5] The guide states that the Do's and Don'ts are applicable 'regardless of model language,' but many of the recommended toolkits and evaluation resources are English-specific, and the overall perspective is acknowledged as largely Western. The limitation is properly acknowledged in Section 9, but the main text (e.g., the abstract and Section 1) presents the guide without this scope restriction. This tension should be resolved by integrating the scope limitation into the framing rather than only in the concluding section.
minor comments (4)
- [Section 3.4] The sentence 'Sasha Luccioni and colleagues have in particular championed the accurate reporting of the carbon emissions of ML systems including LLMs' lacks a citation; please add a reference to enable follow-up.
- [References (Smith et al., 2022)] In the reference for Smith et al. (2022), 'Melissa Hall Melanie Kambadur' should be 'Melissa Hall, Melanie Kambadur'.
- [Section 1 and Section 9] The spelling 'Github' should be 'GitHub' in both locations.
- [Section 2] The Semantic Scholar query 'toolkit OR sheets OR guideline OR principles OR framework OR approach ethics OR ethical OR harms OR fair OR fairness OR risk AND "language models"' is ambiguous due to operator precedence; please clarify the intended Boolean grouping with parentheses.
Circularity Check
No circularity; the guide is a transparent literature synthesis with non-load-bearing self-citations.
full rationale
This paper is a practitioner-oriented synthesis of existing external ethics literature into Do's and Don'ts; it makes no quantitative predictions, fits no parameters, and derives no mathematical claims. The only self-references are (1) the companion whitepaper (Ungless et al., 2024), which is the artifact being summarized rather than an input to the synthesis, and (2) several of the authors' prior research papers cited as relevant resources (e.g., Talat et al. 2021; Goldfarb-Tarrant et al. 2023; Kasirzadeh and Gabriel 2023). These citations are not load-bearing justifications: the recommendations are explicitly 'directly drawn from our extensive literature review' (Section 1), and the methodology describes two systematic searches (ACL Anthology and Semantic Scholar) plus manual review and ad hoc additions. The admitted ad hoc curation is a limitation in verifiable comprehensiveness, not circularity, because the guide's claims do not reduce to the authors' own prior conclusions. No step equates an output to an input by construction. The central contribution—translating the literature into concrete guidance—is self-contained and externally anchored, so the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Manual literature review plus the authors' expertise is sufficient to select and categorize relevant ethical resources for LLM research.
- domain assumption Ethical considerations can be meaningfully mapped onto a linear project lifecycle covering start, data, model, evaluation, and deployment.
- domain assumption The cited prior work is authoritative and the Do's and Don'ts accurately reflect that work.
- domain assumption Advice developed from English, text-only, largely Western sources is transferable to LLM research more broadly.
Cite this review
Pith. "Pith review of The Only Way is Ethics: A Guide to Ethical Research with Large Language Models." pith.science (2026). https://pith.science/paper/73ITDKFL
@misc{pith2026241216022,
author = {Pith},
title = {Pith review of: The Only Way is Ethics: A Guide to Ethical Research with Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/73ITDKFL}},
note = {Machine review of arXiv:2412.16022}
}
read the original abstract
There is a significant body of work looking at the ethical considerations of large language models (LLMs): critiquing tools to measure performance and harms; proposing toolkits to aid in ideation; discussing the risks to workers; considering legislation around privacy and security etc. As yet there is no work that integrates these resources into a single practical guide that focuses on LLMs; we attempt this ambitious goal. We introduce 'LLM Ethics Whitepaper', which we provide as an open and living resource for NLP practitioners, and those tasked with evaluating the ethical implications of others' work. Our goal is to translate ethics literature into concrete recommendations and provocations for thinking with clear first steps, aimed at computer scientists. 'LLM Ethics Whitepaper' distils a thorough literature review into clear Do's and Don'ts, which we present also in this paper. We likewise identify useful toolkits to support ethical work. We refer the interested reader to the full LLM Ethics Whitepaper, which provides a succinct discussion of ethical considerations at each stage in a project lifecycle, as well as citations for the hundreds of papers from which we drew our recommendations. The present paper can be thought of as a pocket guide to conducting ethical research with LLMs.
Figures
Reference graph
Works this paper leans on
-
[1]
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser, and Zeerak Talat. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.290 Mirages. on anthropomorphism in dialogue systems . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4776--4790, Singapore. Association for Computational Linguistics
-
[2]
Abhishek Anand, Negar Mokhberian, Prathyusha Naresh Kumar, Anweasha Saha, Zihao He, Ashwin Rao, Fred Morstatter, and Kristina Lerman. 2024. https://doi.org/10.48550/arXiv.2403.04085 Don't Blame the Data , Blame the Model : Understanding Noise and Bias When Learning from Subjective Annotations . Preprint, arXiv:2403.04085
-
[3]
Markus Anderljung, Joslyn Barnhart, Jade Leung, Anton Korinek, Cullen O'Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, et al. 2023. Frontier ai regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2307.03718
arXiv 2023
-
[4]
Nesrine Bannour, Sahar Ghannay, Aurélie Névéol, and Anne-Laure Ligozat. 2021. https://doi.org/10.18653/v1/2021.sustainlp-1.2 Evaluating the carbon footprint of NLP methods: a survey and analysis of existing tools . In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing , pages 11--21, Virtual. Association for Computation...
-
[5]
Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. https://doi.org/10.1145/3442188.3445922 On the Dangers of Stochastic Parrots : Can Language Models Be Too Big ? In Proceedings of the 2021 ACM Conference on Fairness , Accountability , and Transparency , pages 610--623, Virtual Event Canada. ACM
arXiv 2021
-
[6]
Luciana Benotti, Karën Fort, Min-Yen Kan, and Yulia Tsvetkov. 2023. https://doi.org/10.18653/v1/2023.eacl-tutorials.4 Understanding Ethics in NLP Authoring and Reviewing . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics : Tutorial Abstracts , pages 19--24, Dubrovnik, Croatia. Association for C...
-
[8]
Steven Bird. 2020. https://doi.org/10.18653/v1/2020.coling-main.313 Decolonising Speech and Language Technology . In Proceedings of the 28th International Conference on Computational Linguistics , pages 3504--3519, Barcelona, Spain (Online). International Committee on Computational Linguistics
-
[9]
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022. https://doi.org/10.1145/3531146.3533083 The Values Encoded in Machine Learning Research . In 2022 ACM Conference on Fairness , Accountability , and Transparency , pages 173--184, Seoul Republic of Korea. ACM
arXiv 2022
Show all 102 references
-
[10]
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language ( Technology ) is Power : A Critical Survey of “ Bias ” in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Lingui...
2020 doi
-
[11]
Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. 2021. https://www.microsoft.com/en-us/research/publication/stereotyping-norwegian-salmon-an-inventory-of-pitfalls-in-fairness-benchmark-datasets/ Stereotyping Norwegian Salmon : An Inventory of ...
2021
-
[12]
Su Lin Blodgett and Brendan O'Connor. 2017. http://arxiv.org/abs/1707.00061 Racial Disparity in Natural Language Processing : A Case Study of Social Media African - American English . arXiv:1707.00061 [cs]. ArXiv: 1707.00061
2017 arXiv
-
[13]
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016. https://proceedings.neurips.cc/paper/2016/file/a486cd07e4ac3d270571622f4f316ec5-Paper.pdf Man is to Computer Programmer as Woman is to Homemaker ? Debiasing Word Embeddings . In Advances ...
2016
-
[14]
Joy Buolamwini and Timnit Gebru. 2018. https://proceedings.mlr.press/v81/buolamwini18a.html Gender Shades : Intersectional Accuracy Disparities in Commercial Gender Classification . In Conference on Fairness , Accountability and Transparency , pages 77--91. PMLR. ISSN: 2640-3498
2018
-
[15]
Agostina Calabrese, Michele Bevilacqua, Björn Ross, Rocco Tripodi, and Roberto Navigli. 2021. https://doi.org/10.1145/3447535.3462484 AAA : Fair Evaluation for Abuse Detection Systems Wanted . In 13th ACM Web Science Conference 2021 , pages 243--252, Virtual Event United Kingdom. ACM
2021
-
[16]
Tommaso Caselli, Roberto Cibin, Costanza Conforti, Enrique Encinas, and Maurizio Teli. 2021. https://doi.org/10.18653/v1/2021.nlp4posimpact-1.4 Guiding Principles for Participatory Design -inspired Natural Language Processing . In Proceedings of the 1st Workshop on NLP for Pos...
2021 doi
-
[17]
Yu, Qian Yang, and Xingxu Xie
Yu-Chu Chang, Xu Wang, Jindong Wang, Yuanyi Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Weirong Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qian Yang, and Xingxu Xie. 2023. https://api.semanticscholar.org/CorpusID:259360395 A Survey on Evaluation ...
2023 arXiv
-
[18]
grandma exploit
Anthony Cuthbertson. 2023. https://www.independent.co.uk/tech/chatgpt-microsoft-windows-11-grandma-exploit-b2360213.html ChatGPT “grandma exploit” helps people pirate software . Publication Title: The Independent
2023
-
[19]
David Danks. 2022. Digital ethics as translational ethics. In Applied ethics in a digital world, pages 1--15. IGI Global
2022
-
[20]
Thomas Davidson, Debasmita Bhattacharya, and Ingmar Weber. 2019. https://doi.org/10.18653/v1/W19-3504 Racial Bias in Hate Speech and Abusive Language Detection Datasets . In Proceedings of the Third Workshop on Abusive Language Online , pages 25--35, Florence, Italy. Associati...
2019 doi
- [22]
-
[23]
Gender
Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022. https://doi.org/10.1145/3531146.3534627 Theories of “ Gender ” in NLP Bias Research . In 2022 ACM Conference on Fairness , Accountability , and Transparency , FAccT '22, pages 2083--2102, New York, NY, USA. Associat...
2022
-
[24]
Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser
Emily Dinan, Gavin Abercrombie, A. Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2022. https://doi.org/10.18653/v1/2022.acl-long.284 SafetyKit : First Aid for Measuring Safety in Open -domain Conversational Systems . In Proceedings of the 60th Annual Me...
2022 doi
-
[25]
Smith, and Jesse Dodge
Yanai Elazar, Akshita Bhagia, Ian Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hanna Hajishirzi, Noah A. Smith, and Jesse Dodge. 2024. https://doi.org/10.48550/arXiv.2310.20707 What's In My Big Data ? a...
-
[26]
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024. Bias and fairness in large language models: A survey. Computational Linguistics, pages 1--79
2024
-
[27]
Brown, Nicholas Joseph, Sam McCandlish, Christopher Olah, Jared Kaplan, and Jack Clark
Deep Ganguli, Liane Lovitt, John Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Benjamin Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zachary...
2022 arXiv
-
[28]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2020. http://arxiv.org/abs/1803.09010 Datasheets for Datasets . arXiv:1803.09010 [cs]
2020 arXiv
-
[29]
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. https://doi.org/10.18653/v1/2023.findings-acl.139 This prompt is measuring mask : evaluating bias evaluation in language models . In Findings of the Association for Computational Linguistics : A...
2023 doi
-
[30]
Hila Gonen and Yoav Goldberg. 2019. https://doi.org/10.18653/v1/N19-1061 Lipstick on a Pig : Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them . In Proceedings of the 2019 Conference of the North American Chapter of the Association f...
2019 doi
-
[31]
Luke Guerdan, Amanda Coston, Zhiwei Steven Wu, and Kenneth Holstein. 2023. https://doi.org/10.1145/3593013.3594036 Ground(less) Truth : A Causal Framework for Proxy Labels in Human - Algorithm Decision - Making . In 2023 ACM Conference on Fairness , Accountability , and Transp...
2023
-
[32]
Lucy Havens, Melissa Terras, Benjamin Bach, and Beatrice Alex. 2020. http://arxiv.org/abs/2011.05911 Situated Data , Situated Systems : A Methodology to Engage with Power Relations in Natural Language Processing Research . arXiv:2011.05911 [cs]. ArXiv: 2011.05911
2020 arXiv
-
[33]
Brunskill, Dan Jurafsky, and Joelle Pineau
Peter Henderson, Jie Hu, Joshua Romoff, E. Brunskill, Dan Jurafsky, and Joelle Pineau. 2020. https://www.semanticscholar.org/paper/74b4f16c5ac91e3e7c88ae81cc8c91416b71d151 Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning . ArXiv
2020
- [34]
-
[35]
algorithmic bias is a data problem
Sara Hooker. 2021. https://doi.org/10.1016/j.patter.2021.100241 Moving beyond “algorithmic bias is a data problem” . Patterns, 2(4):100241
2021
-
[36]
Xiaowei Huang, Wenjie Ruan, Wei Huang, Gao Jin, Yizhen Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai, Yanghao Zhang, Sihao Wu, Peipei Xu, Dengyu Wu, André Freitas, and Mustafa A. Mustafa. 2023. https://api.semanticscholar.org/CorpusID:25882308...
2023 arXiv
-
[37]
Nanna Inie and Leon Derczynski. 2021. https://aclanthology.org/2021.hcinlp-1.16 An IDR Framework of Opportunities and Barriers between HCI and NLP . In Proceedings of the First Workshop on Bridging Human – Computer Interaction and Natural Language Processing , pages 101--108, ...
2021
- [38]
-
[39]
Atoosa Kasirzadeh. 2021. Reasons, values, stakeholders: A philosophical framework for explainable artificial intelligence. arXiv preprint arXiv:2103.00752
2021 arXiv
-
[40]
Atoosa Kasirzadeh. 2024. https://openreview.net/pdf?id=AOokh1UYLH Plurality of value pluralism and ai value alignment . In Pluralistic Alignment Workshop at NeurIPS 2024
2024
-
[41]
Atoosa Kasirzadeh and Iason Gabriel. 2023. In conversation with artificial intelligence: aligning language models with human values. Philosophy & Technology, 36(2):27
2023
-
[42]
Anna Kawakami, Amanda Coston, Haiyi Zhu, Hoda Heidari, and Kenneth Holstein. 2024. https://doi.org/10.1145/3613904.3642849 The Situate AI Guidebook : Co - Designing a Toolkit to Support Multi - Stakeholder Early -stage Deliberations Around Public Sector AI Proposals . ArXiv:24...
2024
-
[43]
Sachin Kumar, Vidhisha Balachandran, Lucille Njoo, Antonios Anastasopoulos, and Yulia Tsvetkov. 2022. https://api.semanticscholar.org/CorpusID:252907607 Language Generation Models Can Cause Harm : So What Can We Do About It ? An Actionable Survey . In Conference of the Europea...
2022
-
[44]
Markus Langer, Daniel Oster, Timo Speith, Holger Hermanns, Lena Kästner, Eva Schmidt, Andreas Sesing, and Kevin Baum. 2021. https://doi.org/10.1016/j.artint.2021.103473 What do we want from Explainable Artificial Intelligence ( XAI )? – A stakeholder perspective on XAI and a c...
2021
-
[45]
Brian Larson. 2017. https://doi.org/10.18653/v1/W17-1601 Gender as a Variable in Natural - Language Processing : Ethical Considerations . In Proceedings of the First ACL Workshop on Ethics in Natural Language Processing , pages 1--11, Valencia, Spain. Association for Computati...
2017 doi
-
[46]
Leidner and Vassilis Plachouras
Jochen L. Leidner and Vassilis Plachouras. 2017. https://doi.org/10.18653/v1/W17-1604 Ethical by Design : Ethics Best Practices for Natural Language Processing . In Proceedings of the First ACL Workshop on Ethics in Natural Language Processing , pages 30--40, Valencia, Spain. ...
2017 doi
-
[47]
Dave Lewis, Linda Hogan, David Filip, and P. J. Wall. 2020. https://doi.org/10.13052/jicts2245-800X.823 Global Challenges in the Standardization of Ethics for Trustworthy AI . Journal of ICT Standardization
2020 doi
-
[48]
Calvin Liang. 2021. https://medium.com/@caliang/reflexivity-positionality-and-disclosure-in-hci-3d95007e9916 Reflexivity, positionality, and disclosure in HCI
2021
-
[49]
Liang, Sean A
Calvin A. Liang, Sean A. Munson, and Julie A. Kientz. 2021. https://doi.org/10.1145/3443686 Embracing Four Tensions in Human - Computer Interaction Research with Marginalized People . ACM Transactions on Computer-Human Interaction, 28(2):1--47
2021 doi
-
[50]
Jie Liu et al. 2012. The enterprise risk management and the risk oriented internal audit. Ibusiness, 4(03):287
2012
-
[51]
Zoey Liu, Crystal Richardson, Richard Hatcher, and Emily Prud'hommeaux. 2022. https://doi.org/10.18653/v1/2022.acl-long.272 Not always about you: Prioritizing community needs when developing endangered language technology . In Proceedings of the 60th Annual Meeting of the Asso...
2022 doi
-
[52]
Nicola Lucchi. 2023. https://doi.org/10.1017/err.2023.59 Chatgpt: A case study on copyright challenges for generative artificial intelligence systems . European Journal of Risk Regulation, page 1–23
2023 doi
-
[53]
Nitin Madnani, Anastassia Loukina, Alina von Davier, Jill Burstein, and Aoife Cahill. 2017. https://doi.org/10.18653/v1/W17-1605 Building Better Open - Source Tools to Support Fairness in Automated Scoring . In Proceedings of the First ACL Workshop on Ethics in Natural Languag...
2017 doi
-
[54]
Keoni Mahelona, Gianna Leoni, Suzanne Duncan, and Miles Thompson. 2023. https://blog.papareo.nz/whisper-is-another-case-study-in-colonisation/ OpenAI ’s whisper is another case study in colonisation . Papa Reo
2023
-
[55]
James W Malazita and Korryn Resetar. 2019. Infrastructures of abstraction: how computer science education produces anti-political subjects. Digital Creativity, 30(4):300--312
2019
-
[56]
Moreno Mancosu and Federico Vegetti. 2020. https://doi.org/10.1177/2056305120940703 What You Can Scrape and What Is Right to Scrape : A Proposal for a Tool to Collect Public Facebook Data . Social Media + Society, 6(3):2056305120940703. Publisher: SAGE Publications Ltd
2020 doi
-
[57]
Nina Markl. 2022. https://doi.org/10.18653/v1/2022.ltedi-1.1 Mind the data gap(s): Investigating power in speech and language datasets . In Proceedings of the Second Workshop on Language Technology for Equality , Diversity and Inclusion , pages 1--12, Dublin, Ireland. Associat...
2022 doi
-
[58]
Andrew McNamara, Justin Smith, and Emerson Murphy-Hill. 2018. Does acm’s code of ethics change ethical decision making in software development? In Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of sof...
2018
-
[59]
Boaz Miller. 2021. Is technology value-neutral? Science, Technology, & Human Values, 46(1):53--80
2021
-
[60]
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023. https://doi.org/10.1145/3605943 Recent Advances in Natural Language Processing via Large Pre -trained Language Models : A Survey . ACM Co...
2023 doi
-
[61]
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. https://doi.org/10.1145/3287560.3287596 Model Cards for Model Reporting . Proceedings of the Conference on Fairness, Acco...
2019
-
[62]
Saif Mohammad. 2022. https://doi.org/10.18653/v1/2022.acl-long.573 Ethics Sheets for AI Tasks . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ( Volume 1: Long Papers ) , pages 8368--8379, Dublin, Ireland. Association for Computation...
2022 doi
-
[63]
Lucia Nalbandian. 2022. https://doi.org/10.1186/s40878-022-00305-0 An eye for an ' I :' a critical assessment of artificial intelligence tools in migration and asylum management . Comparative Migration Studies, 10(1):32
2022 doi
-
[64]
Nathan, Predrag V
Lisa P. Nathan, Predrag V. Klasnja, and Batya Friedman. 2007. https://doi.org/10.1145/1240866.1241046 Value scenarios: a technique for envisioning systemic effects of new technologies . In CHI '07 Extended Abstracts on Human Factors in Computing Systems , CHI EA '07, pages 258...
2007
-
[65]
Timothy Niven and Hung-Yu Kao. 2019. https://doi.org/10.18653/v1/P19-1459 Probing neural network comprehension of natural language arguments . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4658--4664, Florence, Italy. Associa...
2019 doi
-
[66]
Anaelia Ovalle, Arjun Subramonian, Vagrant Gautam, Gilbert Gee, and Kai-Wei Chang. 2023. https://doi.org/10.1145/3600211.3604705 Factoring the Matrix of Domination : A Critical Review and Reimagination of Intersectionality in AI Fairness . In Proceedings of the 2023 AAAI / ACM...
2023
-
[67]
So, Maud Texier, and Jeff Dean
David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David R. So, Maud Texier, and Jeff Dean. 2022. https://doi.org/10.1109/MC.2022.3148714 The Carbon Footprint of Machine Learning Training Will Plateau , Then Shrink . Comp...
2022
-
[68]
Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh, Alexandra Sasha Luccioni, and Margaret Mitchell. 2024. Civics: Building a dataset for examining culturally-informed values in large language models. arXiv preprint arXiv:2405.13974
2024 arXiv
-
[69]
Inioluwa Deborah Raji, Morgan Klaus Scheuerman, and Razvan Amironesei. 2021. You can't sit with us: Exclusionary pedagogy in ai ethics education. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 515--525
2021
-
[70]
Inioluwa Deborah Raji, Andrew Smart, Rebecca N White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes. 2020. Closing the AI Accountability Gap : Defining an End -to- End Framework for Internal Algorithmic Auditing . page 12
2020
- [71]
-
[72]
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.acl-main.442 Beyond Accuracy : Behavioral Testing of NLP Models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguis...
2020 doi
-
[73]
Anna Rogers, Timothy Baldwin, and Kobi Leins. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.414 ` Just What do You Think You 're Doing , Dave ?' A Checklist for Responsible Data Use in NLP . In Findings of the Association for Computational Linguistics : EMNLP 2021 , pa...
2021 doi
-
[74]
Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. 2024. https://aclanthology.org/2024.naacl-long.301 XSTest : A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models . In Proceedings of the 2024 Conferenc...
2024
-
[75]
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Talat, Helen Margetts, and Janet Pierrehumbert. 2021. https://doi.org/10.18653/v1/2021.acl-long.4 HateCheck : Functional Tests for Hate Speech Detection Models . In Proceedings of the 59th Annual Meeting of the Association for C...
2021 doi
-
[76]
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971--30004. PMLR
2023
-
[77]
Boaz Shmueli, Jan Fell, Soumya Ray, and Lun-Wei Ku. 2021. https://doi.org/10.18653/v1/2021.naacl-main.295 Beyond Fair Pay : Ethical Implications of NLP Crowdsourcing . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Ling...
2021 doi
-
[78]
Ben Shneiderman. 2020. https://doi.org/10.1080/10447318.2020.1741118 Human- Centered Artificial Intelligence : Reliable , Safe & Trustworthy . International Journal of Human–Computer Interaction, 36(6):495--504. Publisher: Taylor & Francis \_eprint: https://doi.org/10.1080/104...
2020
-
[79]
Mona Sloane, Emanuel Moss, Olaitan Awomolo, and Laura Forlano. 2022. https://doi.org/10.1145/3551624.3555285 Participation Is not a Design Fix for Machine Learning . In Equity and Access in Algorithms , Mechanisms , and Optimization , pages 1--6, Arlington VA USA. ACM
2022
-
[80]
I 'm sorry to hear that
Eric Michael Smith, Melissa Hall Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. http://arxiv.org/abs/2205.09209 " I 'm sorry to hear that": finding bias in language models with a holistic descriptor dataset . Technical Report arXiv:2205.09209, arXiv. ArXiv:2205....
2022 arXiv
-
[81]
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. http://arxiv.org/abs/1908.09203 Release Strat...
2019 arXiv
-
[82]
Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daum \'e III, Jesse Dodge, Isabella Duan, Ellie Evans, Felix Friedrich, Avijit Ghosh, Usman Gohar, Sara Hooker, Yacine Jernite, Ria Kalluri, Alberto Lusoli, Alina Leidinger, ...
2024 arXiv
-
[83]
Ryan Steed, Swetasudha Panda, Ari Kobren, and Michael L. Wick. 2022. https://doi.org/10.18653/v1/2022.acl-long.247 Upstream Mitigation Is Not All You Need : Testing the Bias Transfer Hypothesis in Pre - Trained Language Models . In ACL
2022 doi
-
[84]
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. https://doi.org/10.18653/v1/P19-1355 Energy and Policy Considerations for Deep Learning in NLP . Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645--3650. Conference Name:...
2019 doi
-
[85]
Nishant Subramani, Sasha Luccioni, Jesse Dodge, and Margaret Mitchell. 2023. https://doi.org/10.18653/v1/2023.trustnlp-1.18 Detecting Personal Information in Training Corpora : an Analysis . In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing ( TrustN...
2023 doi
-
[86]
Zeerak Talat, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein. 2021. http://arxiv.org/abs/2101.11974 Disembodied Machine Learning : On the Illusion of Objectivity in NLP . arXiv:2101.11974 [cs]. ArXiv: 2101.11974
2021 arXiv
-
[87]
Bennett, and Min-Yen Kan
Samson Tan, Shafiq Joty, Kathy Baxter, Araz Taeihagh, Gregory A. Bennett, and Min-Yen Kan. 2021. https://doi.org/10.18653/v1/2021.acl-long.321 Reliability Testing for Natural Language Processing Systems . In Proceedings of the 59th Annual Meeting of the Association for Computa...
2021 doi
-
[88]
Ungless, Nikolas Vitsakis, Zeerak Talat, James Garforth, Björn Ross, Arno Onken, Atoosa Kasirzadeh, and Alexandra Birch
Eddie L. Ungless, Nikolas Vitsakis, Zeerak Talat, James Garforth, Björn Ross, Arno Onken, Atoosa Kasirzadeh, and Alexandra Birch. 2024. https://doi.org/10.48550/arXiv.2410.19812 Ethics Whitepaper : Whitepaper on Ethical Research into Large Language Models . arXiv preprint. ArX...
-
[89]
Levent Uzun. 2023. https://doi.org/10.1007/s44206-023-00070-2 Are Concerns Related to Artificial Intelligence Development and Use Really Necessary : A Philosophical Discussion . Digital Society, 2(3):40
2023 doi
-
[90]
Ibo Van de Poel. 2020. Embedding values in artificial intelligence (ai) systems. Minds and machines, 30(3):385--409
2020
-
[91]
Maggie Walter, Raymond Lovett, Bobby Maher, Bhiamie Williamson, Jacob Prehn, Gawaian Bodkin-Andrews, and Vanessa Lee. 2021. https://doi.org/10.1002/ajs4.141 Indigenous Data Sovereignty in the Era of Big Data and Open Data . Australian Journal of Social Issues, 56(2):143--156. ...
2021 doi
- [92]
-
[93]
Laura Weidinger, John F. J. Mellor, M. Rauh, C. Griffin, J. Uesato, Po-Sen Huang, M. Cheng, Mia Glaese, B. Balle, A. Kasirzadeh, Z. Kenton, S. Brown, W. Hawkins, T. Stepleton, C. Biles, A. Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick...
2021
-
[94]
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, et al. 2023. Sociotechnical safety evaluation of generative ai systems. arXiv preprint arXiv:2310.11986
2023 arXiv
-
[95]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al. 2022. Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability,...
2022
-
[96]
ai supply chain
David Gray Widder and Dawn Nafus. 2023. Dislocated accountabilities in the “ai supply chain”: Modularity and developers’ notions of responsibility. Big Data & Society, 10(1):20539517231177620
2023
-
[97]
Langdon Winner. 1980. https://www.jstor.org/stable/20024652 Do Artifacts Have Politics ? Daedalus, 109(1):121--136. Publisher: The MIT Press
1980
-
[98]
Wong, Michael A
Richmond Y. Wong, Michael A. Madaio, and Nick Merrill. 2023. https://doi.org/10.1145/3579621 Seeing Like a Toolkit : How Toolkits Envision the Work of AI Ethics . Proceedings of the ACM on Human-Computer Interaction, 7(CSCW1):1--27
2023 doi
-
[99]
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. 2021. https://doi.org/10.18653/v1/2021.naacl-main.190 Detoxifying Language Models Risks Marginalizing Minority Voices . In Proceedings of the 2021 Conference of the North American Chapter of...
2021 doi
-
[100]
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli. 2024. Hallucination is inevitable: An innate limitation of large language models. arXiv preprint arXiv:2401.11817
2024 arXiv
-
[101]
Jiancheng Yang, Hongwei Bran Li, and Donglai Wei. 2023. The impact of chatgpt and llms on medical imaging stakeholders: perspectives and use cases. Meta-Radiology, page 100007
2023
-
[102]
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang. 2024. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, page 100211
2024
-
[103]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[104]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.