REVIEW 2 major objections 5 minor 91 references
A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper presents a five-step, query-only framework for auditing LLM-based chatbots for dialect bias, and demonstrates it on Amazon Rufus, showing that the chatbot produces lower-quality responses to prompts in minoritized English…
desk verdict A useful, honest framework paper whose case study is suggestive rather than decisive. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework is a five-step procedure: identify the target chatbot, collect realistic user prompts, perturb those prompts across dialects (and optionally across other facets such as typos), repeatedly query the chatbot, and evaluate response quality across dialect groups. The dialect perturbation step relies on Multi-VALUE, a rule-based transformation system whose linguistic rules and feature probabilities come from eWAVE, an atlas of 235 grammatical features across 51 English dialects; the case study also uses a second-stage perturbation that lowercases prompts, removes punctuation, and inserts common typos. The central measurement is the classification of each response as unsure or incorrect, defined in the customer-service context, with statistical significance assessed through paired t-tests and ANOVA under Benjamini-Hochberg multiple-testing correction.
What would settle it
Recruit native speakers of AAE, SgE, IndE, AppE, and ChcE to rate each transformed prompt for grammaticality and semantic equivalence to the SAE original, then rerun the Rufus audit using only prompts that all raters judge valid; if the quality gap between dialect and SAE prompts disappears or shrinks to near zero, the measured harm was an artifact of corrupted prompts rather than dialect bias.
Extended reading notes
Core claim
The authors claim to have established that a widely deployed LLM-based chatbot, Amazon Rufus, exhibits dialect-based quality-of-service harms: responses to prompts written in minoritized English dialects are significantly more likely to be unsure or factually incorrect than responses to semantically equivalent standard-American-English prompts, with the gap widening when typos are present. They further identify a concrete failure pattern: prompts using the zero copula (a grammatical feature common in African American English and Singaporean English, as when "Is this jacket machine washable?" becomes "This jacket machine washable?") caused Rufus to abandon the product page and run an unrelated product search 69% of the time, versus 6% for otherwise identical prompts with the copula. In the same demonstration, a programmatically queryable 'copy' of Rufus built from an extracted system prompt and a base LLM replicated the real chatbot's pattern of lower quality on minoritized dialects with over 90% agreement in unsureness, suggesting that such copies can be useful proxies when direct access is unavailable.
Load-bearing premise
The audit's conclusions rest on the assumption that the automated dialect transformations alter only genuine dialect features, leaving every prompt grammatically correct in its target dialect and semantically equivalent to its standard-English original.
Editorial extensions
If this is right
- Dialect bias documented in base LLMs does propagate to production chatbots, so application-level audits are necessary and cannot be replaced by model audits alone.
- Because the framework requires only query access, external auditors, advocacy groups, and individual users can hold chatbot providers accountable without internal cooperation.
- Typos are not a minor nuisance but an amplifier of dialect-based harm, so robustness evaluations for chatbots should include realistic noisy input alongside dialect variation.
- The zero-copula failure mode identifies a concrete, testable target for improvement: training or retrieval systems could be made robust to omission of the verb 'to be', a well-documented AAE and SgE grammar rule.
- The copy-audit approach, which agreed with the real chatbot on unsureness in over 90% of cases, suggests that cheaper programmatic audits of reconstructed chatbots can approximate full-system audits when direct access is infeasible.
Reading between the lines
- If the zero-copula failure is the main driver of the incorrectness gap, the likely locus is the retrieval-and-indexing layer that maps product-page text onto the question, not just the base LLM; a follow-up audit with retrieval disabled could localize the mechanism.
- The same framework could be extended to non-English languages and to other real-world input facets such as emojis, voice-to-text errors, or code-switching, which the paper names but does not test.
- The authors' critique of the ASPD dataset implies that some previously published estimates of LLM dialect bias may be inflated by formality and profanity confounds; re-analyzing those results with formality-matched stimuli would test that implication.
- The framework's reliance on a single transformation tool, Multi-VALUE, means the case-study magnitudes should be interpreted as lower bounds on real-world harm if human-written dialect prompts differ from rule-based transformations in ways speakers would find natural.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a framework for auditing LLM-based chatbots for dialect-based quality-of-service harms, defined as systems not working equally well for users of different dialects. The framework consists of five steps: target identification, prompt collection, prompt perturbation, response collection, and response evaluation. It emphasizes dynamically generated, semantically equivalent prompts across dialects, measurement of task-relevant response quality rather than representational harm, and a black-box query-only access requirement. The authors motivate the framework with a critical review of prior dialect-bias audit methods, particularly the AAVE/SAE Paired Dataset, and demonstrate it in a case study of Amazon Rufus, a commercial customer-service chatbot. In the case study, prompts sourced from Amazon's own suggestions were transformed into five minoritized English dialects using Multi-VALUE and further perturbed with lowercasing, punctuation removal, and typos. The authors report that Rufus produces significantly more unsure and incorrect responses for prompts in African American English, Indian English, and Singaporean English, and that typos exacerbate this gap. They also report results from a 'copy' audit using GPT-4o-mini with a system prompt purportedly extracted from Rufus.
Significance. If the findings hold, the paper makes a valuable contribution: it shifts dialect-bias auditing from model-level representational harms to application-level quality-of-service harms, which is better aligned with how users actually experience LLM-based chatbots. The framework's query-only requirement, use of dynamically generated prompts, and support for multi-turn and typo-perturbed inputs are practical design choices that lower the barrier for external auditors. The paper also ships open perturbation code, and the case study audits a real, widely deployed commercial chatbot rather than a research LLM, which is an important step for external accountability. The candid treatment of limitations in §6.1, including the fundamental mismatch between generated text and identity and the acknowledged tradeoffs in implementation, is commendable. The central empirical claim, however, is currently supported only by a case study whose dialect-manipulation step lacks the validation that the framework itself recommends, so the strength of the stated conclusions is disproportionate to the evidence.
major comments (2)
- [§5 and §6.1] The case study's central claim—that Rufus produces lower-quality responses to prompts written in minoritized dialects—depends on the Multi-VALUE transforms being semantically equivalent to the SAE originals and grammatically natural in the target dialect. The framework itself recommends in §4.3 that auditors validate transformed prompts with native speakers, but the case study does not perform any such validation on its own prompt set. §6.1 reports that 'many ChcE prompts had zero perturbations and thus matched the SAE originals' and that 'some AppE prompts initially had too many perturbations' and were regenerated based on the authors' own judgment of semantic equivalence. Without an independent check that the actual prompts read as authentic dialect rather than corrupted or stilted text, the observed quality gaps for AAE, IndE, and SgE could reflect a general response to ungrammatical input rather than dialect-specific quality-of-service harm. This is a load-bearing threat to the strongest empirical conclusion of the paper. I recommend either conducting a native-speaker validation of the prompt set (or a representative sample), or re-framing the case study as a demonstration of the framework's mechanics with the dialect-bias conclusion made explicitly conditional on such validation.
- [§5 and §6.1] The analysis does not measure or control for transformation intensity. Multi-VALUE applies features with probabilities of 100%, 60%, or 30% depending on eWAVE feature frequency, and dialects differ substantially in the number of listed features (e.g., AppE has 76 features occurring at least at 'rare' frequency versus ChcE's 32). This means prompts in different dialects receive varying densities of perturbations. The authors acknowledge this in §6.1 but do not include any measure of the number of transformations per prompt in the statistical analysis, nor do they test whether the dialect effect is robust to controlling for transformation intensity. Adding such a control or a sensitivity analysis would substantially strengthen the causal interpretation by addressing the alternative explanation that higher transformation density, rather than dialect membership per se, drives the quality gap.
minor comments (5)
- [§5] The statistical reporting is internally inconsistent: §5 states that paired t-tests with the Benjamini-Hochberg procedure were used, while the caption of Figure 5 reports that 'Results vary statistically significantly across dialect as well as across the combination of dialect and formality per ANOVA tests.' Please clarify the actual test design, including the pairing structure, and report effect sizes or confidence intervals so readers can assess the magnitude of the effects.
- [§5] The zero-copula comparison reports '6% of the time (6/108 prompts)' for prompts with the copula and '69% of the time (25/36 prompts)' for otherwise identical prompts with zero copula, but the denominators are not explained. Clarify how many prompts and which dialects contribute to each denominator, and how the 'otherwise identical' condition was constructed.
- [§6.1] Given that some ChcE prompts were identical to their SAE originals and some AppE prompts were regenerated, please report the number of prompts that were identical to the SAE baseline and the number that were discarded/regenerated. A robustness check that excludes or separately analyzes these prompts would help establish that the results are not driven by these problematic cases.
- [§5] The fidelity assessment of the copy chatbot reports '>90% agreement' with real Rufus on unsureness, but neither the sample size nor the way agreement was computed is specified. State the number of prompts used for the comparison and consider reporting a chance-corrected agreement metric (e.g., Cohen's kappa).
- [§5] There is a typo in the 'Copying Rufus' subsection: 'its base LLM is propriety' should be 'its base LLM is proprietary.'
Circularity Check
No significant circularity: the dialect-bias findings rest on new measurements with external transformation tools, not on a self-referential definition or fitted parameter.
full rationale
I walked the paper's derivation chain from framework design to the Rufus case study. The framework's output (unsureness and incorrectness across dialects) is an operationalization of quality-of-service harms, not a quantity defined by the authors' prior results; the case study then measures it directly by manually querying Rufus and manually annotating responses. Prompt variation is produced by Multi-VALUE/eWAVE, an externally built, linguistically grounded tool, not by a parameter fitted to the outcome. The statistical findings reported in §5 (paired t-tests, ANOVA) are new empirical measurements. The self-citations in the paper ([24], [39], [40]) support background claims about accountability and about the risks of using LLMs for dialect transformation; those claims are also supported by external references and do not carry the quantitative result. The limitations acknowledged in §6.1 (stochastic transformation, zero-perturbation ChcE prompts, discarded AppE prompts, lack of re-validation) are measurement-validity concerns that an external critic could use to question whether the gap reflects dialect rather than text corruption, but they do not make the derivation circular: the dialect factor is not constructed from the measured quality gap, nor is the conclusion forced by definition. No equation or fitted parameter is renamed as a prediction.
Assumptions & free parameters
free parameters (1)
- Typo insertion frequency =
unspecified
assumptions (5)
- domain assumption Multi-VALUE generated prompts are semantically equivalent to the SAE originals and grammatically correct in each target dialect.
- domain assumption eWAVE heuristic probabilities (100%, 60%, 30%) for feature incorporation produce representative dialect text.
- domain assumption Manual annotation of unsureness and incorrectness is valid and consistent.
- domain assumption Single queries per prompt suffice because Rufus responses are deterministic.
- standard math Paired t-tests, ANOVA, and Benjamini-Hochberg correction are appropriate for comparing response quality across dialects.
Cite this review
Pith. "Pith review of A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms." pith.science (2026). https://pith.science/paper/7CBJBS24
@misc{pith2026250604419,
author = {Pith},
title = {Pith review of: A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms},
year = {2026},
howpublished = {\url{https://pith.science/paper/7CBJBS24}},
note = {Machine review of arXiv:2506.04419}
}
read the original abstract
Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses when prompted with text written in minoritized dialects. However, whether and how this bias propagates to systems built on top of LLMs, such as chatbots, is still unclear. We conduct a review of existing approaches for auditing LLMs for dialect bias and show that they cannot be straightforwardly adapted to audit LLM-based chatbots due to issues of substantive and ecological validity. To address this, we present a framework for auditing LLM-based chatbots for dialect bias by measuring the extent to which they produce quality-of-service harms, which occur when systems do not work equally well for different people. Our framework has three key characteristics that make it useful in practice. First, by leveraging dynamically generated instead of pre-existing text, our framework enables testing over any dialect, facilitates multi-turn conversations, and represents how users are likely to interact with chatbots in the real world. Second, by measuring quality-of-service harms, our framework aligns audit results with the real-world outcomes of chatbot use. Third, our framework requires only query access to an LLM-based chatbot, meaning that it can be leveraged equally effectively by internal auditors, external auditors, and even individual users in order to promote accountability. To demonstrate the efficacy of our framework, we conduct a case study audit of Amazon Rufus, a widely-used LLM-based chatbot in the customer service domain. Our results reveal that Rufus produces lower-quality responses to prompts written in minoritized English dialects, and that these quality-of-service harms are exacerbated by the presence of typos in prompts.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
A Language Is a Dialect with an Army and a Navy
1997. A Language Is a Dialect with an Army and a Navy. Language in Society 26, 3 (1997), 469–469. http://www.jstor.org/stable/4168793
-
[2]
Greenhouse Gas Emissions from a Typical Passenger Vehicle
2025. Greenhouse Gas Emissions from a Typical Passenger Vehicle. United States Environmental Protection Agency. https://www.epa.gov/ greenvehicles/greenhouse-gas-emissions-typical-passenger-vehicle
2025
-
[3]
Khan Academy. 2024. Why We’re Deeply Invested in Making AI Better at Math Tutoring (and What We’ve Been Up to Lately). https: //blog.khanacademy.org/why-were-deeply-invested-in-making-ai-better-at-math-tutoring-and-what-weve-been-up-to-lately/
2024
-
[4]
William Agnew, A. Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R. McKee. 2024. The Illusion of Artificial Inclusion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association for Computing Machinery, New York, NY, USA, Article 286, 12 ...
arXiv 2024
-
[5]
Stefan Baack. 2024. A Critical Analysis of the Largest Source for Generative AI Training Data: Common Crawl. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio de Janeiro, Brazil) (FAccT ’24). Association for Computing Machinery, New York, NY, USA, 2199–2208. https://doi.org/10.1145/3630106.3659033
arXiv 2024
-
[6]
Matt Barnum. 2024. We Tested an AI Tutor for Kids. It Struggled With Basic Math. The Wall Street Journal. https://www.wsj.com/tech/ai/ai-is- tutoring-students-but-still-struggles-with-basic-math-694e76d3
2024
-
[8]
Yoav Benjamini and Yosef Hochberg. 1995. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society. Series B (Methodological) 57, 1 (1995), 289–300. http://www.jstor.org/stable/2346101
arXiv 1995
-
[9]
Abeba Birhane, Ryan Steed, Victor Ojewale, Briana Vecchione, and Inioluwa Deborah Raji. 2024. AI auditing: The Broken Bus on the Road to AI Accountability. In 2nd IEEE Conference on Secure and Trustworthy Machine Learning (SATML). IEEE, Toronto, Canada. http://arxiv.org/abs/2401.14462 arXiv:2401.14462 [cs]
arXiv 2024
Show all 91 references
-
[10]
Abeba Birhane, vinay uday prabhu, Sanghyun Han, Vishnu Boddeti, and Sasha Luccioni. 2023. Into the LAION’s Den: Investigating Hate in Multimodal Datasets. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track . https://openreview. ...
2023
-
[12]
Su Lin Blodgett, Solon Barocas, Hal Daumé Iii, and Hanna Wallach. 2020. Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, O...
2020 doi
-
[13]
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016. Demographic Dialectal Variation in Social Media: A Case Study of African-American English. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Association for Computational...
2016 doi
-
[14]
Mark Bovens. 2007. Analysing and Assessing Accountability: A Conceptual Framework. European Law Journal 13, 4 (2007), 447–468. https: //doi.org/10.1111/j.1468-0386.2007.00378.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1468-0386.2007.00378.x
2007
-
[15]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020 arXiv
-
[16]
2005.Language and Identity
Mary Bucholtz and Kira Hall. 2005.Language and Identity. John Wiley & Sons, Ltd, Chapter 16, 369–394. https://doi.org/10.1002/9780470996522.ch16 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/9780470996522.ch16
2005 doi
-
[17]
Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (FAT*) . PMLR, 77–91. https://proceedings.mlr.press/v81/buolamwini18a.html
2018
-
[18]
Wallace Chafe and Deborah Tannen. 1987. The Relation between Written and Spoken Language. Annual Review of Anthropology 16 (1987), 383–407. http://www.jstor.org/stable/2155877
1987
-
[19]
Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan ...
2023 doi
-
[20]
Myra Cheng, Tiziano Piccardi, and Diyi Yang. 2023. CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association fo...
2023 doi
-
[21]
Trishul Chilimbi. 2024. How We Built Rufus, Amazon’s AI-Powered Shopping Assistant. IEEE Spectrum. https://spectrum.ieee.org/amazon-rufus
2024
- [22]
-
[23]
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, a...
2017
-
[24]
Sasha Costanza-Chock, Emma Harvey, Inioluwa Deborah Raji, Martha Czernuszenko, and Joy Buolamwini. 2022. Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem. In 2022 ACM Conference on Fairness, Accountability, and Transparency (FAcc...
2022
-
[25]
Crichton
Scott J. Crichton. 2017. STATE OF LOUISIANA VERSUS WARREN DEMESME. https://www.lasc.org/opinions/2017/17KK0954.sjc.addconc.pdf
2017
-
[26]
Tony Crowley. 2006. The Political Production of a Language. Journal of Linguistic Anthropology 16, 1 (2006), 23–35. https://doi.org/10.1525/jlin. 2006.16.1.023 arXiv:https://anthrosource.onlinelibrary.wiley.com/doi/pdf/10.1525/jlin.2006.16.1.023
2006 doi
-
[27]
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. 2021. Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language Technologies. InProceedings of the 2021 Conference on Empirical Methods in Natural La...
2021
-
[28]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Docu- menting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. InProceedings of the 2021 Conference on Empirical Method...
2021
-
[29]
Franz Faul, Edgar Erdfelder, Albert-Georg Lang, and Axel Buchner. 2007. G*Power 3: a flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav. Res. Methods 39, 2 (May 2007), 175–191
2007
-
[30]
Eve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi, Xavier Yin, and Dan Klein. 2024. Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination. http://arxiv.org/abs/2406.08818 arXiv:2406.08818 [cs]
2024 arXiv
-
[31]
Geoffrey A. Fowler. 2024. TurboTax and H&R Block now use AI for tax advice. It’s awful. The Washington Post. https://www.washingtonpost.com/ technology/2024/03/04/ai-taxes-turbotax-hrblock-chatbot/
2024
-
[32]
Susan Gal and Judith T. Irvine. 1995. The Boundaries of Languages and Disciplines: How Ideologies Construct Difference. Social Research 62, 4 (1995), 967–1001. http://www.jstor.org/stable/40971131
1995
-
[33]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2023. Bias and Fairness in Large Language Models: A Survey. http://arxiv.org/abs/2309.00770 arXiv:2309.00770 [cs]
2023 arXiv
-
[34]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevor Cohn, Yulan He, and Yang Liu (Eds.)...
2020 doi
-
[35]
Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020. Investigating African- American Vernacular English in Transformer-Based Text Generation. InProceedings of the 2020 Conference on Empirical Methods in Natural Lang...
2020 doi
-
[36]
Jeffrey Grogger. 2011. Speech Patterns and Racial Wage Inequality. Journal of Human Resources 46, 1 (2011), 1–25. https://doi.org/10.3368/jhr.46.1.1 arXiv:https://jhr.uwpress.org/content/46/1/1.full.pdf
2011 doi
-
[38]
Kenneth R Hammond. 1998. Ecological validity: Then and now
1998
-
[39]
Emma Harvey, Allison Koenecke, and Rene F Kizilcec. 2024. Towards an Educator-Centered Method for Measuring Bias in Large Language Model-Based Chatbot Tutors. In AI for Education: Bridging Innovation and Responsibility at the 38th AAAI Annual Conference on AI . https: 18 Harve...
2024
-
[40]
Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, and Hanna Wallach. 2024. Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems. arXiv:2411.15662 [cs.CY] https://arxiv.or...
2024 arXiv
-
[41]
Einar Haugen. 1996. Dialect, Language, Nation. American Anthropologist 68, 4 (Aug. 1996), 922 – 935. https://doi.org/10.1525/aa.1966.68.4.02a00040
1996 doi
-
[42]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect. Nature 633, 8028 (01 Sep 2024), 147–154. https://doi.org/10.1038/s41586-024-07856-5
2024 doi
-
[43]
Holleman, Ignace T
Gijs A. Holleman, Ignace T. C. Hooge, Chantal Kemner, and Roy S. Hessels. 2020. The ‘Real-World Approach’ and Its Problems: A Critique of the Term Ecological Validity. Frontiers in Psychology 11 (2020). https://doi.org/10.3389/fpsyg.2020.00721
2020
-
[44]
Nanna Inie, Jeanette Falk, and Raghavendra Selvan. 2025. How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About It. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25) . Association for Comput...
2025
-
[45]
Jacobs and Hanna Wallach
Abigail Z. Jacobs and Hanna Wallach. 2021. Measurement and Fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT). Association for Computing Machinery, New York, NY, USA, 375–385. https://doi.org/10.1145/3442188.3445901
2021
-
[46]
errors,
David Johnson and Lewis VanBrackle. 2012. Linguistic discrimination in writing assessment: How raters react to African American “errors, ” ESL errors, and standard English errors on a state-mandated writing exam.Assessing Writing 17, 1 (2012), 35–54. https://doi.org/10.1016/j....
2012 doi
-
[47]
Taylor Jones, Jessica Rose Kalbfeld, Ryan Hancock, and Robin Clark. 2019. Testifying while black: An experimental study of court reporter accuracy in transcription of African American English. Language 95, 2 (2019), e216–e252
2019
-
[48]
Ecological Validity
John F. Kihlstrom. 2021. Ecological Validity and “Ecological Validity”. Perspectives on Psychological Science 16, 2 (2021), 466–471. https: //doi.org/10.1177/1745691620966791 arXiv:https://doi.org/10.1177/1745691620966791 PMID: 33593121
2021 doi
-
[50]
Bernd Kortmann, Kerstin Lunkenheimer, and Katharina Ehret (Eds.). 2020. eW A VE. https://ewave-atlas.org/
2020
-
[51]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in Large Language Models. In Proceedings of The ACM Collective Intelligence Conference (Delft, Netherlands) (CI ’23). Association for Computing Machinery, New York, NY, USA, 12–24. https://doi.org/10....
2023
-
[52]
Kretzschmar and Charles F
William A. Kretzschmar and Charles F. Meyer. 2012. The idea of Standard American English . Cambridge University Press, 139–158
2012
-
[53]
Colin Lecher. 2024. NYC’s AI Chatbot Tells Businesses to Break the Law. The Markup. https://themarkup.org/news/2024/03/29/nycs-ai-chatbot- tells-businesses-to-break-the-law
2024
-
[54]
Are you accepting new patients?
Tamara G.J. Leech, Amy Irby-Shasanmi, and Anne L. Mitchell. 2019. “Are you accepting new patients?” A pilot field experiment on telephone-based gatekeeping and Black patients’ access to pediatric care. Health Services Research 54, S1 (2019), 234–242. https://doi.org/10.1111/14...
2019
-
[55]
Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman
Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A New Generation of Perspective API: Efficient Multilingual Character-level Transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data ...
2022
-
[56]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in ...
2020
-
[57]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL] https://arxiv.org/abs/1907.11692
2019 arXiv
-
[58]
Alexandra Luccioni and Joseph Viviano. 2021. What‘s in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Langu...
2021 doi
-
[59]
Hanjia Lyu, Jiebo Luo, Jian Kang, and Allison Koenecke. 2025. Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (Athens, Greece) (FAccT ’25). ...
2025
-
[60]
Rajiv Mehta. 2024. How customers are making more informed shopping decisions with Rufus, Amazon’s generative AI-powered shopping assistant. Amazon. https://www.aboutamazon.com/news/retail/how-to-use-amazon-rufus
2024
-
[61]
Rajiv Mehta and Trishul Chilimbi. 2024. Amazon announces Rufus, a new generative AI-powered conversational shopping experience. Amazon. https://www.aboutamazon.com/news/retail/amazon-rufus
2024
-
[62]
Pepper Miller and Kristen DiCerbo. 2024. LLM Based Math Tutoring: Challenges and Dataset. https://doi.org/10.35542/osf.io/5zwv3
2024 doi
-
[63]
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023. Auditing large language models: a three-layered approach. AI and Ethics (May 2023). https://doi.org/10.1007/s43681-023-00289-2 A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms 19
2023 doi
-
[64]
Roberto Navigli, Simone Conia, and Björn Ross. 2023. Biases in Large Language Models: Origins, Inventory, and Discussion. J. Data and Information Quality 15, 2, Article 10 (June 2023), 21 pages. https://doi.org/10.1145/3597307
2023 doi
-
[65]
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 6464 (Oct. 2019), 447–453. https://doi.org/10.1126/science.aax2342
2019 doi
-
[66]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[67]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[68]
Thomas Purnell, William Idsardi, and John Baugh. 1999. Perceptual and Phonetic Experiments on American English Dialect Identification. Journal of Language and Social Psychology 18, 1 (1999), 10–30. https://doi.org/10.1177/0261927X99018001002 arXiv:https://doi.org/10.1177/02619...
1999 doi
-
[69]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving Language Understanding by Generative Pre-Training. https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf
2018
-
[70]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
2019
-
[71]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research 21, 140 (2020), 1–67. htt...
2020
-
[72]
Inioluwa Deborah Raji and Joy Buolamwini. 2019. Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (AIES) . Association for Computing M...
2019
-
[76]
Alene Rhea, Kelsey Markey, Lauren D’Arinzo, Hilke Schellmann, Mona Sloane, Paul Squires, and Julia Stoyanovich. 2022. Resume Format, LinkedIn URLs and Other Unexpected Influences on AI Personality Prediction in Hiring: Results of an Audit. In Proceedings of the 2022 AAAI/ACM C...
2022
-
[77]
Shalaleh Rismani, Renee Shelby, Andrew Smart, Edgar Jatho, Joshua Kroll, AJung Moon, and Negar Rostamzadeh. 2023. From Plane Crashes to Algorithmic Harm: Applicability of Safety Engineering Frameworks for Responsible ML. In Proceedings of the 2023 CHI Conference on Human Facto...
2023
-
[78]
Data and Discrimination: Converting Critical Concerns into Productive Inquiry,
Christian Sandvig, Kevin Hamilton, K. Karahalios, and Cédric Langbort. 2014. Auditing Algorithms: Research Methods for Detecting Discrimination on Internet Platforms. In "Data and Discrimination: Converting Critical Concerns into Productive Inquiry, ” a preconference at the 64...
2014
-
[79]
Rogers, Inna Goncearenco, Giuseppe Sarli, Igor Galynker, Denis Peskoff, Marine Carpuat, Jules White, Shyamal Anadkat, Alexander Hoyle, and Philip Resnik
Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, Pranav Sandeep Dulepet, Saurav Vidyadhara, Dayeon Ki, Sweta Agrawal, Chau Pham, Gerson Kroiz, Feileen Li, Hudson Tao, Ashay S...
2024 arXiv
-
[80]
Andrew D Selbst and Solon Barocas. 2023. Unfair Artificial Intelligence: How FTC Intervention Can Overcome the Limitations of Discrimination Law. University of Pennsylvania Law Review 171 (2023)
2023
-
[81]
Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Rostamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023. Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction. In Proceedin...
2023
-
[82]
Hong Shen, Alicia DeVos, Motahhare Eslami, and Kenneth Holstein. 2021. Everyday Algorithm Auditing: Understanding the Power of Everyday Users in Surfacing Harmful Algorithmic Behaviors. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021), 1–29. https: //d...
2021 doi
-
[83]
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The Woman Worked as a Babysitter: On Biases in Language Generation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on N...
2019 doi
-
[84]
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng. 2021. Societal Biases in Language Generation: Progress and Challenges. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural...
2021 doi
-
[85]
Smitherman
G. Smitherman. 1986. Talkin and Testifyin: The Language of Black America . Wayne State University Press. https://books.google.com/books?id= HXD7pYv80bUC
1986
-
[86]
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2023. Dialect-robust Evaluation of Generated Text. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguist...
2023 doi
-
[87]
Latanya Sweeney. 2013. Discrimination in online ad delivery. Commun. ACM 56, 5 (May 2013), 44–54. https://doi.org/10.1145/2447976.2447990
2013
-
[88]
Briana Vecchione, Karen Levy, and Solon Barocas. 2021. Algorithmic Auditing and Social Justice: Lessons from the History of Audit Studies. In Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO) . ACM, – NY USA, 1–9. https://doi.org/10.1145/3465416.3483294
2021
-
[89]
Pranav Narayanan Venkit, Mukund Srinath, and Shomir Wilson. 2022. A Study of Implicit Bias in Pretrained Language Models against People with Disabilities. In Proceedings of the 29th International Conference on Computational Linguistics , Nicoletta Calzolari, Chu-Ren Huang, Han...
2022
-
[90]
Dickerson
Angelina Wang, Jamie Morgenstern, and John P. Dickerson. 2025. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence (17 Feb 2025). https://doi.org/10.1038/s42256-025-00986-z
2025 doi
-
[91]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Courtney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, William...
2022
-
[92]
Marcia Farr Whiteman. 2013. Dialect influence in writing. In Writing. Routledge, 153–166. A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms 21
2013
-
[93]
Maranke Wieringa. 2020. What to account for when accounting for algorithms: a systematic literature review on algorithmic accountability. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT*) . Association for Computing Machinery, New York,...
2020
-
[94]
Nathan Matias
Lucas Wright, Roxana Mika Muenster, Briana Vecchione, Tianyao Qu, Pika (Senhuang) Cai, Alan Smith, Comm 2450 Student Investigators, Jacob Metcalf, and J. Nathan Matias. 2024. Null Compliance: NYC Local Law 144 and the challenges of algorithm accountability. In The 2024 ACM Con...
2024
-
[95]
Meg Young, Michael Katell, and P.M. Krafft. 2022. Confronting Power and Corporate Capture at the FAccT Conference. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing Machiner...
2022
-
[96]
Jiahao Yu, Yuhang Wu, Dong Shu, Mingyu Jin, Sabrina Yang, and Xinyu Xing. 2024. Assessing Prompt Injection Risks in 200+ Custom GPTs. InICLR 2024 Workshop on Secure and Trustworthy Large Language Models (ICLR Workshops) . arXiv. http://arxiv.org/abs/2311.11538 arXiv:2311.11538 [cs]
2024 arXiv
-
[97]
Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2020. Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593 [cs.CL] https://arxiv.org/abs/1909.08593
2020 arXiv
-
[98]
Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. Multi-VALUE: A Framework for Cross-Dialectal English NLP. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (ACL) . Assoc...
2023 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.