REVIEW 3 major objections 6 minor 12 references
My LLM might Mimic AAE -- But When Should it?
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Black Americans judge LLM-generated African American English as authentic as human speech, and want the choice of when it appears.
desk verdict A useful community-based evaluation of LLM AAE, but the 'on par' conclusion is overstated relative to the paper's own numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is a paired continuation evaluation: human-transcribed prefixes (from CORAAL, Twitter, and NPR) are completed either by a human or by an LLM prompted with in-context examples, and Black American annotators rate the suffix on six Likert scales (coherence, AAE features, Black-sounding, White-sounding, mocking, offensive). The in-context prompting using CORAAL ground truth as chat history is what allowed the models to produce coherent AAE rather than refusal or off-topic output.
What would settle it
A direct adversarial test would be to have Black American annotators rate LLM-generated AAE against naturally occurring casual AAE speech (not interview transcripts) in matched contexts; if annotators judge the LLM output as significantly less authentic or more stereotyped than natural speech, the paper's parity claim would be falsified.
Extended reading notes
Core claim
The paper's central discovery is that Black Americans perceive LLM-generated AAE as comparable in authenticity to transcribed Black American speech from the CORAAL corpus, and sometimes as more AAE-heavy or more "Black sounding" than the human baseline. This holds across three LLMs (GPT 4o-mini, Llama 3, Mixtral) on judgments of coherence, presence of AAE features, and sounding like a Black American. At the same time, the survey shows a clear contextual preference: formal or task-specific settings call for MUSE, while personal or casual settings are open to AAE, and users want the autonomy to choose.
Load-bearing premise
The claim that LLM AAE is "on par" with human AAE rests on treating the CORAAL interview transcripts as the ground-truth representative of authentic AAE; if those transcripts are not representative, such as because interview speech differs from casual AAE or contains transcription artifacts, the parity conclusion is weakened.
Editorial extensions
If this is right
- If LLMs can produce authentic AAE on demand, then user-facing assistants can offer AAE as a selectable mode without risking mockery, provided the user opts in.
- The results imply that defaulting to MUSE in formal contexts aligns with user expectations, so product designers should keep MUSE as default and make AAE an explicit user choice.
- The finding that LLM output was sometimes judged more "Black sounding" than human transcripts suggests a calibration target: models may be over-performing AAE features, and "on par" is not "identical."
- The lack of perceived offensiveness suggests that fears of LLM AAE automatically being minstrelsy are not supported by Black American judgments in this sample.
Reading between the lines
- The paper's parity result could be extended by testing LLM-generated AAE against spontaneous casual speech rather than interview transcripts, which may be a higher bar for authenticity.
- If Black Americans want AAE as an opt-in feature, designers should treat AAE generation as a user-controlled setting rather than a default, and may need a "dialect intensity" control to avoid over-AAE output.
- The survey's strong preference for MUSE in formal settings suggests that any product feature offering AAE should be paired with context controls, since use of AAE in formal contexts may expose users to linguistic discrimination.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a two-part study with Black American participants: a scenario-based survey (n=104) on when participants would want LLMs to use African American English (AAE), and an annotation study (n=228; 8,654 judgments) in which Black American annotators rated human and LLM continuations on coherence, AAE features, Black-sounding, White-sounding, mocking, and offensiveness. The LLMs evaluated are GPT-4o-mini, Llama-3-70B-Instruct, and Mixtral-8x7B, with human baselines drawn from CORAAL transcribed interviews, a Twitter AAE corpus, and NPR MUSE interviews. The paper finds that participants prefer MUSE in formal settings but want the option to use AAE in casual settings, and it claims that appropriately prompted LLM outputs are perceived as authentic AAE 'on par' with transcribed Black American speech while generally not being seen as mocking or offensive.
Significance. If the preference findings hold, the survey provides useful, community-grounded guidance for designing language technologies that give Black Americans control over when AAE is used. The annotation study is valuable as a participatory evaluation of LLM AAE output by the relevant speech community, and the authors are transparent about researcher positionality and data limitations. The release of data and code is a concrete strength. The main weakness is that the headline authenticity-parity claim is not actually established by the reported statistics on the most diagnostic dimensions, where LLM outputs significantly exceed the human AAE baseline; the paper needs either an equivalence-based analysis or a carefully qualified conclusion.
major comments (3)
- [§4.2.1, Table 3; Abstract; §5] The statement that LLM outputs have 'a level of AAE authenticity on par with transcripts of Black American speech' is not supported by the study's most diagnostic measures. For CORAAL, the AAE Features mean for GPT is 1.18 (p<0.001) versus the human mean of 0.18, and for Llama it is 0.86 (p<0.01); for Black Sounding, GPT's mean is 1.01 (p<0.01) versus the human mean of 0.39. These are significant differences on the exact dimensions that define AAE authenticity, and the paper itself notes in §5 that LLM outputs were 'often perceived as more AAE-heavy' than the baseline. The absence of significant differences on other dimensions (Coherence, White Sounding, Mocking, Offensive) is not evidence of parity, because a non-significant difference is not an equivalence test. The central claim should be reframed as 'at least as AAE-strong as the baseline, within a range annotators found acceptable,' or supported by an explicit equivalence or non-inferiority analysis with pre-specified bounds.
- [§4.2.1, Table 3; §1 Contribution 2] The claim that annotators 'did not consider them to be mocking or offensive' is contradicted by the Llama Tweets condition. In Table 3, the Mocking mean for Llama Tweets continuations is 0.14 (p<0.001) while the human Tweets baseline is -0.79; a mean above zero indicates slight agreement that the text sounds like mocking. The Offensive mean for Llama Tweets is -0.07 (p<0.001) versus the human baseline of -0.96, which is neutral rather than disagreement. The abstract and Contribution 2 need to be qualified to acknowledge this exception, and §5's statement that participants 'generally disagreed that the machine-generated text ... was offensive to or mocking of Black Americans' should explicitly exclude or discuss the Llama Tweets case.
- [§3.2, Table 3, Tweets columns] The human Tweets baseline is rated by the same annotators as lacking AAE features (µ=-0.57) and not Black-sounding (µ=-0.30), so it does not function as an 'authentic AAE' ground truth in the way the CORAAL baseline does. Because the Tweets human baselines are not perceived as AAE, comparisons of LLM Tweets continuations to these baselines cannot support the general claim that LLM output is on par with human AAE; they only show that LLM continuations are more AAE-like than a baseline that annotators already judged to be non-AAE. The paper should either restrict the parity claim to CORAAL or analyze the Tweets condition separately as a test of whether LLM output can exceed a weak AAE baseline.
minor comments (6)
- [§5, first paragraph] The phrase 'our the language in our human AAE corpus' should read 'the language in our human AAE corpus.'
- [§4.2, Results Analysis Approach] The text says 'significant' means a false discovery rate of 5% and then states that Bonferroni corrections were applied with 10 comparisons per judgment. FDR and Bonferroni are different multiplicity controls; please clarify which method was actually used for Tables 3 and 4 and whether the 48 between-corpus comparisons in Table 4 received a correction.
- [§A.5.1, §A.5.2, §A.5.3] The GPT instruction text contains the typo 'do not use of the strings' in three places; it should be 'do not use the strings.'
- [§4.2.1, final paragraph] The word 'equivalentally' should be 'equivalently.'
- [Figure 2] The annotation labels such as '2 - Strongest Agreement' should be explained in the caption, since the Likert scale is elsewhere described as ranging from -2 (Strongly Disagree) to +2 (Strongly Agree).
- [§3.2] The paper does not report inter-annotator agreement for the six Likert judgments; reporting a measure such as Krippendorff's alpha would strengthen confidence in the reliability of the aggregate scores.
Circularity Check
No significant circularity: the authenticity comparison rests on external human baselines and independent Black American annotator judgments, not on fitted parameters or self-referential definitions.
full rationale
This paper is an empirical study, not a derivation, and I found no step where a claimed result is equivalent to its input by construction. The central authenticity claim is that Black American annotators rated LLM-generated AAE continuations as on par with transcribed human AAE speech from CORAAL. That comparison is anchored in external data: CORAAL transcripts, an X/Twitter AAE corpus, and NPR MUSE interviews, with judgments collected from separate Black American annotators who were blind to whether a suffix was human- or LLM-produced (§3.2). No parameter is fitted to the annotation scores and then renamed as a prediction; the Likert ratings are raw elicited judgments. The survey preferences (§4.1) are likewise directly measured participant choices. The paper does cite Cunningham et al. (2024), which shares an author (Hal Daumé III), but that citation appears in related work as background on AAE speakers' labor with language technologies and is not load-bearing for the paper's own survey or annotation results. The Limitations section (§7) honestly flags that CORAAL is a transcription and may contain artifacts that affect perceived authenticity, and that annotators were not told whether text was human or machine generated; these are validity and interpretation concerns, not evidence of circularity. A skeptic could argue that the 'on par' conclusion is not fully supported because Table 3 shows statistically significant differences on AAE Features and Black Sounding for some LLMs relative to the CORAAL baseline, but that is an argument about inference from the data, not a circularity where the output is predetermined by the input. The study is self-contained against external benchmarks and its judgments come from independent community annotators, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Likert responses are treated as interval data for computing means and t-tests.
- domain assumption CORAAL transcribed interviews represent authentic AAE.
- domain assumption The Prolific sample is reasonably representative of Black Americans with AAE familiarity.
- domain assumption Annotating single exchange pairs captures authentic language use.
Cite this review
Pith. "Pith review of My LLM might Mimic AAE -- But When Should it?." pith.science (2026). https://pith.science/paper/5OKXBLQD
@misc{pith2026250204564,
author = {Pith},
title = {Pith review of: My LLM might Mimic AAE -- But When Should it?},
year = {2026},
howpublished = {\url{https://pith.science/paper/5OKXBLQD}},
note = {Machine review of arXiv:2502.04564}
}
abstract
We examine the representation of African American English (AAE) in large language models (LLMs), exploring (a) the perceptions Black Americans have of how effective these technologies are at producing authentic AAE, and (b) in what contexts Black Americans find this desirable. Through both a survey of Black Americans ($n=$ 104) and annotation of LLM-produced AAE by Black Americans ($n=$ 228), we find that Black Americans favor choice and autonomy in determining when AAE is appropriate in LLM output. They tend to prefer that LLMs default to communicating in Mainstream U.S. English in formal settings, with greater interest in AAE production in less formal settings. When LLMs were appropriately prompted and provided in context examples, our participants found their outputs to have a level of AAE authenticity on par with transcripts of Black American speech. Select code and data for our project can be found here: https://github.com/smelliecat/AAEMime.git
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[5]
Exploring the role of grammar and word choice in bias toward African American English (AAE) in hate speech classifica- tion. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 789–798. Jazmia Henry
work page 2022
-
[7]
Dialect prejudice predicts AI decisions about people’s character, employability, and criminality. Preprint, arXiv:2403.00742. Yolanda Holt
-
[8]
Mix- tral of experts. Preprint, arXiv:2401.04088. Tyler Kendall and Charlie Farrington
-
[10]
Lever- aging syntactic dependencies in disambiguation: The case of African American English. In Proceedings of the 2024 Joint International Conference on Compu- tational Linguistics, Language Resources and Evalu- ation (LREC-COLING 2024), pages 10403–10415, Torino, Italia. ELRA and ICCL. John R Rickford, Greg J Duncan, Lisa A Gennetian, Ray Yun Gou, Rebec...
work page 2024
-
[11]
Disambiguation of morpho-syntactic features of African American English -- the case of habitual be
Disambiguation of morpho-syntactic features of African American English–the case of habitual be. arXiv preprint arXiv:2204.12421. Hanna L Smokoski
-
[12]
Llamafac- tory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372. A Appendix A.1 Preparation of the CORAAL Corpus (Black American transcribed interviews) Our prompt texts or prefixes needed to have authentic AAE to the degree possible and cover a broad range of the different variations of AAE spoken in the wild. (Lane...
arXiv 2015
-
[2016]
arXiv preprint arXiv:1608.08868
Demographic dialectal variation in social media: A case study of African-American English. arXiv preprint arXiv:1608.08868. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al
-
[2019]
Interview: A Large-Scale Open-Source Corpus of Media Dialog
AI- Based Digital Assistants. Business & Information Systems Engineering, 61:535–544. Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, and Julian McAuley. 2020a. Large-scale modeling of media dialog with discourse patterns and knowledge grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process- ing (EMNLP). *Equa...
work page Pith review arXiv 2020
Show all 12 references
-
[2020]
Advances in neural information processing systems, 33:1877–1901
Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anasta- sios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica
1901
-
[2021]
Accessed: 2024- 06-01
Aave corpora. Accessed: 2024- 06-01. Jane H Hill
2024
-
[2022]
It’s Kind of Like Code-Switching
“It’s Kind of Like Code-Switching”: Black Older Adults’ Experiences with a V oice Assistant for Health Information Seek- ing. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY , USA. Association for Computing Machinery. Cami...
2022
-
[2024]
Preprint, arXiv:2403.04132
Chatbot arena: An open platform for evaluating LLMs by human prefer- ence. Preprint, arXiv:2403.04132. Jay L. Cunningham, Su Lin Blodgett, Hal Daumé III, Christina Harrington, Hanna Wallach, and Michael Madaio
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.