Pith. sign in

REVIEW 3 major objections 6 minor 12 references

My LLM might Mimic AAE -- But When Should it?

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Black Americans judge LLM-generated African American English as authentic as human speech, and want the choice of when it appears.

desk verdict A useful community-based evaluation of LLM AAE, but the 'on par' conclusion is overstated relative to the paper's own numbers. read the letter →

arxiv 2502.04564 v2 pith:5OKXBLQD submitted 2025-02-06 cs.CL

classification cs.CL
keywords AfricanAmericanEnglishlargelanguagemodelsAAEauthenticityuserpreferencesin-contextlearninglinguisticjudgmentsdialectprejudiceBlackperspectives
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether Black Americans want large language models to produce African American English (AAE) and whether current models do it well. Through a survey of 104 Black Americans and annotation of LLM outputs by 228 Black Americans, it finds that people want the option to switch between Mainstream U.S. English (MUSE) and AAE, with MUSE preferred in formal settings and AAE welcomed in casual ones. When models were prompted with in-context examples of AAE, annotators rated the generated text as authentic as transcribed human speech, and generally not mocking or offensive.

What carries the argument

The mechanism is a paired continuation evaluation: human-transcribed prefixes (from CORAAL, Twitter, and NPR) are completed either by a human or by an LLM prompted with in-context examples, and Black American annotators rate the suffix on six Likert scales (coherence, AAE features, Black-sounding, White-sounding, mocking, offensive). The in-context prompting using CORAAL ground truth as chat history is what allowed the models to produce coherent AAE rather than refusal or off-topic output.

What would settle it

A direct adversarial test would be to have Black American annotators rate LLM-generated AAE against naturally occurring casual AAE speech (not interview transcripts) in matched contexts; if annotators judge the LLM output as significantly less authentic or more stereotyped than natural speech, the paper's parity claim would be falsified.

Watch

Extended reading notes

Core claim

The paper's central discovery is that Black Americans perceive LLM-generated AAE as comparable in authenticity to transcribed Black American speech from the CORAAL corpus, and sometimes as more AAE-heavy or more "Black sounding" than the human baseline. This holds across three LLMs (GPT 4o-mini, Llama 3, Mixtral) on judgments of coherence, presence of AAE features, and sounding like a Black American. At the same time, the survey shows a clear contextual preference: formal or task-specific settings call for MUSE, while personal or casual settings are open to AAE, and users want the autonomy to choose.

Load-bearing premise

The claim that LLM AAE is "on par" with human AAE rests on treating the CORAAL interview transcripts as the ground-truth representative of authentic AAE; if those transcripts are not representative, such as because interview speech differs from casual AAE or contains transcription artifacts, the parity conclusion is weakened.

Editorial extensions

If this is right

  • If LLMs can produce authentic AAE on demand, then user-facing assistants can offer AAE as a selectable mode without risking mockery, provided the user opts in.
  • The results imply that defaulting to MUSE in formal contexts aligns with user expectations, so product designers should keep MUSE as default and make AAE an explicit user choice.
  • The finding that LLM output was sometimes judged more "Black sounding" than human transcripts suggests a calibration target: models may be over-performing AAE features, and "on par" is not "identical."
  • The lack of perceived offensiveness suggests that fears of LLM AAE automatically being minstrelsy are not supported by Black American judgments in this sample.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's parity result could be extended by testing LLM-generated AAE against spontaneous casual speech rather than interview transcripts, which may be a higher bar for authenticity.
  • If Black Americans want AAE as an opt-in feature, designers should treat AAE generation as a user-controlled setting rather than a default, and may need a "dialect intensity" control to avoid over-AAE output.
  • The survey's strong preference for MUSE in formal settings suggests that any product feature offering AAE should be paired with context controls, since use of AAE in formal contexts may expose users to linguistic discrimination.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper reports a two-part study with Black American participants: a scenario-based survey (n=104) on when participants would want LLMs to use African American English (AAE), and an annotation study (n=228; 8,654 judgments) in which Black American annotators rated human and LLM continuations on coherence, AAE features, Black-sounding, White-sounding, mocking, and offensiveness. The LLMs evaluated are GPT-4o-mini, Llama-3-70B-Instruct, and Mixtral-8x7B, with human baselines drawn from CORAAL transcribed interviews, a Twitter AAE corpus, and NPR MUSE interviews. The paper finds that participants prefer MUSE in formal settings but want the option to use AAE in casual settings, and it claims that appropriately prompted LLM outputs are perceived as authentic AAE 'on par' with transcribed Black American speech while generally not being seen as mocking or offensive.

Significance. If the preference findings hold, the survey provides useful, community-grounded guidance for designing language technologies that give Black Americans control over when AAE is used. The annotation study is valuable as a participatory evaluation of LLM AAE output by the relevant speech community, and the authors are transparent about researcher positionality and data limitations. The release of data and code is a concrete strength. The main weakness is that the headline authenticity-parity claim is not actually established by the reported statistics on the most diagnostic dimensions, where LLM outputs significantly exceed the human AAE baseline; the paper needs either an equivalence-based analysis or a carefully qualified conclusion.

major comments (3)
  1. [§4.2.1, Table 3; Abstract; §5] The statement that LLM outputs have 'a level of AAE authenticity on par with transcripts of Black American speech' is not supported by the study's most diagnostic measures. For CORAAL, the AAE Features mean for GPT is 1.18 (p<0.001) versus the human mean of 0.18, and for Llama it is 0.86 (p<0.01); for Black Sounding, GPT's mean is 1.01 (p<0.01) versus the human mean of 0.39. These are significant differences on the exact dimensions that define AAE authenticity, and the paper itself notes in §5 that LLM outputs were 'often perceived as more AAE-heavy' than the baseline. The absence of significant differences on other dimensions (Coherence, White Sounding, Mocking, Offensive) is not evidence of parity, because a non-significant difference is not an equivalence test. The central claim should be reframed as 'at least as AAE-strong as the baseline, within a range annotators found acceptable,' or supported by an explicit equivalence or non-inferiority analysis with pre-specified bounds.
  2. [§4.2.1, Table 3; §1 Contribution 2] The claim that annotators 'did not consider them to be mocking or offensive' is contradicted by the Llama Tweets condition. In Table 3, the Mocking mean for Llama Tweets continuations is 0.14 (p<0.001) while the human Tweets baseline is -0.79; a mean above zero indicates slight agreement that the text sounds like mocking. The Offensive mean for Llama Tweets is -0.07 (p<0.001) versus the human baseline of -0.96, which is neutral rather than disagreement. The abstract and Contribution 2 need to be qualified to acknowledge this exception, and §5's statement that participants 'generally disagreed that the machine-generated text ... was offensive to or mocking of Black Americans' should explicitly exclude or discuss the Llama Tweets case.
  3. [§3.2, Table 3, Tweets columns] The human Tweets baseline is rated by the same annotators as lacking AAE features (µ=-0.57) and not Black-sounding (µ=-0.30), so it does not function as an 'authentic AAE' ground truth in the way the CORAAL baseline does. Because the Tweets human baselines are not perceived as AAE, comparisons of LLM Tweets continuations to these baselines cannot support the general claim that LLM output is on par with human AAE; they only show that LLM continuations are more AAE-like than a baseline that annotators already judged to be non-AAE. The paper should either restrict the parity claim to CORAAL or analyze the Tweets condition separately as a test of whether LLM output can exceed a weak AAE baseline.
minor comments (6)
  1. [§5, first paragraph] The phrase 'our the language in our human AAE corpus' should read 'the language in our human AAE corpus.'
  2. [§4.2, Results Analysis Approach] The text says 'significant' means a false discovery rate of 5% and then states that Bonferroni corrections were applied with 10 comparisons per judgment. FDR and Bonferroni are different multiplicity controls; please clarify which method was actually used for Tables 3 and 4 and whether the 48 between-corpus comparisons in Table 4 received a correction.
  3. [§A.5.1, §A.5.2, §A.5.3] The GPT instruction text contains the typo 'do not use of the strings' in three places; it should be 'do not use the strings.'
  4. [§4.2.1, final paragraph] The word 'equivalentally' should be 'equivalently.'
  5. [Figure 2] The annotation labels such as '2 - Strongest Agreement' should be explained in the caption, since the Likert scale is elsewhere described as ranging from -2 (Strongly Disagree) to +2 (Strongly Agree).
  6. [§3.2] The paper does not report inter-annotator agreement for the six Likert judgments; reporting a measure such as Krippendorff's alpha would strengthen confidence in the reliability of the aggregate scores.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the authenticity comparison rests on external human baselines and independent Black American annotator judgments, not on fitted parameters or self-referential definitions.

full rationale

This paper is an empirical study, not a derivation, and I found no step where a claimed result is equivalent to its input by construction. The central authenticity claim is that Black American annotators rated LLM-generated AAE continuations as on par with transcribed human AAE speech from CORAAL. That comparison is anchored in external data: CORAAL transcripts, an X/Twitter AAE corpus, and NPR MUSE interviews, with judgments collected from separate Black American annotators who were blind to whether a suffix was human- or LLM-produced (§3.2). No parameter is fitted to the annotation scores and then renamed as a prediction; the Likert ratings are raw elicited judgments. The survey preferences (§4.1) are likewise directly measured participant choices. The paper does cite Cunningham et al. (2024), which shares an author (Hal Daumé III), but that citation appears in related work as background on AAE speakers' labor with language technologies and is not load-bearing for the paper's own survey or annotation results. The Limitations section (§7) honestly flags that CORAAL is a transcription and may contain artifacts that affect perceived authenticity, and that annotators were not told whether text was human or machine generated; these are validity and interpretation concerns, not evidence of circularity. A skeptic could argue that the 'on par' conclusion is not fully supported because Table 3 shows statistically significant differences on AAE Features and Black Sounding for some LLMs relative to the CORAAL baseline, but that is an argument about inference from the data, not a circularity where the output is predetermined by the input. The study is self-contained against external benchmarks and its judgments come from independent community annotators, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on assumptions about the representativeness of the CORAAL baseline, the interval nature of Likert scores, and the generalizability of the Prolific sample. None are fitted parameters or invented entities.

assumptions (4)
  • domain assumption Likert responses are treated as interval data for computing means and t-tests.
    The paper maps 5-point Likert scores to -2..+2 and applies t-tests to compare means, which assumes interval-level measurement (Section 4.2).
  • domain assumption CORAAL transcribed interviews represent authentic AAE.
    The study uses CORAAL as the human baseline for judging LLM authenticity (Sections 3.2 and A.1). If this baseline is unrepresentative, the comparison is invalid.
  • domain assumption The Prolific sample is reasonably representative of Black Americans with AAE familiarity.
    The study generalizes from an online convenience sample of 104 survey and 228 annotator participants to the broader Black American population (Section 3.1, Limitations).
  • domain assumption Annotating single exchange pairs captures authentic language use.
    The paper acknowledges that single pairs lose nuance of full dialogue, which could affect authenticity and coherence judgments (Limitations).

how reviews work

0 comments
Cite this review

Pith. "Pith review of My LLM might Mimic AAE -- But When Should it?." pith.science (2026). https://pith.science/paper/5OKXBLQD

@misc{pith2026250204564,
  author       = {Pith},
  title        = {Pith review of: My LLM might Mimic AAE -- But When Should it?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5OKXBLQD}},
  note         = {Machine review of arXiv:2502.04564}
}
abstract

We examine the representation of African American English (AAE) in large language models (LLMs), exploring (a) the perceptions Black Americans have of how effective these technologies are at producing authentic AAE, and (b) in what contexts Black Americans find this desirable. Through both a survey of Black Americans ($n=$ 104) and annotation of LLM-produced AAE by Black Americans ($n=$ 228), we find that Black Americans favor choice and autonomy in determining when AAE is appropriate in LLM output. They tend to prefer that LLMs default to communicating in Mainstream U.S. English in formal settings, with greater interest in AAE production in less formal settings. When LLMs were appropriately prompted and provided in context examples, our participants found their outputs to have a level of AAE authenticity on par with transcripts of Black American speech. Select code and data for our project can be found here: https://github.com/smelliecat/AAEMime.git

Figures

Figures reproduced from arXiv: 2502.04564 by the authors.

Figure 1
Figure 1. This heatmap depicts participant (n = 104) pref￾erences (horizontal axis) for the use of language varieties in seven scenarios (vertical axis). A greater number of partici￾pants preferred either for the system to use MUSE or to allow them to select between MUSE and AAE. There were some ex￾ceptions: e.g., auto-detection was considered more acceptable in SMS, and MUSE was preferred for email. a preference gradient tha… view at source ↗
Figure 2
Figure 2. Examples of response continuations generated by [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Sample question from the survey on participants preference in a realistic scenario. [PITH_FULL_IMAGE:figures/full_fig_p019_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Sample question from annotation task where participants are asked to consider the highlighted, underlined part of the [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Left: Bar Plot of Gender Distribution Among Respondents: This graph displays the count of survey participants according to their gender identification, including Female, Male, Non-Binary, Undisclosed, and Other. The largest groups are Female and Male, with significant …
Figure 6
Figure 6. Figure 6: Bar Plot of Gender Distribution Across Age Groups: This graph presents a breakdown of gender identities among survey respondents segmented by age groups ranging from 18 to 64 and over. The categories include Female, Male, and Non-Binary, as well as respondents who pref…
Figure 7
Figure 7. Figure 7: Left: Bar Plot of Survey Respondents by Region: This graph displays the number of survey respondents categorized by their geographic regions within the United States—South, Northeast, West, and Midwest. The South shows the highest participation with 53 respondents, fol…
Figure 8
Figure 8. Figure 8: Bar Plot of Education Level Distribution Among Respondents: This graph shows the diverse educational backgrounds of survey participants, ranging from high school diplomas to doctorate degrees. Each bar represents the count of individuals with specific educational quali…
Figure 9
Figure 9. Figure 9: Bar Plot of Language Proficiency Preferences: This graph quantifies participant preferences for language proficiency in different varieties, focusing on Mainstream U.S. English (MUSE) and African American English (AAE). The bars represent the number of participants pro…
Figure 10
Figure 10. Figure 10: Bar Plot of Perceived Benefits: This graph illustrates the various benefits identified by participants when African American English (AAE) is incorporated into chatbot interactions. Each bar represents specific advantages such as enhanced cultural representation, pers…
Figure 11
Figure 11. Figure 11: Bar Plot of Participant Concerns: This graph illustrates the range of selected concerns among participants regarding the integration of African American English (AAE) into chatbot technology. Each bar represents a distinct set of issues, from perpetuating stereotypes …
Figure 12
Figure 12. Figure 12: Bar Plot of Terminology Preferences for AAE: This graph presents the count of participants’ preferences for various terms used to describe African American English. Each bar represents the popularity of terms such as ‘African American English’, ‘African American Verna…
Figure 13
Figure 13. Figure 13: Bar Plot of Contextual Preferences for Using AAE: This graph displays the frequency of preferences among participants for using African American English (AAE) across various social and professional contexts. Each bar indicates the count of participants who prefer usin…
Figure 14
Figure 14. Figure 14: Bar Plot of Preferred Self-Identification Terms: This graph illustrates the distribution of preferred self-identification terms among respondents, highlighting the diversity within racial and ethnic identities. The terms range from ‘Black’ and ‘African American’ to mo…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

12 extracted references · 7 canonical work pages

  1. [5]

    In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 789–798

    Exploring the role of grammar and word choice in bias toward African American English (AAE) in hate speech classifica- tion. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 789–798. Jazmia Henry

  2. [7]

    Preprint, arXiv:2403.00742

    Dialect prejudice predicts AI decisions about people’s character, employability, and criminality. Preprint, arXiv:2403.00742. Yolanda Holt

  3. [8]

    Preprint, arXiv:2401.04088

    Mix- tral of experts. Preprint, arXiv:2401.04088. Tyler Kendall and Charlie Farrington

  4. [10]

    In Proceedings of the 2024 Joint International Conference on Compu- tational Linguistics, Language Resources and Evalu- ation (LREC-COLING 2024), pages 10403–10415, Torino, Italia

    Lever- aging syntactic dependencies in disambiguation: The case of African American English. In Proceedings of the 2024 Joint International Conference on Compu- tational Linguistics, Language Resources and Evalu- ation (LREC-COLING 2024), pages 10403–10415, Torino, Italia. ELRA and ICCL. John R Rickford, Greg J Duncan, Lisa A Gennetian, Ray Yun Gou, Rebec...

  5. [11]

    Disambiguation of morpho-syntactic features of African American English -- the case of habitual be

    Disambiguation of morpho-syntactic features of African American English–the case of habitual be. arXiv preprint arXiv:2204.12421. Hanna L Smokoski

  6. [12]

    ground truth

    Llamafac- tory: Unified efficient fine-tuning of 100+ language models. arXiv preprint arXiv:2403.13372. A Appendix A.1 Preparation of the CORAAL Corpus (Black American transcribed interviews) Our prompt texts or prefixes needed to have authentic AAE to the degree possible and cover a broad range of the different variations of AAE spoken in the wild. (Lane...

  7. [2016]

    arXiv preprint arXiv:1608.08868

    Demographic dialectal variation in social media: A case study of African-American English. arXiv preprint arXiv:1608.08868. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

  8. [2019]

    Interview: A Large-Scale Open-Source Corpus of Media Dialog

    AI- Based Digital Assistants. Business & Information Systems Engineering, 61:535–544. Bodhisattwa Prasad Majumder, Shuyang Li, Jianmo Ni, and Julian McAuley. 2020a. Large-scale modeling of media dialog with discourse patterns and knowledge grounding. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process- ing (EMNLP). *Equa...

Show all 12 references
  1. [2020]

    Advances in neural information processing systems, 33:1877–1901

    Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901. Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anasta- sios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica

  2. [2021]

    Accessed: 2024- 06-01

    Aave corpora. Accessed: 2024- 06-01. Jane H Hill

  3. [2022]

    It’s Kind of Like Code-Switching

    “It’s Kind of Like Code-Switching”: Black Older Adults’ Experiences with a V oice Assistant for Health Information Seek- ing. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, CHI ’22, New York, NY , USA. Association for Computing Machinery. Cami...

  4. [2024]

    Preprint, arXiv:2403.04132

    Chatbot arena: An open platform for evaluating LLMs by human prefer- ence. Preprint, arXiv:2403.04132. Jay L. Cunningham, Su Lin Blodgett, Hal Daumé III, Christina Harrington, Hanna Wallach, and Michael Madaio

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.