{"id":"b079685f-29cf-489f-adb1-4114bf904a35","arxiv_id":"2502.04564","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Black Americans favor letting users choose when LLMs speak African American English, and rated LLM-generated AAE as comparable to human transcripts, although some outputs were seen as mocking.","lead":"This study asked 104 Black American adults how they want AI chatbots to use African American English, and had 228 Black Americans judge whether LLM-generated AAE sounded authentic. Most people wanted the option to switch between standard English and AAE, depending on formality, and rated well-prompted LLM output as similar to real Black speech.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'on par' authenticity claim is contradicted by the paper's own Table 3 data: LLMs score significantly above the human CORAAL baseline on AAE Features and Black Sounding, so the central parity conclusion is an interpretive leap rather than a tested result.","rationale":"The reader's conditional verdict is appropriate, but the specific load-bearing weakness is not only that CORAAL may be an imperfect ground truth; it is that the paper's own results show systematic divergence on the two measures that most directly operationalize authenticity. Treating non-significant t-tests as evidence of parity is a logical error; the correct statistical approach is an equivalence test. The paper has real strengths: community-based annotation, careful prompting, and transparent limitations. However, the headline in the abstract and conclusion—'on par'—overstates what the data show. I therefore keep the conditional verdict rather than rejecting the paper, because the survey findings and the annotation dataset remain valuable and the claim can be tested with existing data. My concern is partial overlap with the reader's weakest_assumption: the reader emphasized CORAAL representativeness, whereas I emphasize that even accepting CORAAL as ground truth, the parity assertion is not supported by the construct-defining ratings.","tokens_in":20426,"tokens_out":5694,"duration_ms":62233,"concrete_test":"Run a two one-sided tests (TOST) equivalence analysis on the existing Table 3 data for the AAE Features and Black Sounding judgments, comparing each LLM to the human CORAAL baseline with a pre-registered equivalence bound of ±0.25 Likert points (roughly a quarter of a scale step, justified by the judgment scale). If the 90% confidence intervals for the human-vs-LLM differences fall outside the bounds on either dimension, the 'on par' claim is statistically rejected; if they fall inside, the parity claim survives on the study's own measures.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs produce AAE 'on par' with transcribed Black American speech (Abstract; §5)—is not supported by the study's most diagnostic measures. In Table 3 (§4.2.1), for the CORAAL corpus, GPT's AAE Features mean is 1.18 vs. human 0.18 (p<0.001), Llama's is 0.86 (p<0.01), and GPT's Black Sounding mean is 1.01 vs. human 0.39 (p<0.01). These are significant differences on the exact dimensions that define AAE authenticity. The paper's own §5 acknowledges LLM outputs are 'more AAE-heavy' than the baseline, so the 'on par' framing cannot follow from these data. The non-significant differences on Coherence, White Sounding, Mocking, and Offensive are consistent with many alternative hypotheses and are not positive evidence of authenticity; the Tweets human baseline is itself rated as lacking AAE features (μ=-0.57) and not Black-sounding (μ=-0.30), so the 'human ground truth' behaves differently across the two AAE corpora. Additionally, the comparison is asymmetric: human suffixes are orthographic transcripts with disfluencies and transcription artifacts (acknowledged in §7), while LLM suffixes are polished continuations generated specifically to be AAE. Because the central contribution is the authenticity parity claim, and the data on construct-defining dimensions point away from parity, the conclusion should be reframed or verified with an equivalence analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a two-part study with Black American participants: a scenario-based survey (n=104) on when participants would want LLMs to use African American English (AAE), and an annotation study (n=228; 8,654 judgments) in which Black American annotators rated human and LLM continuations on coherence, AAE features, Black-sounding, White-sounding, mocking, and offensiveness. The LLMs evaluated are GPT-4o-mini, Llama-3-70B-Instruct, and Mixtral-8x7B, with human baselines drawn from CORAAL transcribed interviews, a Twitter AAE corpus, and NPR MUSE interviews. The paper finds that participants prefer MUSE in formal settings but want the option to use AAE in casual settings, and it claims that appropriately prompted LLM outputs are perceived as authentic AAE 'on par' with transcribed Black American speech while generally not being seen as mocking or offensive.","tokens_in":20736,"tokens_out":5170,"duration_ms":52097,"significance":"If the preference findings hold, the survey provides useful, community-grounded guidance for designing language technologies that give Black Americans control over when AAE is used. The annotation study is valuable as a participatory evaluation of LLM AAE output by the relevant speech community, and the authors are transparent about researcher positionality and data limitations. The release of data and code is a concrete strength. The main weakness is that the headline authenticity-parity claim is not actually established by the reported statistics on the most diagnostic dimensions, where LLM outputs significantly exceed the human AAE baseline; the paper needs either an equivalence-based analysis or a carefully qualified conclusion.","major_comments":[{"comment":"The statement that LLM outputs have 'a level of AAE authenticity on par with transcripts of Black American speech' is not supported by the study's most diagnostic measures. For CORAAL, the AAE Features mean for GPT is 1.18 (p<0.001) versus the human mean of 0.18, and for Llama it is 0.86 (p<0.01); for Black Sounding, GPT's mean is 1.01 (p<0.01) versus the human mean of 0.39. These are significant differences on the exact dimensions that define AAE authenticity, and the paper itself notes in §5 that LLM outputs were 'often perceived as more AAE-heavy' than the baseline. The absence of significant differences on other dimensions (Coherence, White Sounding, Mocking, Offensive) is not evidence of parity, because a non-significant difference is not an equivalence test. The central claim should be reframed as 'at least as AAE-strong as the baseline, within a range annotators found acceptable,' or supported by an explicit equivalence or non-inferiority analysis with pre-specified bounds.","section":"§4.2.1, Table 3; Abstract; §5"},{"comment":"The claim that annotators 'did not consider them to be mocking or offensive' is contradicted by the Llama Tweets condition. In Table 3, the Mocking mean for Llama Tweets continuations is 0.14 (p<0.001) while the human Tweets baseline is -0.79; a mean above zero indicates slight agreement that the text sounds like mocking. The Offensive mean for Llama Tweets is -0.07 (p<0.001) versus the human baseline of -0.96, which is neutral rather than disagreement. The abstract and Contribution 2 need to be qualified to acknowledge this exception, and §5's statement that participants 'generally disagreed that the machine-generated text ... was offensive to or mocking of Black Americans' should explicitly exclude or discuss the Llama Tweets case.","section":"§4.2.1, Table 3; §1 Contribution 2"},{"comment":"The human Tweets baseline is rated by the same annotators as lacking AAE features (µ=-0.57) and not Black-sounding (µ=-0.30), so it does not function as an 'authentic AAE' ground truth in the way the CORAAL baseline does. Because the Tweets human baselines are not perceived as AAE, comparisons of LLM Tweets continuations to these baselines cannot support the general claim that LLM output is on par with human AAE; they only show that LLM continuations are more AAE-like than a baseline that annotators already judged to be non-AAE. The paper should either restrict the parity claim to CORAAL or analyze the Tweets condition separately as a test of whether LLM output can exceed a weak AAE baseline.","section":"§3.2, Table 3, Tweets columns"}],"minor_comments":[{"comment":"The phrase 'our the language in our human AAE corpus' should read 'the language in our human AAE corpus.'","section":"§5, first paragraph"},{"comment":"The text says 'significant' means a false discovery rate of 5% and then states that Bonferroni corrections were applied with 10 comparisons per judgment. FDR and Bonferroni are different multiplicity controls; please clarify which method was actually used for Tables 3 and 4 and whether the 48 between-corpus comparisons in Table 4 received a correction.","section":"§4.2, Results Analysis Approach"},{"comment":"The GPT instruction text contains the typo 'do not use of the strings' in three places; it should be 'do not use the strings.'","section":"§A.5.1, §A.5.2, §A.5.3"},{"comment":"The word 'equivalentally' should be 'equivalently.'","section":"§4.2.1, final paragraph"},{"comment":"The annotation labels such as '2 - Strongest Agreement' should be explained in the caption, since the Likert scale is elsewhere described as ranging from -2 (Strongly Disagree) to +2 (Strongly Agree).","section":"Figure 2"},{"comment":"The paper does not report inter-annotator agreement for the six Likert judgments; reporting a measure such as Krippendorff's alpha would strengthen confidence in the reliability of the aggregate scores.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper's preference-survey contribution is solid and likely publishable after revision. The main risk is the overstatement of the authenticity parity claim in the abstract and conclusion, which conflicts with significant differences in Table 3 on AAE Features and Black Sounding. I would encourage the editor to require either a reframing of the claim or an equivalence analysis before acceptance, rather than treating this as a purely cosmetic issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading for the survey and annotation design, but the central claim that LLM AAE is on par with human transcripts doesn't follow from the data. The stress-test note is right: Table 3 shows significant differences on AAE Features for GPT and Llama, and on Black Sounding for GPT, relative to the CORAAL human baseline. The LLM outputs are more AAE-heavy, not equivalent. The paper acknowledges this in §5, but the abstract and contributions still say \"on par.\" That's an interpretive leap, and the non-significant differences on coherence, mocking, and offensiveness don't rescue it.\n\nWhat's genuinely new and good: this is the first study to combine a preference survey of Black Americans with community-based annotation of LLM-generated AAE. The survey finding—users want autonomy and MUSE in formal settings, with openness to AAE in casual contexts—is a concrete, useful result. The annotation effort is serious: 228 Black American annotators, 8,654 judgments, and a transparent interface. The paper is also honest about its own limitations, including the fact that LLM outputs were more AAE-heavy than the baseline.\n\nThe soft spots beyond the overclaim: the Tweets human baseline is rated as lacking AAE features and not Black-sounding, so the \"human ground truth\" behaves inconsistently across corpora. And the comparison is asymmetric—human transcripts have disfluencies and transcription artifacts, while LLM outputs are polished continuations. That alone could explain some of the perceived differences. The survey is a convenience sample and descriptive, which the authors mostly acknowledge. The GitHub repo offers only select code and data, limiting replication.\n\nBottom line: the survey results and the annotation dataset are solid contributions and worth peer review. The \"on par\" authenticity claim needs reframing or an equivalence test before it can be taken at face value. I'd send this to a serious referee, with the expectation of revision. I'd cite it for the community-based evaluation template and the preference findings, not for the parity conclusion.","headline":"A useful community-based evaluation of LLM AAE, but the 'on par' conclusion is overstated relative to the paper's own numbers.","tokens_in":21263,"tokens_out":2438,"would_cite":true,"duration_ms":26016,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Black Americans judge LLM-generated African American English as authentic as human speech, and want the choice of when it appears.","keywords":["African American English","large language models","AAE authenticity","user preferences","in-context learning","linguistic judgments","dialect prejudice","Black American perspectives"],"falsifier":"A direct adversarial test would be to have Black American annotators rate LLM-generated AAE against naturally occurring casual AAE speech (not interview transcripts) in matched contexts; if annotators judge the LLM output as significantly less authentic or more stereotyped than natural speech, the paper's parity claim would be falsified.","tokens_in":217,"feed_emoji":"🗣️","tokens_out":4558,"duration_ms":97545,"temperature":0.7,"pith_summary":"This paper asks whether Black Americans want large language models to produce African American English (AAE) and whether current models do it well. Through a survey of 104 Black Americans and annotation of LLM outputs by 228 Black Americans, it finds that people want the option to switch between Mainstream U.S. English (MUSE) and AAE, with MUSE preferred in formal settings and AAE welcomed in casual ones. When models were prompted with in-context examples of AAE, annotators rated the generated text as authentic as transcribed human speech, and generally not mocking or offensive.","feed_headline":"LLM AAE is as authentic as human speech, Black Americans say","feed_subtitle":"Black American annotators found LLM AAE on par with human transcripts - and want control over when it appears.","key_machinery":"The mechanism is a paired continuation evaluation: human-transcribed prefixes (from CORAAL, Twitter, and NPR) are completed either by a human or by an LLM prompted with in-context examples, and Black American annotators rate the suffix on six Likert scales (coherence, AAE features, Black-sounding, White-sounding, mocking, offensive). The in-context prompting using CORAAL ground truth as chat history is what allowed the models to produce coherent AAE rather than refusal or off-topic output.","core_discovery":"The paper's central discovery is that Black Americans perceive LLM-generated AAE as comparable in authenticity to transcribed Black American speech from the CORAAL corpus, and sometimes as more AAE-heavy or more \"Black sounding\" than the human baseline. This holds across three LLMs (GPT 4o-mini, Llama 3, Mixtral) on judgments of coherence, presence of AAE features, and sounding like a Black American. At the same time, the survey shows a clear contextual preference: formal or task-specific settings call for MUSE, while personal or casual settings are open to AAE, and users want the autonomy to choose.","pith_inferences":["The paper's parity result could be extended by testing LLM-generated AAE against spontaneous casual speech rather than interview transcripts, which may be a higher bar for authenticity.","If Black Americans want AAE as an opt-in feature, designers should treat AAE generation as a user-controlled setting rather than a default, and may need a \"dialect intensity\" control to avoid over-AAE output.","The survey's strong preference for MUSE in formal settings suggests that any product feature offering AAE should be paired with context controls, since use of AAE in formal contexts may expose users to linguistic discrimination."],"forward_implications":["If LLMs can produce authentic AAE on demand, then user-facing assistants can offer AAE as a selectable mode without risking mockery, provided the user opts in.","The results imply that defaulting to MUSE in formal contexts aligns with user expectations, so product designers should keep MUSE as default and make AAE an explicit user choice.","The finding that LLM output was sometimes judged more \"Black sounding\" than human transcripts suggests a calibration target: models may be over-performing AAE features, and \"on par\" is not \"identical.\"","The lack of perceived offensiveness suggests that fears of LLM AAE automatically being minstrelsy are not supported by Black American judgments in this sample."],"supporting_citations":[{"why":"Supplies the CORAAL corpus, which provides the human AAE baseline and the in-context training examples used for prompting the LLMs.","marker":"Kendall and Farrington, 2023"},{"why":"Introduces in-context learning, the prompting technique the paper uses to get LLMs to produce AAE continuations.","marker":"Brown et al., 2020"},{"why":"Documents dialect prejudice in LLMs, motivating the paper's test of whether LLM-generated AAE is perceived as mocking or offensive.","marker":"Hofmann et al., 2024"},{"why":"Provides prior evidence of racial disparities in automatic speech recognition, framing why AAE representation in language technology matters.","marker":"Koenecke et al., 2020"}],"fun_headline_variants":["Black Americans rate LLM AAE as authentic as human speech","LLM AAE authenticity on par with human speech, but users want choice","When should LLMs mimic AAE? Black Americans want control","LLM AAE passes authenticity bar, yet formal settings demand MUSE","Black users: LLM AAE is real, but only for casual contexts"],"cache_read_input_tokens":23424,"weakest_assumption_plain":"The claim that LLM AAE is \"on par\" with human AAE rests on treating the CORAAL interview transcripts as the ground-truth representative of authentic AAE; if those transcripts are not representative, such as because interview speech differs from casual AAE or contains transcription artifacts, the parity conclusion is weakened.","fun_headline_variants_meta":{"raw":{"variants":["Black Americans rate LLM AAE as authentic as human speech","LLM AAE authenticity on par with human speech, but users want choice","When should LLMs mimic AAE? Black Americans want control","LLM AAE passes authenticity bar, yet formal settings demand MUSE","Black users: LLM AAE is real, but only for casual contexts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1326,"prompt_tokens":845,"completion_tokens":481,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":461,"completion_tokens_details":{"reasoning_tokens":387}},"tokens_in":461,"tokens_out":481,"duration_ms":4798,"temperature":1.0,"reasoning_tokens":387,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T22:16:17.147494+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct adversarial test would be to have Black American annotators rate LLM-generated AAE against naturally occurring casual AAE speech (not interview transcripts) in matched contexts; if annotators judge the LLM output as significantly less authentic or more stereotyped than natural speech, the paper's parity claim would be falsified.","supporting_citations":[],"review_version":1}