{"id":"26cffcf5-422e-41aa-9694-1f3c5a1efb1c","arxiv_id":"2501.03441","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"African American English text chatbots underperformed standard English, but spoken chatbots with an African American voice and mild AAE improved warmth, similarity, and engagement for AAE-speaking evaluators.","lead":"This study built text and spoken chatbots that use African American English and an African American voice, then had AAE-speaking students rate them. Text chatbots in AAE performed worse than standard English, while spoken chatbots with an African American voice and mild AAE were rated higher on warmth, similarity, and engagement.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Spoken benefit is not shown to come from AAE elements: best spoken condition uses AA accent with SAE text (§4.4.2), and single-voice-per-condition confounds accent with speaker identity.","rationale":"The reader identified static evaluation as the weakest assumption. I agree that is a limitation and is honestly flagged in Section 7. However, a more immediate threat to the central claim is internal: the spoken benefit attributed to 'AAE elements' is never isolated. The best spoken condition is AA accent plus SAE text, and the paper does not report whether Low/Medium AAE text adds anything beyond the AA voice. Since the central claim explicitly includes AAE elements, this is load-bearing. The single-voice-per-condition design compounds the problem because accent and speaker identity are perfectly confounded; a different AA speaker might not produce the same preference. The absence of significance testing with n=8 means the observed mean advantages could be within sampling error. These concerns do not require rejecting the paper: the text-chatbot underperformance pattern and the transparency of the limitations section are useful, and the released code/data allow the proposed re-analyses. Conditional acceptance remains appropriate, with the condition that the spoken benefit be shown to be robust to voice identity and to AAE-text-level contrasts.","tokens_in":18144,"tokens_out":5831,"duration_ms":52695,"concrete_test":"Re-analyze the existing Figure 4 data with evaluator as a random effect; compute the S-vs-L contrast on Warmth, Similarity to Self, and Engagement Preference with pre-specified equivalence bounds. If the contrast is nonsignificant or negligible, the 'AAE elements' component of the central claim fails. Separately, to test the voice-identity confound, run the spoken evaluation with 3-5 AA and 3-5 SA voice exemplars matched for age/gender; if the AA-voice advantage does not replicate across voices, the accent effect is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract attributes the spoken benefit to 'an African American voice and AAE elements,' but Section 4.4.2 reports that the most effective spoken chatbot 'pairs an AA accent with SAE dialect features' and outperforms the SAE baseline on all dimensions. No pairwise contrast is reported between that AA-accent/SAE condition (S) and the AA-accent/Low-AAE condition (L) for Warmth, Similarity to Self, or Engagement Preference, so the improvement claimed for L over the SAE baseline cannot be separated from the effect of the AA voice alone. The design also uses exactly one AA voice exemplar (Appendix B: ATL_se0_ag2_f_02_1) and one SA voice, leaving 'African American voice' confounded with speaker identity, pitch, and prosody. With only 8 spoken evaluators and no significance tests, the observed spoken benefit could be voice-identity or sampling noise. Section 7's static-evaluation limitation is acknowledged and real, but it is secondary: even within the static evaluation, the marginal contribution of AAE elements to the spoken chatbot is unidentified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops text-based and spoken chatbots that vary in African American English (AAE) dialect intensity and, for spoken chatbots, use an African American (AA) accent. The authors evaluate these chatbots with AAE-speaking participants on metrics such as comprehension, warmth, trustworthiness, similarity to self, and engagement preference, comparing them to standard English baselines. The main reported finding is a modality contrast: text-based AAE chatbots generally underperform their SAE counterparts, while spoken chatbots with an AA accent and AAE elements outperform the SAE baseline on warmth, similarity to self, and engagement preference, with the best spoken condition pairing an AA accent with SAE text.","tokens_in":18368,"tokens_out":2350,"duration_ms":24310,"significance":"If the result holds, the paper would provide useful, actionable evidence for dialect personalization in conversational AI, showing that voice-based personalization can benefit AAE-speaking users even when text-based dialect rendering does not. The study is notable for systematically varying AAE intensity across three LLM families, using multi-turn dialogues from five application domains, and for centering evaluation on self-identified AAE speakers rather than on generic crowdworkers. The authors also release their code and data, which strengthens reproducibility. However, the inferential base is thin: the evaluation uses small, partially non-overlapping rater groups, no significance tests, and a single voice exemplar per accent condition, so the central modality-contrast claim is not yet established at the level the abstract states.","major_comments":[{"comment":"The spoken benefit is not attributable to AAE elements because the best spoken condition pairs an AA accent with SAE dialect features. The abstract claims that spoken chatbots 'benefit from an African American voice and AAE elements,' but the paper does not report a direct pairwise contrast between the AA-accent/SAE condition (S) and the AA-accent/Low-AAE condition (L) on Warmth, Similarity to Self, or Engagement Preference. Without that contrast, the observed gains over the SAE baseline could be driven entirely by the AA voice.","section":"§4.4.2, Fig. 4"},{"comment":"The evaluation uses only 12 evaluators for text chatbots and 8 for spoken chatbots, and no significance tests or effect sizes are reported. Many of the confidence intervals in Figures 3 and 4 overlap substantially, so the qualitative claims that Low AAE 'enhances' warmth, similarity, and engagement preference, or that High AAE 'largely fails,' are not backed by inferential evidence. The authors should report pairwise tests or at minimum bootstrap confidence intervals for the key comparisons.","section":"§4.4, Figs. 3 and 4"},{"comment":"The text and spoken evaluations were conducted by different, only partially overlapping groups of evaluators, as the authors acknowledge. The manual verification that findings were 'largely consistent' is not quantified and does not address whether the text-versus-spoken contrast could be explained by rater variation. This is load-bearing because the central claim is a modality contrast, not just a within-modality effect.","section":"§7, Evaluator Differences"},{"comment":"The spoken condition uses exactly one AA voice exemplar (ATL_se0_ag2_f_02_1) and one SA voice exemplar, so 'African American voice' is confounded with speaker identity, pitch, and prosody. The authors should either use multiple voice exemplars or explicitly frame the result as specific to the chosen voice; otherwise the claimed voice-based benefit may not generalize beyond a single speaker.","section":"§3.2, Appendix B"},{"comment":"The abstract's wording overstates the findings. It says spoken chatbots benefit from 'an African American voice and AAE elements,' but Section 4.4.2 states that the most effective spoken configuration pairs an AA accent with SAE dialect features. The abstract and conclusion should be revised to say that the spoken benefit is associated with the AA voice, while the marginal contribution of AAE text elements in spoken output remains unclear.","section":"Abstract and §4.4.2"}],"minor_comments":[{"comment":"The title contains a typo: 'V oice' should be 'Voice'.","section":"Title"},{"comment":"There are inconsistent spellings of 'AAVE' as 'AA VE' and 'AAVE,' and inconsistent spacing in terms such as 'T ext' and 'V oice.' A copyedit pass would improve readability.","section":"Throughout"},{"comment":"The feature-tagging accuracy is reported as 91% for Claude and 86% for GPT-4o, but there is no inter-annotator agreement or error analysis for the gold test set itself. Since the test set is used to validate the tagging approach, reporting agreement would strengthen the claim that the automatic tagger is reliable.","section":"§4.3"},{"comment":"The 'Engagement Preference' metric asks whether the evaluator would prefer the AAE chatbot instead of the 'Original Chatbot,' but 'Original Chatbot' is not defined in the table. For spoken chatbots it is unclear whether the baseline is the SAE-accented or the SAE-dialect/SA-accent version.","section":"Table 7"},{"comment":"The claim that 'none of the studied models achieved strong representation of the grounding persona' is based on a single 'Text Persona Adherence' item, and the annotator comments about 'young male' AAE are anecdotal. Reporting representative examples and a coding scheme for those comments would make this observation more credible.","section":"§4.4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for a CS/CL venue and the authors are transparent about limitations. The main issue is that the abstract and conclusion make a stronger claim than the experimental design supports, particularly regarding the role of AAE elements in the spoken condition. I would encourage the editor to ask for either additional pairwise analysis or a tempering of the central claim, not for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on dialect personalization or sociolinguistic HCI. The paper gives the first controlled comparison I know of that varies AAE intensity in both text and spoken chatbot responses and gets judgments from screened AAE-speaking users. That alone is a real contribution. The authors also separate response generation from dialect rendering, use multi-turn dialogues across five domains, validate a feature tagger on literature-sourced examples at 91% accuracy, and release code and data. Those are concrete and reproducible pieces.\n\nThe text-side result is the most defensible: across three LLMs and three intensity levels, AAE text generally underperforms SAE text with these evaluators, and high AAE hurts. That is consistent and useful. The spoken-side claim is softer than the abstract suggests. Section 4.4.2 reports that the best spoken configuration pairs an AA accent with SAE text, outperforming the SAE baseline on everything. So the apparent boost from Low AAE over the SA baseline cannot be separated from the effect of the AA voice alone; no pairwise S-versus-L contrast is reported. On top of that, one AA voice exemplar is used, which confounds accent with speaker identity. With eight spoken evaluators and no significance tests, the spoken benefit is suggestive, not demonstrated. The abstract's phrase 'AAE elements, improving performance and preference' is not supported; the conclusion's more modest statement about preferring an African American voice is closer to the evidence.\n\nThe limitations section is honest: static third-party evaluation, non-overlapping evaluator groups, and small samples are all acknowledged. The authors even report the manual check on overlapping evaluators, which is more than many papers do. Minor point: they use Claude both for generating AAE and for tagging AAE features, which could inflate internal consistency, though the tagger is validated externally, so this is not load-bearing.\n\nOverall, this is a solid empirical start with a clear design and transparent reporting. It deserves peer review, but revision should include significance tests or effect sizes, explicit S-versus-L contrasts, multiple voices per accent condition, and a reworded abstract. I would not cite the spoken claim as established, but I would cite the text result and the released resources.\n\nIf I were the editor, I would send this to a serious venue with a request for revision, not desk-reject it.","headline":"A genuinely useful controlled comparison of AAE intensity in chatbots, but the spoken benefit is mostly the voice, not the dialect text—the abstract overstates what the design can separate.","tokens_in":18866,"tokens_out":2749,"would_cite":true,"duration_ms":29412,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adding an African American voice to spoken chatbots improves how AAE-speaking users rate them, while written AAE chatbots underperform their standard-English counterparts.","keywords":["African American English","dialect personalization","chatbot evaluation","text-to-speech","spoken dialogue systems","linguistic similarity","large language models","human-computer interaction"],"falsifier":"Run the same text-versus-speech comparison as a live interaction study in which AAE-speaking participants converse freely with the chatbots rather than watch recordings; if the spoken chatbot's advantage over text on warmth or engagement disappears or reverses, the modality-contrast claim is refuted. A more targeted check would compare the AA-accented SAE spoken chatbot against the SA-accented SAE baseline while controlling whether raters know the voice's intended ethnicity, testing whether the effect depends on explicit recognition rather than felt rapport.","tokens_in":17978,"feed_emoji":"🗣️","tokens_out":5874,"duration_ms":51334,"temperature":0.7,"pith_summary":"This paper tries to establish whether chatbots that use African American English (AAE) actually serve AAE-speaking users better than standard-English chatbots do. The authors build text chatbots by translating responses into AAE at three intensity levels with three large language models, and spoken chatbots by voicing those responses with an African American accent using text-to-speech. When AAE-speaking university students evaluated pre-recorded dialogues, the text AAE chatbots underperformed the standard-English baseline, but spoken chatbots with an African American voice and mild AAE scored higher on warmth, similarity to self, and engagement preference. The upshot, if correct, is that dialect personalization works in chatbots only when the voice carries it, and heavy dialect expression backfires.","feed_headline":"A Black-accented voice improves chatbot ratings","feed_subtitle":"Spoken chatbots with mild AAE beat standard English; heavy dialect and text-only AAE fall short.","key_machinery":"The central mechanism is the text-versus-speech contrast built from two separable components: an SAE-to-AAE translation function $E(I, D_a, D_b)$ implemented by prompting an LLM at three intensity levels (Low, Medium, High), and a text-to-speech stage using the F5 model conditioned on a reference clip of an African American speaker from the Corpus of Regional African American Language. Dialect expression is deliberately separated from response generation so that AAE changes only surface style, not content. Evaluation is static: 100 ten-turn SODA dialogues across five chatbot domains are translated, voiced, and rated by AAE-speaking judges on fifteen metrics covering dialect expression, speech quality, user alignment, and engagement. The comparison of the same content across text and speech, with dialect intensity varied, is what isolates modality as the decisive factor.","core_discovery":"The central claim is that linguistic personalization to African American English is modality-dependent. In text, every AAE chatbot—across three LLM families and three AAE intensity levels—scored at or below the Standard American English baseline on trustworthiness, role appropriateness, and engagement preference with AAE-speaking evaluators. In speech, the same dialect content paired with an African American accent improved key outcomes: the best configuration, an AA voice speaking standard English or low-intensity AAE, outperformed the standard-voice baseline on warmth, similarity to self, communication ease, role appropriateness, and engagement preference. The effect is non-monotonic: High AAE, which mostly means heavy phonetic rewriting, lowered inoffensiveness and naturalness and fell back toward or below baseline. The paper concludes that the voice, not the dialect text, is what carries personalization benefit, and that technology limits—especially text-to-speech trained mostly on standard English—still constrain authentic AA speech generation.","pith_inferences":["If the modality contrast holds in live use, designers should prioritize voice and accent personalization over dialect-heavy text generation; mild AAE plus an AA voice is a safer default than heavy phonetic rewriting.","The paper's annotator comments suggest LLMs default to a young male AAE persona; a testable extension is whether controlling persona explicitly changes the text-chatbot results, since persona mismatch rather than dialect per se could drive underperformance.","The non-monotonic effect of dialect intensity predicts that an adaptive system mirroring each user's own AAE level turn by turn would outperform any fixed level; that is a direct, testable consequence the paper leaves implicit.","Because the TTS model is trained mostly on standard English, the clarity drop at High AAE may be a technology ceiling rather than a user preference; as accented TTS improves, the optimal dialect level could shift upward."],"forward_implications":["Spoken chatbots with an African American voice and low-intensity AAE outperform a standard-voice, standard-English baseline on warmth, similarity to self, and engagement preference for AAE-speaking evaluators.","Text-based AAE chatbots do not outperform the SAE baseline on any measured characteristic; higher AAE intensity makes them worse on trustworthiness, role appropriateness, and engagement.","Heavy AAE expression, dominated by phonetic changes, is perceived as less inoffensive, less natural, and less clear, so extreme dialect levels are counterproductive.","The best spoken configuration uses an AA accent with SAE or Low/Medium AAE text, indicating the accent carries most of the personalization benefit in voice interactions.","Modality determines whether dialect personalization helps: the same AAE content that fails in text succeeds in speech."],"supporting_citations":[{"why":"Defines the AAE feature inventory used to build and evaluate dialect expression in chatbot responses.","marker":"Rickford (1999)"},{"why":"SODA supplies the 100 multi-turn dialogues across five domains used as the evaluation data.","marker":"Kim et al. (2023)"},{"why":"F5 text-to-speech model converts chatbot responses into speech conditioned on a speaker reference.","marker":"Chen et al. (2024)"},{"why":"Corpus of Regional African American Language provides the audio reference for the African American voice.","marker":"Kendall and Farrington (2023)"},{"why":"Prior AAE generation and evaluation work whose metrics and approach inform the chatbot evaluation design.","marker":"Deas et al. (2023)"},{"why":"Motivates separating dialect from content and supplies bias and inoffensiveness considerations for AAE chatbot design.","marker":"Fleisig et al. (2024)"},{"why":"Earlier African American-sounding TTS model that supports the feasibility and persona approach for accented voices.","marker":"Pinhanez et al. (2024)"},{"why":"LibriSpeech provides the Standard American voice references for the user and the SAE baseline chatbot.","marker":"Panayotov et al. (2015)"}],"fun_headline_variants":["AA voice lifts chatbot ratings, but text dialect falls flat","In speech, African American voice helps; in text, dialect hurts","Voice beats text: mild AAE wins, heavy AAE loses","For chatbots, accent matters more than dialect words","Spoken AAE outperforms standard, but only with mild dialect"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that AAE-speaking university students' ratings of pre-recorded, third-party dialogues predict how the same chatbots would be experienced in live, interactive use.","fun_headline_variants_meta":{"raw":{"variants":["AA voice lifts chatbot ratings, but text dialect falls flat","In speech, African American voice helps; in text, dialect hurts","Voice beats text: mild AAE wins, heavy AAE loses","For chatbots, accent matters more than dialect words","Spoken AAE outperforms standard, but only with mild dialect"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1366,"prompt_tokens":861,"completion_tokens":505,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":420}},"tokens_in":477,"tokens_out":505,"duration_ms":5136,"temperature":1.0,"reasoning_tokens":420,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:51:57.858238+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same text-versus-speech comparison as a live interaction study in which AAE-speaking participants converse freely with the chatbots rather than watch recordings; if the spoken chatbot's advantage over text on warmth or engagement disappears or reverses, the modality-contrast claim is refuted. A more targeted check would compare the AA-accented SAE spoken chatbot against the SA-accented SAE baseline while controlling whether raters know the voice's intended ethnicity, testing whether the effect depends on explicit recognition rather than felt rapport.","supporting_citations":[],"review_version":1}