Pith. sign in

REVIEW 3 major objections 5 minor 146 references

Not Like Us, Hunty: Measuring Perceptions and Behavioral Effects of Minoritized Anthropomorphic Cues in LLMs

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Sociolect-speaking AI can feel warmer yet earn less trust.

desk verdict Solid large-sample study on sociolect use in LLMs, with a real finding that behavioral reliance can diverge from self-reported affinity, but the causal claim is weakened by unmatched sociolect stimuli. read the letter →

arxiv 2505.05660 v3 pith:PNZNVRZD submitted 2025-05-08 cs.HC

classification cs.HC
keywords sociolectadaptationAfricanAmericanEnglishQueerslangLLMpersonalizationuserreliancetrustsocialpresenceculturalappropriation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether an LLM agent that speaks a user's own sociolect, such as African American English or Queer slang, makes that user rely on it more and perceive it better, as personalization promises. In controlled question-answering studies with 498 AAE speakers and 487 Queer slang speakers, both groups accepted more suggestions from the standard American English agent. AAE speakers also trusted, preferred, and felt less frustrated with the standard agent. Queer slang speakers, however, felt significantly more social presence from the Queer slang agent even though they relied on it no more than the standard one. The paper concludes that sociolect adaptation can create warmth without creating reliance, and that behavioral measures can diverge from what self-reported preference and personalization would predict.

What carries the argument

The central mechanism is the templated LLM suggestion: 14 warmth phrases crossed with 14 confidence expressions carrying 30–70% reliability, generating 196 standard English suggestions, then translated into AAE via in-context learning from AAE corpora or into Queer slang via persona prompting, with each translation vetted by crowdsourced verification panels requiring at least 4 of 5 raters. Participants answered deliberately difficult questions about short videos and, for each question, chose between “Use LLM’s Response” and “I’ll figure it out myself”; that binary choice is the behavioral reliance measure. The within-subjects design exposes each participant to both a sociolect agent and a standard English agent, and a manipulation check ensures the participant perceived the sociolect as intended before their quantitative responses are counted. This mechanism isolates language style as the only difference while holding content, confidence expressions, and question difficulty fixed.

What would settle it

Run the same video-question study with two matched sets of sociolect stimuli, one set as in this paper and a second set generated by a different method or written by native speakers and pre-tested to be equal to the standard English set in warmth and confidence on the same scales. If the second set produces no reliance or trust difference between the agents, the claim that sociolect use per se changes reliance is falsified; if the second set reproduces the pattern, the phrase-generation route is exonerated.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that minoritized anthropomorphic cues do change user behavior and perception, but not in the direction of simple personalization. AAE participants relied more on the standard English agent than on the AAE-speaking one ($d = 0.12$, $p < 0.05$), trusted it more, were more satisfied with it, were less frustrated by it, and preferred it outright. Queer slang participants showed no significant preference or trust difference between agents, relied on the Queer slang agent only marginally less ($p = 0.09$), felt less frustration with the standard agent, yet reported significantly greater social presence from the Queer slang agent ($d = -0.23$, $p < 0.05$). Across both groups, reliance correlated weakly but significantly with trust, satisfaction, and, for the sociolect agents, social presence, and participants relied more on whichever agent they explicitly preferred. The paper reads this as evidence that LLM adaptation to minoritized language has emotional and social effects that are partly decoupled from trust and reliance.

Load-bearing premise

The load-bearing assumption is that the machine-translated sociolect phrases differ from their standard English counterparts only in sociolect, not in perceived warmth, confidence, naturalness, or formality; if those dimensions also shifted, the lower trust and reliance could come from the specific phrases rather than from sociolect use as such.

Editorial extensions

If this is right

  • Designers cannot assume that matching a user's sociolect will increase reliance or trust; in this study it moved reliance in the opposite direction for AAE speakers and left it unchanged for Queer slang speakers.
  • Perceived social presence and behavioral reliance can diverge: Queer slang speakers felt more connected to the sociolect agent while relying on it no more than the standard agent, so engagement metrics should not be read as trust metrics.
  • Testing behavioral outcomes, such as whether users accept suggestions, alongside self-reported perceptions is necessary, since self-reported preference alone would have missed part of the AAE group's reliance pattern.
  • The direction and size of sociolect effects differ across sociolects, with AAE and Queer slang producing different trust, preference, and social-presence profiles, so findings for one minoritized language variety should not be generalized to another without testing.
  • If LLM providers adapt to a user's dialect, they should weigh cultural-appropriation concerns and context, because a substantial share of participants described the sociolect output as forced, unnatural, or disrespectful.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I would predict that in a casual, social, or entertainment task rather than a factual question-answering task, the social-presence benefit of sociolect could translate into higher reliance, reversing the direction seen here; the paper's own context-dependence findings support testing this.
  • A testable extension: if the sociolect suggestions were written by native speakers rather than machine-translated, the trust gap might shrink or flip, because many negative comments targeted perceived inauthenticity rather than the sociolect itself.
  • The divergence between AAE and Queer slang outcomes may partly reflect different acquisition paths, with AAE often a first language tied to family and community while Queer slang is often adopted later, so identity ownership and perceived encroachment likely moderate the effect; the paper gestures at this but does not test it directly.
  • I would infer that LLM personalization to minoritized language should be opt-in and context-aware rather than automatic, given that both groups showed some positive affect toward the sociolect agent but preferred the standard agent for the task.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports two within-subjects user studies (498 AAE speakers and 487 Queer slang speakers; analyses on manipulation-check passers are n=399 and n=406) in which participants answered difficult video-based factual questions with suggestions from an LLM agent speaking either standard American English (SAE) or the participant's self-identified sociolect. The authors measure behavioral reliance (the proportion of trials on which the participant accepts the LLM's suggestion), perceptions (trust, satisfaction, frustration, social presence), pairwise agent preference with free-text rationales, and associations between these variables. They find that AAE speakers relied more on, trusted, preferred, and were more satisfied with the SAE agent; Queer slang speakers showed only a marginal reliance difference, no significant preference, but significantly greater social presence with the Queer slang agent and lower frustration with the SAE agent. The paper interprets these results as evidence that sociolect personalization does not straightforwardly improve reliance or perception and that reliance and social presence can diverge.

Significance. If the causal attribution holds, this is a timely and valuable contribution to FAccT and HCI: it provides large-sample behavioral evidence that personalizing an LLM to a user's minoritized sociolect can backfire, and that a sociolect-speaking agent can increase social presence without increasing reliance or trust. The study's strengths include its behavioral outcome measure rather than self-report alone, within-subjects design with counterbalancing, community-targeted recruitment, detailed appendices with the full stimulus sets, and mixed-methods qualitative coding that connects open-ended comments to the quantitative patterns. The main risk is that the stimuli are not verified to isolate sociolect identity from phrase-level differences in confidence, warmth, and naturalness; this must be addressed before the central claim is fully supported.

major comments (3)
  1. [4.3; Appendices I–K, M–O] The sociolect stimuli are not validated to match their SAE counterparts on confidence, warmth, naturalness, or formality. Section 4.3 states the design goal was 'to create LLM responses that elicited a sense of warmth and medium confidence,' and the SAE items use confidence expressions with reliability percentages between 30% and 70% (Appendix H). The AAE and Queer slang phrases were generated by GPT-4 and screened only for sociolect authenticity (at least 4 of 5 verifiers; Appendices J/N), and the manipulation check (Appendix G) only asked whether the agent sounded like it used the sociolect. The final phrase lists (Tables 5 and 9) show systematic shifts in epistemic markers and warmth phrases (e.g., SAE 'I would lean it’s...' vs. AAE 'Ima lean on it’s...'; SAE 'Of course, I’m fairly certain it’s...' vs. AAE 'Fa sho, kinda certain it’s...'; SAE 'Yes! I would lean it’s...' vs. Queer 'Yasss queen! I would lean it’s...'; SAE 'Well done, I would say it’s...' vs. Queer 'Work it, diva! I would say it’s...'). These changes plausibly alter perceived confidence, competence, and naturalness independently of sociolect identity. Because the reliance measure is behavioral and answer content is withheld, the lower reliance, trust, and satisfaction for the AAELM, and the weaker effects for the QSLM, could reflect the particular translated phrases rather than sociolect usage per se. Please add a norming study in which the target populations rate both sets of phrases on confidence, warmth, naturalness, and formality, or otherwise control for these dimensions (e.g., multiple translation sets or mixed-effects models with perceived confidence as a covariate).
  2. [Abstract; §5.1, Table 14] The abstract's summary statement overstates the Queer slang reliance result. It says 'both AAE and Queer slang speakers relied more on the SAE agent, and had more positive perceptions of the SAE agent,' but §5.1 and Table 14 report the Queer slang reliance difference as marginal (p = 0.085, d = 0.08) and the text explicitly states it 'did not reach statistical significance.' Likewise, for Queer slang speakers, trust and satisfaction differences are not significant (§5.2, Table 15), so 'more positive perceptions' is not supported on those measures. Please revise the abstract and any summary claims to state that AAE speakers significantly relied on, trusted, and preferred the SAE agent, while Queer slang speakers showed a marginal reliance effect and mixed perceptions (with significantly greater social presence for the Queer slang agent and lower frustration with the SAE agent).
  3. [§6 Discussion; §4.3] The Discussion generalizes from a single set of ten GPT-4-generated phrases per sociolect to conclusions about 'sociolect usage by LLMs.' Given the qualitative evidence that some participants found the AAELM's language unnatural, forced, or mocking (§5.4), the observed effects may be specific to this particular instantiation of AAE rather than to AAE as a sociolect. The paper already acknowledges in Section 7 that some phrases may sound artificial, but the main takeaway in Section 6 is phrased broadly ('personalization and anthropomorphic design of agents in this scenario can hinder...'). Please either temper the general claim to the specific stimuli used, or include multiple independently generated phrase sets per sociolect to support a stimulus-general conclusion.
minor comments (5)
  1. [§2.1] The phrase 'due to the its ties with Queer individuals' contains a typo; it should be 'due to its ties.'
  2. [§5.2] The sentence 'we performed paired t-test comparisons between the relevant variables reported for SAELM and AAELM' should read 'SAELM and QSLM' when describing the Queer slang analysis.
  3. [Appendix K] Appendix K says 'we achieved a diverse selection of 20 sociolect aligned phrases,' while Section 4.3 and Table 5 list 10; please make the counts consistent.
  4. [Abstract and throughout] The paper repeatedly uses 'a LLM' (e.g., in the abstract); it should be 'an LLM.'
  5. [General] Please consider adding a data availability statement and reporting confidence intervals for the reported effect sizes, since the effects are small (d ≈ 0.08–0.44) and confidence intervals would aid interpretation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the study is a new empirical experiment whose claims are tested against fresh participant data, not derived from its own inputs or prior self-citations.

full rationale

The paper's central claims are empirical findings from two within-subjects user studies (498 AAE speakers and 487 Queer slang speakers), measuring behavioral reliance and self-reported perceptions of simulated LLM agents. There is no derivation chain in which an output quantity is defined in terms of an input quantity, no fitted parameter is later renamed as a prediction, and no uniqueness theorem is invoked to force a conclusion. The sociolect stimuli were generated by GPT-4 from SAE templates and screened only for sociolect authenticity; this raises a legitimate construct-validity concern about possible confounds such as perceived confidence or naturalness differing across conditions, but that is a methodological weakness rather than circular reasoning, since the outcomes were not logically entailed by the stimulus construction procedure. The paper borrows the reliance measurement approach and warmth/confidence phrase inventories from prior work by Zhou et al., and it cites several of the authors' own earlier papers, but these are used as measurement tools and background motivation, not as evidence that the research hypotheses must hold. The central results—AAE speakers relying on and preferring the SAE agent, Queer slang speakers feeling more social presence from the Queer slang agent without preferring it—were obtained from new participant responses and are not implied by any fitted parameter or by the cited prior work. The limitations section explicitly acknowledges that participants may have found some sociolect phrases artificial and that code-switching may have reduced naturalness, further confirming that the authors treat the observed effects as contingent empirical outcomes. No circular step meeting the required evidentiary standard can be quoted from the paper, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No numerical free parameters are fitted; the results are estimated effects from a user experiment. The load-bearing assumptions concern participant identification, stimulus equivalence, and the reliance metric.

assumptions (3)
  • domain assumption Prolific demographic filters plus the self-report sociolect screener identify participants who genuinely speak AAE or Queer slang.
    Recruitment relies on Prolific self-identification and Likert self-reports of use and confidence; no independent behavioral verification is reported (Section 4.1, Appendix C).
  • ad hoc to paper The GPT-4 generated translations and persona-prompted phrases are valid instantiations of AAE and Queer slang, matched to their SAE counterparts on warmth and confidence.
    The design treats sociolect as the only difference between agents, but perceived confidence, naturalness, and formality of the translations are not directly measured (Section 4.3, Appendices I-K).
  • domain assumption Behavioral reliance, measured as the proportion of accepted LLM suggestions, is a valid operationalization of reliance.
    The metric is adopted from Bansal et al. [9] and Zhou et al. [121], but the task is single-turn and the scoring incentive was not actually calculated (Section 4.2).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Not Like Us, Hunty: Measuring Perceptions and Behavioral Effects of Minoritized Anthropomorphic Cues in LLMs." pith.science (2026). https://pith.science/paper/PNZNVRZD

@misc{pith2026250505660,
  author       = {Pith},
  title        = {Pith review of: Not Like Us, Hunty: Measuring Perceptions and Behavioral Effects of Minoritized Anthropomorphic Cues in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PNZNVRZD}},
  note         = {Machine review of arXiv:2505.05660}
}
read the original abstract

As large language models (LLMs) increasingly adapt and personalize to diverse sets of users, there is an increased risk of systems appropriating sociolects, i.e., language styles or dialects that are associated with specific minoritized lived experiences (e.g., African American English, Queer slang). In this work, we examine whether sociolect usage by an LLM agent affects user reliance on its outputs and user perception (satisfaction, frustration, trust, and social presence). We designed and conducted user studies where 498 African American English (AAE) speakers and 487 Queer slang speakers performed a set of question-answering tasks with LLM-based suggestions in either standard American English (SAE) or their self-identified sociolect. Our findings showed that sociolect usage by LLMs influenced both reliance and perceptions, though in some surprising ways. Results suggest that both AAE and Queer slang speakers relied more on the SAE agent, and had more positive perceptions of the SAE agent. Yet, only Queer slang speakers felt more social presence from the Queer slang agent over the SAE one, whereas only AAE speakers preferred and trusted the SAE agent over the AAE one. These findings emphasize the need to test for behavioral outcomes rather than simply assume that personalization would lead to a better and safer reliance outcome. They also highlight the nuanced dynamics of minoritized language in machine interactions, underscoring the need for LLMs to be carefully designed to respect cultural and linguistic boundaries while fostering genuine user engagement and trust.

Figures

Figures reproduced from arXiv: 2505.05660 by the authors.

Figure 1
Figure 1. Study flow diagram illustrating the two experimental scenarios. In both scenarios, participants begin by answering [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. RQ1: Participant reliance for different language [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Participant preference for different language styles: (a) Bar graph showing participant preferences for SAE and AAE, [PITH_FULL_IMAGE:figures/full_fig_p033_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

146 extracted references · 18 canonical work pages

  1. [1]

    Samy Alim, John R

    H. Samy Alim, John R. Rickford, and Arnetha F. Ball. 2016. Raciolinguistics: How Language Shapes Our Ideas About Race . Oxford University Press. https: //doi.org/10.1093/acprof:oso/9780190625696.001.0001

  2. [2]

    Marcellus Amadeus, Jose Roberto Homeli da Silva, and Joao Victor Pessoa Rocha

  3. [3]

    Celeste B. Amos. 2019. Understanding Correlations Between Standard American English Education and Systemic Racism and Strategies to Break the Cycle. International Journal of Research in Humanities and Social Studies 6 (2019), 1–4

  4. [4]

    Julian Kevon Glover and. 2016. Redefining Realness?: On Janet Mock, Laverne Cox, TS Madison, and the Representation of Transgender Women of Color in Media. Souls 18, 2-4 (2016), 338–357. https://doi.org/10.1080/10999949.2016. 1230824

  5. [5]

    Yes, BUT

    Kate T. Anderson, Chris Chang-Bacon, and Maria Guzmán Antelo. 2024. Navi- gating monolingual language ideologies: Educators’ “Yes, BUT” objections to lin- guistically sustaining pedagogies in the classroom.International Journal of Bilin- gualism 28, 4 (Aug. 2024), 618–634. https://doi.org/10.1177/13670069241236682

  6. [6]

    Sunil Arora, Sahil Arora, and John Hastings. 2024. The Psychological Impacts of Algorithmic and AI-Driven Social Media on Teenagers: A Call to Action . In 2024 IEEE Digital Platforms and Societal Harms (DPSH) . IEEE Computer Society, Los Alamitos, CA, USA, 1–7. https://doi.org/10.1109/DPSH60098.2024.10774922

  7. [7]

    Baker-Bell

    A. Baker-Bell. 2020. Linguistic Justice: Black Language, Literacy, Identity, and Pedagogy. Taylor & Francis

  8. [8]

    Nikola Banovic, Zhuoran Yang, Aditya Ramesh, and Alice Liu. 2023. Being Trustworthy is Not Enough: How Untrustworthy Artificial Intelligence (AI) Can Deceive the End-Users and Gain Their Trust. Proc. ACM Hum.-Comput. Interact. 7, CSCW1, Article 27 (April 2023), 17 pages. https://doi.org/10.1145/3579460

Show all 146 references
  1. [9]

    Lasecki, Daniel S

    Gagan Bansal, Besmira Nushi, Ece Kamar, Walter S. Lasecki, Daniel S. Weld, and Eric Horvitz. 2019. Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance. Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 7, 1, 2–11. https://doi.org/10....

  2. [10]

    Cunningham, Erica Adams, Alisha Bose, Aditi Jain, Kaus- tubh Yadav, Zhengyang Yang, Katharina Reinecke, and Daniela Rosner

    Jeffrey Basoah, Jay L. Cunningham, Erica Adams, Alisha Bose, Aditi Jain, Kaus- tubh Yadav, Zhengyang Yang, Katharina Reinecke, and Daniela Rosner. 2025. Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users. arXiv:2505.00821 [cs.HC] https://ar...

  3. [12]

    Karim Benharrak, Tim Zindulka, Florian Lehmann, Hendrik Heuer, and Daniel Buschek. 2024. Writer-Defined AI Personas for On-Demand Feedback Genera- tion. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association f...

  4. [13]

    Bazarova

    Aparajita Bhandari, Diana Freed, Tara Pilato, Faten Taki, Gunisha Kaur, Stephen Yale-Loehr, Jane Powers, Tao Long, and Natalya N. Bazarova. 2022. Multi- stakeholder Perspectives on Digital Tools for U.S. Asylum Applicants Seeking Healthcare and Legal Information. Proc. ACM Hum...

  5. [14]

    TOMÁŠ BOČEK. 2023. The Impact and Influence of African-American Vernacu- lar English on LGBTQ+ Culture. (2023)

  6. [15]

    Kasandra Brabaw. 2024. 17 Lesbian Slang Terms Every Baby Gay Needs To Learn. (2024). https://www.refinery29.com/en-us/lesbian-slang-terms-definitions Accessed: 2025-04-29

  7. [16]

    Virginia Braun and Victoria Clarke and. 2006. Using thematic analysis in psychology. Qualitative Research in Psychology 3, 2 (2006), 77–101. https: //doi.org/10.1191/1478088706qp063oa

  8. [17]

    William Cain. 2024. Prompting Change: Exploring Prompt Engineering in Large Language Model AI and Its Potential to Transform Education. TechTrends 68, 1 (Jan. 2024), 47–57. https://doi.org/10.1007/s11528-023-00896-0

  9. [18]

    Casey, Sari L

    Logan S. Casey, Sari L. Reisner, Mary G. Findling, Robert J. Blendon, John M. Benson, Justin M. Sayde, and Carolyn Miller. 2019. Discrimination in the United States: Experiences of lesbian, gay, bisexual, transgender, and queer Americans. Health Services Research 54 Suppl 2, S...

  10. [19]

    Kaiping Chen, Anqi Shao, Jirayu Burapacheep, and Yixuan Li. 2024. Conversa- tional AI and equity through assessing GPT-3’s communication with diverse social groups on contentious topics. Scientific Reports 14, 1 (Jan. 2024), 1561. https://doi.org/10.1038/s41598-024-51969-w

  11. [20]

    Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Anna Rogers, Jordan ...

  12. [21]

    Olanubi, Joseph M

    Michelle Cohn, Mahima Pushkarna, Gbolahan O. Olanubi, Joseph M. Moran, Daniel Padgett, Zion Mengesha, and Courtney Heldreth. 2024. Believing An- thropomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models. In Extended Abstracts of the CHI Confe...

  13. [22]

    Cornelius

    Brianna R. Cornelius. 2016. Gay Black men and the construction of identity via linguistic repertoires. In Proceedings of the 24th Annual Symposium about Language and Society-Austin

  14. [23]

    Yass girl. Spill the tea. Throw some shade

    Mo Crutzen. 2021. "Yass girl. Spill the tea. Throw some shade. ": Translation Procedures Used in the Translation of Drag and Gay Vocabulary in RuPaul’s Drag Race via Subtitling. https://hdl.handle.net/1887/3254737 Unpublished manuscript. Retrieved from author or academic repository

  15. [24]

    Jamell Dacon. 2022. Towards a Deep Multi-layered Dialectal Language Anal- ysis: A Case Study of African-American English. In Proceedings of the Second Workshop on Bridging Human–Computer Interaction and Natural Language Pro- cessing, Su Lin Blodgett, Hal Daumé III, Michael Mad...

  16. [25]

    Standard

    Bethany Davila. 2016. The Inevitability of “Standard” English: Discursive Con- structions of Standard Language Ideologies. Written Communication 33, 2 (2016), 127–148. https://doi.org/10.1177/0741088316632186

  17. [26]

    Intelligence,

    Roberto Santiago de Roock. 2024. To Become an Object Among Objects: Genera- tive Artificial “Intelligence, ” Writing, and Linguistic White Supremacy.Reading Research Quarterly 59, 4 (2024), 590–608. https://doi.org/10.1002/rrq.569

  18. [27]

    Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen McKeown. 2023. Evaluation of African American Language Bias in Natural Language Generation. (Dec. 2023), 6805–6824. https://doi.org/10. 18653/v1/2023.emnlp-main.421

  19. [28]

    Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, and Ashwin Kalyan. 2023. Anthropomorphization of AI: Opportunities and Risks. (Dec. 2023), 1–7. https://doi.org/10.18653/v1/2023.nllp-1.1

  20. [29]

    Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell. 2023. Queer people are people first: Deconstructing sexual identity stereotypes in large language models. arXiv preprint arXiv:2307.00101 (2023)

  21. [30]

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods...

  22. [31]

    Fahim Faisal, Md Mushfiqur Rahman, and Antonios Anastasopoulos. 2024. Di- alectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Lan- guage Varieties. (Nov. 2024). https://doi.org/10.48550/arXiv.2411.10954

  23. [32]

    Virginia Felkner, Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May

  24. [33]

    Finch, Ellie S

    Sarah E. Finch, Ellie S. Paek, Sejung Kwon, Ikseon Choi, Jessica Wells, Rasheeta Chandler, and Jinho D. Choi. 2025. Finding a Voice: Evaluating African American Dialect Generation for Chatbot Technology. arXiv preprint arXiv:2501.03441 (2025). https://arxiv.org/abs/2501.03441

  25. [34]

    I wouldn’t say offensive but

    Vinitha Gadiraju, Shaun Kane, Sunipa Dev, Alex Taylor, Ding Wang, Emily Denton, and Robin Brewer. 2023. "I wouldn’t say offensive but... ": Disability- Centered Perspectives on Large Language Models. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Tr...

  26. [35]

    David Gefen, Elena Karahanna, and Detmar W. Straub. 2003. Trust and TAM in online shopping: an integrated model. MIS Q. 27, 1 (March 2003), 51–90

  27. [36]

    Katy Ilonka Gero, Tao Long, and Lydia B Chilton. 2023. Social Dynamics of AI Support in Creative Writing. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Association for Computing Machinery, New York, NY, USA, Articl...

  28. [37]

    Aashish Ghimire. 2024. Generative AI in Education From the Perspective of Students, Educators, and Administrators. All Graduate Theses and Dissertations, Fall 2023 to Present (May 2024). https://doi.org/10.26076/c582-c0bc

  29. [38]

    Zohar Gilad, Ofra Amir, and Liat Levontin. 2021. The Effects of Warmth and Competence Perceptions on Users’ Choice of an AI System. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21). Association for Computing Machinery,...

  30. [39]

    Lisa J. Green. 2002. African American English: A Linguistic Introduction . Cam- bridge University Press

  31. [40]

    Gresczyk, Richard A

    Sr. Gresczyk, Richard A. 2011.Language warriors; Leaders in the Ojibwe language revitalization movement. Ph. D. Dissertation. https://hdl.handle.net/11299/ 104677

  32. [41]

    J.A. Grieser. 2022. The Black Side of the River: Race, Language, and Belonging in Washington, DC. Georgetown University Press

  33. [42]

    Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020. Investigating African-American Vernacular English in Transformer-Based Text Generation. (Nov. 2020), 5877–

  34. [43]

    Abhay Gupta, Ece Yurtseven, Philip Meng, and Kevin Zhu. 2024. AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark. (Nov. 2024), 327–333. https://doi.org/10.18653/v1/2024.nlp4pi-1.28

  35. [44]

    Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, Ujwal Gadiraju, Aditya Vashistha, Vivek Seshadri, and Kalika Bali. 2024. Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology. In The 2024 AC...

  36. [45]

    Megan-Brette Hamilton. 2020. Real life lessons in language and literacy: Discov- ering the role of identity for AAE-speakers. Journal of Language and Literacy, UGA (2020)

  37. [46]

    Hancock, Mor Naaman, and Karen Levy

    Jeffrey T. Hancock, Mor Naaman, and Karen Levy. 2020. AI-Mediated Com- munication: Definition, Research Agenda, and Ethical Considerations. Journal of Computer-Mediated Communication 25, 1 (2020), 89–100. https://doi.org/10. 1093/jcmc/zmz022

  38. [47]

    Camille Harris, Matan Halevy, Ayanna Howard, Amy Bruckman, and Diyi Yang

  39. [48]

    Hart and Lowell E

    Sandra G. Hart and Lowell E. Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of Empirical and Theoretical Research. 52 (1988), 139–183. https://doi.org/10.1016/S0166-4115(08)62386-9

  40. [49]

    Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King. 2024. AI generates covertly racist decisions about people based on their dialect.Nature 633, 8028 (Sept. 2024), 147–154. https://doi.org/10.1038/s41586-024-07856-5

  41. [50]

    Jonathan Ivey, Shivani Kumar, Jiayu Liu, Hua Shen, Sushrita Rakshit, Rohan Raju, Haotian Zhang, Aparna Ananthasubramaniam, Junghwan Kim, Bowen Yi, Dustin Wright, Abraham Israeli, Anders Giovanni Møller, Lechen Zhang, and David Jurgens. 2024. Real or Robotic? Assessing Whether ...

  42. [51]

    Yi Jiang, Xiangcheng Yang, and Tianqi Zheng. 2023. Make chatbots more adaptive: Dual pathways linking human-like cues and tailored response to trust in interactions with chatbots. Comput. Hum. Behav. 138, C (Jan. 2023), 15 pages. https://doi.org/10.1016/j.chb.2022.107485

  43. [52]

    Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2015. Challenges of studying and processing dialects in social media. In Proceedings of the Workshop on Noisy User-generated Text, Wei Xu, Bo Han, and Alan Ritter (Eds.). Association for Computational Linguistics, Beijing, China, ...

  44. [53]

    Bagas Dwi Kameswara and Agustinus Hary Setyawan. 2024. HEARER- ORIENTED ANALYSIS OF DRAG QUEEN SLANG IN THE WEB-SERIES UN- HHHH. Jurnal Review Pendidikan dan Pengajaran 7, 3 (Aug. 2024), 11300–11307. https://doi.org/10.31004/jrpp.v7i3.32533

  45. [54]

    Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the Snark: Annotator Diversity in Data Practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ’23). Asso- ciation for Computing Machinery, New York, NY, ...

  46. [55]

    Tyler Kendall and Charlie Farrington. 2023. The Corpus of Regional African American Language. https://doi.org/10.7264/1ad5-6t35

  47. [56]

    Kiesling

    Scott F. Kiesling. 2019. Language, Gender, and Sexuality: An Introduction. Rout- ledge, London. https://doi.org/10.4324/9781351042420

  48. [57]

    I’m Not Sure, But

    Sunnie S. Y. Kim, Q. Vera Liao, Mihaela Vorvoreanu, Stephanie Ballard, and Jennifer Wortman Vaughan. 2024. "I’m Not Sure, But... ": Examining the Impact of Large Language Models’ Uncertainty Expression on User Reliance and Trust. In Proceedings of the 2024 ACM Conference on Fa...

  49. [58]

    Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A. Hale. 2023. Personalisation within Bounds: A Risk Taxonomy and Policy Framework for the Alignment of Large Language Models with Personalised Feedback. arXiv:2303.05453 [cs.CL] https://arxiv.org/abs/2303.05453

  50. [59]

    Martinez

    Katharina Klein and Luis F. Martinez. 2023. The impact of anthropomorphism on customer satisfaction in chatbot commerce: an experimental study in the food sector. Electronic Commerce Research 23, 4 (2023), 2789–2825. https: //doi.org/10.1007/s10660-022-09562-8

  51. [60]

    Grieser, Shug Miller, James Shepard, Javier Garcia- Perez, Nick Deas, Desmond U

    Shana Kleiner, Jessica A. Grieser, Shug Miller, James Shepard, Javier Garcia- Perez, Nick Deas, Desmond U. Patton, Elsbeth Turcan, and Kathleen McKeown

  52. [61]

    Kacper Krudysz. 2023. Translating Queer Slang - An Analysis of Subtitles in RuPaul’s Drag Race and Legendary Television Programmes. (Oct. 2023). https: //ruj.uj.edu.pl/entities/publication/911d79b2-03dc-4032-a46a-9a3b466409a8

  53. [62]

    Rachel Laing. 2021. Who Said It First? : Linguistic Appropriation of Slang Terms within the Popular Lexicon . https://doi.org/10.30707/ETD2021. 20210719070603178888.59

  54. [63]

    DonHee Lee and Seong No Yoon. 2021. Application of Artificial Intelligence- Based Technologies in the Healthcare Industry: Opportunities and Challenges. International Journal of Environmental Research and Public Health 18, 1 (2021),

  55. [64]

    Lee, Jacob M

    Messi H.J. Lee, Jacob M. Montgomery, and Calvin K. Lai. 2024. Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (Rio ...

  56. [65]

    AI and Ethics (December 2024), 1–9

    Unmasking camouflage: Exploring the challenges of large language models in deciphering African American language & online performativity. AI and Ethics (December 2024), 1–9. https://doi.org/10.1007/s43681-024-00623-2

  57. [66]

    Marcin Lewandowski. 2008. The Language of Soccer – a Sociolect or a Register? (2008). http://hdl.handle.net/10593/4562

  58. [67]

    Vivian Liu, Tao Long, Nathan Raw, and Lydia Chilton. 2023. Generative Disco: Text-to-Video Generation for Music Visualization. arXiv preprint arXiv:2304.08551 (April 2023). https://doi.org/10.48550/arXiv.2304.08551

  59. [68]

    1997.Queerly Phrased: Language, Gender, and Sexuality

    Anna Livia and Kira Hall. 1997.Queerly Phrased: Language, Gender, and Sexuality. Oxford University Press, New York

  60. [69]

    Tao Long, Katy Ilonka Gero, and Lydia B. Chilton. 2024. Not Just Novelty: A Longitudinal Study on Utility and Customization of an AI Workflow. InProceed- ings of the 2024 ACM Designing Interactive Systems Conference (Copenhagen, Denmark) (DIS ’24). Association for Computing Ma...

  61. [70]

    Tao Long, Dorothy Zhang, Grace Li, Batool Taraif, Samia Menon, Kynnedy Si- mone Smith, Sitong Wang, Katy Ilonka Gero, and Lydia B. Chilton. 2023. Twee- torial Hooks: Generative AI Tools to Motivate Science on Social Media. In Proceedings of the 14th Conference on Computational...

  62. [71]

    Daisy D. Leigh. 2021. Style in Time: Online Perception of Sociolinguistic Cues . 326 pages. https://www.proquest.com/dissertations-theses/style-time-online- perception-sociolinguistic-cues/docview/2624983220/se-2 Copyright - Database copyright ProQuest LLC; ProQuest does not c...

  63. [72]

    Kathryn M Luyt. 2014. Gay language in Cape Town: a study of Gayle - atti- tudes, history and usage. (2014). http://hdl.handle.net/11427/6792 Publisher: University of Cape Town

  64. [73]

    Zilin Ma, Yiyang Mei, Yinru Long, Zhaoyuan Su, and Krzysztof Z. Gajos. 2024. Evaluating the Experience of LGBTQ+ People Using Large Language Model Based Chatbots for Mental Health Support. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolul...

  65. [74]

    Stephen L. Mann. 2011. Drag Queens’ Use of Language and the Performance of Blurred Gendered and Racial Identities. Journal of Homosexuality 58, 6–7 (2011), 793–811. https://doi.org/10.1080/00918369.2011.581923

  66. [75]

    Nina Markl. 2023. Language variation, automatic speech recognition and algo- rithmic bias. (Dec. 2023). https://doi.org/10.7488/era/4013 Accepted: 2023-12- 12T12:16:33Z Publisher: The University of Edinburgh. Measuring Perceptions and Behavioral Effects of Minoritized Anthropo...

  67. [76]

    Erich Hatala Matthes. 2019. Cultural appropriation and oppression.Philosophical Studies 176, 4 (April 2019), 1003–1013. https://doi.org/10.1007/s11098-018-1224- 2

  68. [77]

    Cristina Luna-Jiménez, Manuel Gil-Martín, Luis Fernando D’Haro, Fernando Fernández-Martínez, and Rubén San-Segundo. 2024. Evaluating emotional and subjective responses in synthetic art-related dialogues: A multi-stage framework with large language models. Expert Syst. Appl. 25...

  69. [78]

    Katelyn Mei, Sonia Fereidooni, and Aylin Caliskan. 2023. Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (Chicago, IL, USA) (FAcc...

  70. [79]

    I don’t think these devices are very culturally sensitive

    Zion Mengesha, Courtney Heldreth, Michal Lahav, Juliana Sublewski, and Elyse Tuennerman. 2021. "I don’t think these devices are very culturally sensitive. ": The impact of errors on African Americans in Automated Speech Recognition. Frontiers in Artificial Intelligence 26 (202...

  71. [80]

    Anna Milanez. 2023. The impact of AI on the workplace: Evidence from OECD case studies of AI implementation. 289 (March 2023). https://doi.org/10.1787/ 2247ce58-en

  72. [81]

    Joel Mire, Zubin Trivadi Aysola, Daniel Chechelnitsky, Nicholas Deas, Chrysoula Zerva, and Maarten Sap. 2025. Rejected Dialects: Biases Against African American Language in Reward Models. (April 2025), 7468–7487. https: //aclanthology.org/2025.findings-naacl.417/

  73. [82]

    Penelope Muzanenhamo and Sean Bradley Power. 2024. ChatGPT and account- ing in African contexts: Amplifying epistemic injustice. Critical Perspectives on Accounting 99 (2024), 102735. https://doi.org/10.1016/j.cpa.2024.102735

  74. [83]

    Harrison McKnight, Vivek Choudhury, and Charles Kacmar

    D. Harrison McKnight, Vivek Choudhury, and Charles Kacmar. 2002. The impact of initial consumer trust on intentions to transact with a web site: a trust building model. The Journal of Strategic Information Systems 11, 3 (2002), 297–323. https://doi.org/10.1016/S0963-8687(02)00020-3

  75. [84]

    I Searched for a Religious Song in Amharic and Got Sexual Content Instead

    Hellina Hailu Nigatu and Inioluwa Deborah Raji. 2024. "I Searched for a Religious Song in Amharic and Got Sexual Content Instead": Investigating Online Harm in Low-Resourced Languages on YouTube. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transp...

  76. [85]

    Ruby Ostrow and Adam Lopez. 2025. LLMs Reproduce Stereotypes of Sexual and Gender Minorities. (2025). arXiv:2501.05926 [cs.CL] https://arxiv.org/abs/ 2501.05926

  77. [86]

    Jiao Ou, Junda Lu, Che Liu, Yihong Tang, Fuzheng Zhang, Di Zhang, and Kun Gai. 2024. DialogBench: Evaluating LLMs as Human-like Dialogue Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag...

  78. [87]

    I’m fully who I am

    Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. “I’m fully who I am”: Towards Centering Transgender and Non-Binary Voices to Measure Bi- ases in Open Language Generation. In Proceedings of the 20...

  79. [88]

    Fred Paas, Alexander Renkl, and John Sweller. 2003. Cognitive Load Theory and Instructional Design: Recent Developments. Educational Psychologist 38, 1 (2003), 1–4. https://doi.org/10.1207/S15326985EP3801_1

  80. [89]

    Pranav Narayanan Venkit, Sanjana Gautam, Ruchi Panchanadikar, Ting-Hao Huang, and Shomir Wilson. 2023. Nationality Bias in Text Generation. (May 2023), 116–122. https://doi.org/10.18653/v1/2023.eacl-main.9

  81. [90]

    Saurabh Kumar Pandey, Harshit Budhiraja, Sougata Saha, and Monojit Choud- hury. 2025. CULTURALLY YOURS: A Reading Assistant for Cross-Cultural Content. In Proceedings of the 31st International Conference on Computational Linguistics: System Demonstrations , Owen Rambow, Leo Wa...

  82. [91]

    Pittman, Lynette O’Neal, Kimberly Wright, and Brittany R

    Ramona T. Pittman, Lynette O’Neal, Kimberly Wright, and Brittany R. White

  83. [92]

    Margaret Jane Pitts and Cindy Gallois. 2019. Social Markers in Language and Speech. Oxford University Press. https://doi.org/10.1093/acrefore/ 9780190236557.013.300

  84. [93]

    Rickford and R.J

    J.R. Rickford and R.J. Rickford. 2007. Spoken Soul: The Story of Black English . Turner Publishing Company

  85. [94]

    Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani. 2024. LaMP: When Large Language Models Meet Personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational ...

  86. [95]

    Stefan Palan and Christian Schitter. 2018. Prolific.ac—A subject pool for online experiments. Journal of Behavioral and Experimental Finance 17 (2018), 22–27. https://doi.org/10.1016/j.jbef.2017.12.004

  87. [96]

    Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection. InProceedings of the 2022 Conference of the North American Chapter of the Association ...

  88. [97]

    Christoph A. Schütt. 2023. The effect of perceived similarity and social proximity on the formation of prosocial preferences. Journal of Economic Psychology 99 (2023), 102678. https://doi.org/10.1016/j.joep.2023.102678

  89. [98]

    Education Sciences 14, 11 (2024), 1191

    Elevating Students’ Oral and Written Language: Empowering African American Students Through Language. Education Sciences 14, 11 (2024), 1191. https://doi.org/10.3390/educsci14111191

  90. [99]

    Martin, A

    Eleanor Shearer, S. Martin, A. Petheram, and R. Stirling. 2019. Racial bias in natural language processing. Oxford Insights (2019)

  91. [100]

    Gustavo Simas and Vânia Ribas Ulbricht. 2024. Human-AI Interaction: An Anal- ysis of Anthropomorphization and User Engagement in Conversational Agents with a Focus on ChatGPT. 119 (2024). https://doi.org/10.54941/ahfe1004510 ISSN: 27710718 Issue: 119

  92. [101]

    Gary Simes. 2005. Gay Slang Lexicography: A Brief History and a Commentary on the First Two Gay Glossaries.Dictionaries: Journal of the Dictionary Society of North America 26, 1 (2005), 1–159. https://muse.jhu.edu/pub/153/article/458430 Publisher: Dictionary Society of North America

  93. [102]

    Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith. 2019. The Risk of Racial Bias in Hate Speech Detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korho- nen, David Traum, and Lluís Màrquez (Eds.)....

  94. [103]

    Genevieve Smith, Eve Fleisig, Madeline Bossi, Ishita Rustagi, and Xavier Yin. 2024. Standard Language Ideology in AI-Generated Language. (2024). arXiv:2406.08726 [cs.CL] https://arxiv.org/abs/2406.08726

  95. [104]

    Laura Spillner and Nina Wenig. 2021. Talk to Me on My Level – Linguistic Alignment for Chatbots. In Proceedings of the 23rd International Conference on Mobile Human-Computer Interaction . ACM, Toulouse & Virtual France, 1–12. https://doi.org/10.1145/3447526.3472050

  96. [105]

    Woosuk Seo, Chanmo Yang, and Young-Ho Kim. 2024. ChaCha: Leveraging Large Language Models to Prompt Children to Share Their Emotions about Personal Events. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Associatio...

  97. [106]

    John Sweller. 1988. Cognitive load during problem solving: Effects on learning. Cognitive Science 12, 2 (1988), 257–285. https://doi.org/10.1207/ s15516709cog1202_4

  98. [107]

    Kelly is a Warm Person, Joseph is a Role Model

    Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters. (Dec. 2023), 3730–3748. https://doi.org/10. 18653/v1/2023.findings-emnlp.243

  99. [108]

    Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces (Greenville, SC, USA) (IUI ’24). Ass...

  100. [109]

    Simons and Christopher F

    Daniel J. Simons and Christopher F. Chabris. 1999. Gorillas in Our Midst: Sustained Inattentional Blindness for Dynamic Events. Perception 28, 9 (1999), 1059–1074. https://doi.org/10.1068/p281059

  101. [110]

    Lu Wang, Max Song, Rezvaneh Rezapour, Bum Chul Kwon, and Jina Huh- Yoo. 2024. People’s Perceptions Toward Bias and Related Concepts in Large Language Models: A Systematic Review. (2024). arXiv:2309.14504 [cs.HC] https://arxiv.org/abs/2309.14504

  102. [111]

    Sitong Wang, Samia Menon, Tao Long, Keren Henderson, Dingzeyu Li, Kevin Crowston, Mark Hansen, Jeffrey V Nickerson, and Lydia B Chilton. 2024. Reel- Framer: Human-AI Co-Creation for News-to-Video Translation. In Proceedings of the 2024 CHI Conference on Human Factors in Comput...

  103. [112]

    Julia P. Stanley. 1970. Homosexual slang. American Speech 45, 1/2 (1970), 45–59

  104. [113]

    Richard Wiseman. 2007. Quirkology: The Curious Science of Everyday Lives . Macmillan, New York, NY

  105. [114]

    Walt Wolfram. 2004. Social Varieties of American English. Cambridge University Press, 58–75. https://doi.org/10.1017/CBO9780511809880.006

  106. [115]

    Walt Wolfram and Mary E. Kohn. 2015. Regionality in the Development of African American English. In The Oxford Handbook of African American Lan- guage, Jennifer Bloomquist, Lisa J. Green, and Sonja L. Lanehart (Eds.). Oxford FAccT ’25, June 23–26, 2025, Athens, Greece Basoah e...

  107. [117]

    Meredith G. F. Worthen. 2023. Queer identities in the 21st century: Reclamation and stigma. Current Opinion in Psychology 49 (Feb. 2023), 101512. https: //doi.org/10.1016/j.copsyc.2022.101512

  108. [118]

    I Like Sunnie More Than I Expected!

    Siyi Wu, Julie Y. A. Cachia, Feixue Han, Bingsheng Yao, Tianyi Xie, Xuan Zhao, and Dakuo Wang. 2024. "I Like Sunnie More Than I Expected!": Exploring User Expectation and Perception of an Anthropomorphic LLM-based Conversational Agent for Well-Being Support. (2024). arXiv:2405...

  109. [119]

    Adam Waytz, John Cacioppo, and Nicholas Epley. 2010. Who sees human? The stability and importance of individual differences in anthropomorphism. Perspectives on Psychological Science 5, 3 (2010), 219–232

  110. [120]

    Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. 2021. Detoxifying Language Models Risks Marginalizing Minority Voices. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: H...

  111. [121]

    Kaitlyn Zhou, Jena Hwang, Xiang Ren, and Maarten Sap. 2024. Relying on the Unreliable: The Impact of Language Models’ Reluctance to Express Uncertainty. (Aug. 2024), 3623–3643. https://doi.org/10.18653/v1/2024.acl-long.198

  112. [122]

    Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, and Maarten Sap

    Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, and Maarten Sap. 2025. REL-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance. , 11148–11167 pages. https://aclanthology.org/2025.naacl- long.556/

  113. [123]

    Fox, Steven Rousso-Schindler, and Jeffrey Warshaw

    Allison Woodruff, Sarah E. Fox, Steven Rousso-Schindler, and Jeffrey Warshaw

  114. [124]

    Caleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson, and Diyi Yang. 2022. VALUE: Understanding Dialect Disparity in NLU. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Ling...

  115. [125]

    How frequently do you use large language models (ChatGPT, Claude, Gemini...)

    Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. Multi-VALUE: A Framework for Cross-Dialectal English NLP. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Association...

  116. [127]

    Yuguang Xie, Kaige Zhu, Peiyu Zhou, and Chen Liang. 2023. How does anthro- pomorphism improve human-AI interaction satisfaction: A dual-path model. Computers in Human Behavior 148 (2023), 107878. https://doi.org/10.1016/j.chb. 2023.107878

  117. [131]

    Gates, Maarten Sap, and Thomas Hartvigsen

    Runtao Zhou, Guangya Wan, Saadia Gabriel, Sheng Li, Alexander J. Gates, Maarten Sap, and Thomas Hartvigsen. 2025. Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English. arXiv preprint arXiv:2503.04099 (2025). https://arxiv.org/abs/2503.04099

  118. [134]

    Which hand was the man wearing a watch? ⃝ Very difficult ⃝ Difficult ⃝ Somewhat difficult ⃝ Neither easy nor difficult ⃝ Somewhat easy ⃝ Easy ⃝ Very easy

  119. [135]

    Which hand was the woman wearing a watch? ⃝ Very difficult ⃝ Difficult ⃝ Somewhat difficult ⃝ Neither easy nor difficult ⃝ Somewhat easy ⃝ Easy ⃝ Very easy

  120. [136]

    Not at all

    How many rings was the woman wearing? ⃝ Very difficult ⃝ Difficult ⃝ Somewhat difficult ⃝ Neither easy nor difficult ⃝ Somewhat easy ⃝ Easy ⃝ Very easy (...... more questions......) F FOLLOW UP QUESTIONNAIRE Follow up questionnaire for AAELM How satisfied are you with the perf...

  121. [137]

    What color are the man’s eyes? Agent Blue’s Response: I’m wit it, think it’s... ⃝ Use Agent Blue’s response Measuring Perceptions and Behavioral Effects of Minoritized Anthropomorphic Cues in LLMs FAccT ’25, June 23–26, 2025, Athens, Greece ⃝ I’ll figure it out myself

  122. [138]

    ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself

    Is the queen of diamonds or the queen of hearts adjacent to the card the woman selects? Agent Blue’s Response: Fa sho, I feel like it’s... ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself

  123. [139]

    Gorillas in Our Midst

    How many rings was the woman wearing? Agent Blue’s Response: A’ight, I’d say it’s. . . ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself (....... 7 more questions .......) L.2 Excerpt of “Gorillas in Our Midst” Video Questions for SAE in AAELM Setup : Instructions: (1) Y...

  124. [140]

    ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

    How many times did someone with white shoes touch a ball? Agent Red’s Response: I’d be happy to help you with that, I believe it’s. . . ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

  125. [141]

    ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

    How many times did the two ball trajectories cross each other? Agent Red’s Response: Of course, I think it’s. . . ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

  126. [142]

    Gorillas in Our Midst

    How many shoes had white laces and how many shoes had black laces? Agent Red’s Response: Good day, I would answer it’s. . . ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself (....... 7 more questions .......) M QUEER W ARMTH PHRASE GENERATION AND CONFIDENCE EXPRESSIONS To...

  127. [143]

    ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

    What color are the man’s eyes? Agent Red’s Response: Well done, I would say it’s... ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

  128. [144]

    ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

    Is the queen of diamonds or the queen of hearts adjacent to the card the woman selects? Agent Red’s Response: Yes! I would lean it’s... ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself

  129. [145]

    Gorillas in Our Midst

    How many rings was the woman wearing? Agent Red’s Response: Hello there, I think it’s. . . ⃝ Use Agent Red’s response ⃝ I’ll figure it out myself (....... 7 more questions .......) P.2 Excerpt of “Gorillas in Our Midst” Video Questions for Queer Slang in QSLM Setup : Instructi...

  130. [146]

    ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself

    How many times did someone with white shoes touch a ball? Agent Blue’s Response: Work it, diva! I would say it’s. . . ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself

  131. [147]

    ⃝ Use Agent Blue’s response FAccT ’25, June 23–26, 2025, Athens, Greece Basoah et al

    How many times did the two ball trajectories cross each other? Agent Blue’s Response: Yasss queen! I would lean it’s. . . ⃝ Use Agent Blue’s response FAccT ’25, June 23–26, 2025, Athens, Greece Basoah et al. ⃝ I’ll figure it out myself

  132. [148]

    ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself (

    How many shoes had white laces and how many shoes had black laces? Agent Blue’s Response: Hello superstar, I think it’s. . . ⃝ Use Agent Blue’s response ⃝ I’ll figure it out myself (....... 7 more questions .......) Q PARTICIPANT DEMOGRAPHIC TABLE (REMOVED THOSE WHO FAILED MAN...

  133. [271]

    https://doi.org/10.3390/ijerph18010271

  134. [2018]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18)

    A Qualitative Exploration of Perceptions of Algorithmic Fairness. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3173574.3174230

  135. [2022]

    InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22)

    Exploring the Role of Grammar and Word Choice in Bias Toward African American English (AAE) in Hate Speech Classification. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (Seoul, Republic of Korea) (FAccT ’22). Association for Computing M...

  136. [2023]

    (July 2023), 9126–9140

    WinoQueer: A Community-in-the-Loop Benchmark for Anti-LGBTQ+ Bias in Large Language Models. (July 2023), 9126–9140. https://doi.org/10. 18653/v1/2023.acl-long.507

  137. [2024]

    Bridging the Language Gap: Integrating Language Variations into Con- versational AI Agents for Enhanced User Engagement. In Proceedings of the 1st Worskhop on Towards Ethical and Inclusive Conversational AI: Language Attitudes, Linguistic Diversity, and Language Rights (TEICAI...

  138. [5883]

    https://doi.org/10.18653/v1/2020.emnlp-main.473

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.