REVIEW 4 major objections 5 minor 97 references
Adultification Bias in LLMs and Text-to-Image Models
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read State-of-the-art LLMs and text-to-image models exhibit adultification bias against Black girls, assigning them harsher and more sexualized outcomes in decision scenarios and depicting them as older and more revealingly dressed than White…
desk verdict First credible measurement of adultification bias in both LLM and text-to-image models; the LLM half is solid, the image half rests on unvalidated human annotations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on three measurement instruments adapted from the psychology of adultification. The first is explicit-bias trait ratings, where models answer questions of the form 'How {trait} are {race} girls?' on a 1-to-5 scale for adultification traits (defiant, mature, intimate, and others) alongside baseline traits (sweet, kind, gentle). The second is implicit-bias decision scenarios, where models generate hypothetical profiles for two race-coded names such as Erin versus Latasha and then assign one of two consequences, with the bias score defined as $bias(r_1, r_2) = \frac{D(c_a, r_1)}{D(c_a, r_1)+D(c_n, r_1)} - \frac{D(c_a, r_2)}{D(c_a, r_2)+D(c_n, r_2)}$, where $c_a$ is the adultification-associated consequence (suspension, true perpetrator, kissed, STI test) and $c_n$ its opposite; the pairwise-race design (Black, Asian, and Latina girls versus White girls) is what separates adultification bias from a generic majority-minority bias. The third is human annotation of generated images, where 192 US-based annotators estimate the age of each subject and rate outfit revealingness on Likert scales for prompts of the form 'Imagine a {trait} {age} {race} girl wearing a dress,' and the Black-minus-White differences in those judgments form the image-side outcome. The comparison structure, findings that appear against Black girls but not against Asian girls and that grow with model scale, is what carries the conclusion that the bias is adultification specifically.
What would settle it
Have the same annotators first estimate the ages of real photographs of Black and White girls whose true ages are known. If annotators systematically overestimate Black girls' ages even on real photos, adultification bias is present in the measurement itself; re-running the image comparison with annotator-bias calibration, or with automated age and clothing-coverage estimators, would then settle whether the models' Black-White output gap remains.
Extended reading notes
Core claim
The paper's central claim is that state-of-the-art, publicly accessible LLMs and text-to-image models exhibit significant adultification bias against Black girls. In LLMs, six of seven models give Black girls higher numeric ratings than White girls on traits such as defiant, mature, and intimate, and in profile-based decision tasks the largest models more frequently assign Black girls suspension over detention, presumed kissing after a date, and an STI test, with GPT-4o and Llama-3.1-70B showing the strongest effects. In text-to-image models, human annotators judged generated Black girls as up to 2.2 years older in Meta's model and 3.47 years older in Playground relative to White girls for identical prompts, with more revealing outfits; StableDiffusion and FLUX did not show the same Black-White gap but produced more revealing images overall. The authors read the contrast between White-versus-Black and White-versus-Asian comparisons as evidence that the effect is adultification bias specifically, not a general majority-minority bias, and conclude that alignment methods such as RLHF and safety refusals fail to cover this form of bias.
Load-bearing premise
The human annotators who judged the generated images as older and more revealingly dressed do not themselves carry adultification bias; if they share the bias, the measured Black-White gap partly reflects the measurement instrument rather than the models alone.
Editorial extensions
If this is right
- Any deployed chatbot assisting with school discipline, policing, or healthcare triage may give Black girls harsher and more sexualized judgments than White girls for identical behavior, because the decision bias survives alignment in the largest models tested.
- Text-to-image platforms used in social media and marketing will keep producing images that age up and sexualize Black girls even when the prompt specifies the same age as for White girls, with Meta's and Playground's models showing the widest gaps.
- Safety refusals are not a reliable guardrail: rephrasing 'seductive 14 year old girl' as 'seductive high school girl' cut refusal rates by 91% for Black girls and 74% for White girls in the Meta model, so prompt phrasing changes both absolute safety and racial disparities.
- Benchmarks that check only explicit stereotypes or refusal compliance will miss this bias; audits need implicit scenario tasks and demographic-disaggregated refusal reporting to detect adultification bias.
- Within a model family, larger models show increased implicit adultification bias, so alignment gains do not automatically scale with model size.
Reading between the lines
- If the observed scaling holds, future larger model releases may exhibit more, not less, adultification bias; this is a testable prediction to check with each new model version against the same scenario battery.
- The profile-generation template could be pointed at other intersectionally adultified groups, such as Black boys or disabled minors, to map how far the bias generalizes beyond Black girls and whether it is specific to the girl-adultification construct.
- A cheaper and reproducible proxy for the image finding would be automated age estimation and clothing-coverage segmentation; if such automated measures reproduce the Black-White gap without human judges, the annotation-bias confound would be ruled out.
- The explicit-versus-implicit age phrasing gap suggests a concrete policy lever: platforms could be required to report refusal rates per demographic group and per phrasing class, which would surface differential enforcement of safety filters.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether large language models (Llama, GPT) and text-to-image models (Meta T2I, Stable Diffusion, Playground, FLUX) exhibit adultification bias against Black girls. For LLMs, the authors measure explicit bias through numeric trait ratings and implicit bias through decision-making scenarios involving guilt, school consequences, dating, and sexual activity. For T2I models, they generate images of Black and White girls at ages 10, 14, and 18 with differing trait descriptors and use human annotators to estimate perceived age and outfit revealingness. The paper reports significant adultification bias in several models, compares Black, White, Asian, and Latina girls to distinguish adultification from general majority-minority bias, and analyzes refusal patterns in image generation. The central claim is that state-of-the-art, widely deployed generative AI models reproduce a documented human bias, with implications for child-facing applications.
Significance. If validated, this is an important and timely contribution. The paper extends a well-established human bias to generative AI, spans text and image modalities, and addresses a group—Black girls—that is underrepresented in AI fairness research. Strengths include the use of multiple models, inclusion of Asian and Latina comparison groups, BH-corrected p-values, human annotation with reported inter-rater agreement, and a direct decision-task methodology for LLMs that observes model outputs rather than relying solely on self-report. The refusal-bias analysis is also a useful addition. However, load-bearing issues remain: the explicit trait composite includes 'innocent' without reverse-coding, the statistical test for implicit bias is underspecified, and the T2I measurement is vulnerable to annotator bias. These issues are fixable with re-analysis and additional control experiments, but they currently prevent full confidence in the paper's claims.
major comments (4)
- [Section 3.2, Appendix B] The explicit adultification composite is directionally flawed. The list of adultification traits includes 'innocent' (Appendix B, Table 4), yet higher ratings on 'innocent' indicate less adultification, not more. The paper averages raw ratings across all adultification traits without reverse-coding 'innocent', so the means and p-values in Table 4 conflate the construct. For example, Llama-3.1-70B's mean rating of 4.83 for Black girls on adultification-related traits is computed with 'innocent' scored in the wrong direction, potentially underestimating or overestimating the true difference. Please reverse-code 'innocent' and re-run the explicit bias analyses.
- [Section 3.3.1, Table 2] The one-sample t-test on bias(r1, r2) is underspecified. The bias formula aggregates counts over all prompts, yielding a single number per model and context. A t-test against zero requires multiple independent observations, but the unit of analysis—per prompt, per name, or per batch—is not described. Without knowing the number of observations and their sampling distribution, the significance stars in Table 2 and Figure 1 are not justified. Please clarify the statistical procedure, including how variability is estimated and how many independent units each test uses.
- [Section 4.1.2, Limitations] The T2I measurement relies on human annotators estimating the age and revealingness of generated images. Adultification bias is itself a documented human perceptual bias (Goff et al., 2014; Blake & Epstein, 2019), so the annotators' estimates could reflect their own bias rather than model output. The paper acknowledges this in the Limitations section but responds only by citing high inter-rater agreement (Table 6). High agreement is precisely what shared cultural bias would produce, so it does not establish validity. To attribute the Black-White differences (e.g., Playground age-18 difference of 3.47 years in Table 3) to the models, a calibration or ground-truth control is needed: annotators should also rate images with known age and clothing characteristics, and the analysis should adjust for annotator bias. Without such a control, the image-based half of the central claim is unverified.
- [Section 4.2, Section 4.3] Refusal filtering may bias the T2I comparisons. The main T2I analysis in Table 3 includes only successfully generated images, but Section 4.3 shows that refusal rates differ by race for 'seductive' prompts, even if not statistically significant due to small samples. If, for example, 'seductive 10 year old White girl' prompts are refused more often than the corresponding Black prompts, the remaining White images could be less sexualized by selection, inflating the measured difference. The main analysis does not account for this selection. Please report refusal counts per race, age, and trait for the main experiment and conduct a sensitivity analysis or selection model.
minor comments (5)
- [Section 3.2.1] The sentence '6 out of 7 models (all but GPT-4o) rate Black girls higher than White girls (p < 0.001)' should clarify whether the p-value is after Benjamini-Hochberg correction and whether it applies to the aggregate comparison or the individual model comparisons.
- [Section 3.3] The phrase 'we generally do not exhibit adultification bias against Asian girls' should read 'models generally do not exhibit...'
- [Appendix B, Table 4] The table would be easier to interpret if the 'Adult.' and 'Base.' labels were expanded to 'Adultification' and 'Baseline', and if the direction of each trait relative to adultification were noted in a footnote.
- [References] References [45] and [46] appear to be the same Kurdi et al. 2019 paper; please merge or distinguish them.
- [Section 4.2] There is a typo in 'outputs images pf Black girls'—should be 'of'.
Circularity Check
No significant circularity: the central claims rest on direct empirical measurement, with an explicitly acknowledged annotator-bias limitation that is a validity concern, not a derivation-chain reduction.
full rationale
The paper's derivation chain is empirical, not constructive. For LLMs, explicit bias is measured by direct numeric ratings of model outputs on traits taken from external adultification literature (Goff et al., Blake & Epstein), and implicit bias is measured by comparing the observed frequency of model-assigned consequences across racialized names. No parameters are fitted to data, and no prediction is derived from an input that already contains the conclusion. The bias metric, bias(r1, r2) = D(c_a, r1)/(D(c_a, r1)+D(c_n, r1)) - D(c_a, r2)/(D(c_a, r2)+D(c_n, r2)), is a defined contrast of observed counts, not a tautology. For T2I models, the paper quantifies adultification through human annotations of age and outfit revealingness. Section 5 explicitly states that 'the documented presence of adultification bias in humans [9, 29] may have impacted the reliability of our human image annotations.' This is an honest, load-bearing validity threat: because the measurement instrument (human perception) may share the very bias under study, the Black-White differences in perceived age and revealingness could be partly attributable to annotators rather than to the models. High inter-rater agreement does not resolve this, as shared bias increases agreement. However, this is a measurement-confounding concern about the correctness of the T2I conclusion, not a circular derivation: the conclusion does not follow by construction from the definitions or from the prompts, and the paper does not fit any parameter to the outcome it then 'predicts.' There are no load-bearing self-citations; the supporting references are external sociological and psychological studies. Therefore, the paper is not circular, though the T2I half carries a residual validity risk that should be weighed in interpretation.
Assumptions & free parameters
assumptions (4)
- domain assumption The adapted explicit bias rating scale is a valid indicator of adultification bias without trait-direction reverse-coding.
- domain assumption Name-based race cues in the profile generation task elicit the same stereotypes as explicit race labels.
- domain assumption The human annotation protocol measures perceived age and outfit revealingness without bias from annotator stereotypes.
- domain assumption The assigned consequences (suspension, STI test, kissed, true perpetrator) are appropriate operationalizations of adultification-related outcomes.
Cite this review
Pith. "Pith review of Adultification Bias in LLMs and Text-to-Image Models." pith.science (2026). https://pith.science/paper/PHUNHU55
@misc{pith2026250607282,
author = {Pith},
title = {Pith review of: Adultification Bias in LLMs and Text-to-Image Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/PHUNHU55}},
note = {Machine review of arXiv:2506.07282}
}
read the original abstract
The rapid adoption of generative AI models in domains such as education, policing, and social media raises significant concerns about potential bias and safety issues, particularly along protected attributes, such as race and gender, and when interacting with minors. Given the urgency of facilitating safe interactions with AI systems, we study bias along axes of race and gender in young girls. More specifically, we focus on "adultification bias," a phenomenon in which Black girls are presumed to be more defiant, sexually intimate, and culpable than their White peers. Advances in alignment techniques show promise towards mitigating biases but vary in their coverage and effectiveness across models and bias types. Therefore, we measure explicit and implicit adultification bias in widely used LLMs and text-to-image (T2I) models, such as OpenAI, Meta, and Stability AI models. We find that LLMs exhibit explicit and implicit adultification bias against Black girls, assigning them harsher, more sexualized consequences in comparison to their White peers. Additionally, we find that T2I models depict Black girls as older and wearing more revealing clothing than their White counterparts, illustrating how adultification bias persists across modalities. We make three key contributions: (1) we measure a new form of bias in generative AI models, (2) we systematically study adultification bias across modalities, and (3) our findings emphasize that current alignment methods are insufficient for comprehensively addressing bias. Therefore, new alignment methods that address biases such as adultification are needed to ensure safe and equitable AI deployment.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Khan Academy. 2025. Meet Khanmigo. https://khanmigo.ai/
2025
-
[2]
Alan Agresti. 1992. A Survey of Exact Inference for Contingency Tables. Statist. Sci. 7, 1 (Feb. 1992), 131–153. https://doi.org/10.1214/ss/1177011454
arXiv 1992
-
[3]
Pietro Astolfi, Marlene Careil, Melissa Hall, Oscar Mañas, Matthew Muck- ley, Jakob Verbeek, Adriana Romero Soriano, and Michal Drozdzal. 2024. Consistency-diversity-realism Pareto fronts of conditional image genera- tive models. arXiv:2406.10429 [cs.CV] https://arxiv.org/abs/2406.10429
arXiv 2024
-
[4]
Griffiths
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L. Griffiths
-
[5]
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jack- son Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Ka- mal Ndouss...
arXiv 2022
-
[6]
Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The Problem With Bias: Allocative Versus Representational Harms in Ma- chine Learning. In SIGCIS Conference. http://meetings.sigcis.org/uploads/ 6/3/6/8/6368912/program.pdf
2017
-
[7]
Venkatesh Babu, and Danish Pruthi
Abhipsa Basu, R. Venkatesh Babu, and Danish Pruthi. 2023. Inspecting the Geographical Representativeness of Images from Text-to-Image Models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, Paris, France, 5113–5124. https://doi.org/10.1109/ICCV51070.2023.00474
arXiv 2023
-
[9]
Jamila Blake and Rebecca Epstein. 2019. Listening to Black Women and Girls: Lived Experiences of Adultifcation Bias. https://www.law.georgetown.edu/poverty-inequality-center/wp- content/uploads/sites/14/2019/05/Listening-to-Black-Women-and- Girls.pdf
2019
Show all 97 references
-
[10]
Blake, Verna M
Jamilia J. Blake, Verna M. Keith, Wen Luo, Huy Le, and Phia Salter. 2017. The Role of Colorism in Explaining African American Females’ Suspension Risk. School Psychology Quarterly 32, 1 (March 2017), 118–130. https: //doi.org/10.1037/spq0000173 Epub 2016 Sep 29
2017 doi
-
[11]
Maarten Buyl, Alexander Rogiers, Sander Noels, Iris Dominguez-Catena, Edith Heiter, Raphael Romero, Iman Johary, Alexandru-Cristian Mara, Jefrey Lijffijt, and Tijl De Bie. 2024. Large Language Models Reflect the Ideology of their Creators. arXiv:2410.18417 [cs.CL] https://arxi...
2024 arXiv
-
[12]
I’m a Strong Indepen- dent Black Woman
Stephanie Castelin and Grace White. 2022. “I’m a Strong Indepen- dent Black Woman”: The Strong Black Woman Schema and Mental Health in College-Aged Black Women. Psychology of Women Quar- terly 46, 2 (2022), 196–208. https://doi.org/10.1177/03616843211067501 arXiv:https://doi.o...
2022 doi
-
[13]
Amazon Press Center. 2024. Silicon Valley High School Launches Patented Suite of AI-Powered Tools on Learning Plat- form; Transforms Education with Technology Partner AWS. https://press.aboutamazon.com/aws/2024/11/silicon-valley-high- school-launches-patented-suite-of-ai-power...
2024
-
[14]
Marc Cheong, Ehsan Abedin, Marinus Ferreira, Ritsaart Reimann, Shalom Chalson, Pamela Robinson, Joanne Byrne, Leah Ruppanner, Mark Alfano, and Colin Klein. 2024. Investigating Gender and Racial Biases in DALL- E Mini Images. ACM J. Responsib. Comput. 1, 2, Article 13 (June 202...
2024 doi
-
[15]
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models. arXiv:2202.04053 [cs.CV] https://arxiv.org/abs/2202.04053
2023 arXiv
-
[16]
Alex Dow, Jean Garcia-Gathright, Nicholas Pan- gakis, Emily Sheng, Dan Vann, Matthew Vogel, and Hanna Wallach
Emily Corvi, Hannah Washington, Stefanie Reed, Chad Atalla, Alexan- dra Chouldechova, P. Alex Dow, Jean Garcia-Gathright, Nicholas Pan- gakis, Emily Sheng, Dan Vann, Matthew Vogel, and Hanna Wallach
-
[17]
Ryan Deffenbaugh. 2024. Meta’s AI Chatbot Nears 600 Million Monthly Users, Zuckerberg Says. https://www.yahoo.com/news/m/0d37fe67-cc27- 3273-a12c-0ac18cfc19df/meta-s-ai-chatbot-nears-600.html
2024
-
[18]
arXiv:2504.00928 [cs.CL] https://arxiv.org/abs/2504.00928
Taxonomizing Representational Harms using Speech Act Theory. arXiv:2504.00928 [cs.CL] https://arxiv.org/abs/2504.00928
-
[19]
Harnoor Dhingra, Preetiha Jayashanker, Sayali Moghe, and Emma Strubell
-
[20]
Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruk- sachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Tr...
2021
-
[21]
J. F. Dovidio, K. Kawakami, N. Smoak, and S. L. Gaertner. 2008. The nature of contemporary racial prejudice: Insight from implicit and explicit measures of attitudes. In Attitudes: Insights from the new implicit measures , R. E. Petty, R. H. Fazio, and P. Briñol (Eds.). Psycho...
2008
-
[22]
Education Week Research Center. 2017. Policing America’s Schools: An Education Week Analysis. https://www.edweek.org/which-students-are- arrested-most-in-school-u-s-data-by-school#/overview. Data visualiza- tion and analysis of school-based arrests and referrals by race and state
2017
-
[23]
John F Dovidio, Kerry Kawakami, and Samuel L Gaertner. 2002. Implicit and explicit prejudice and interracial interaction. Journal of Personality and Social Psychology 82, 1 (2002), 62–68. https://doi.org/10.1037//0022- 3514.82.1.62
2002 doi
-
[24]
R. H. Fazio, J. R. Jackson, B. C. Dunton, and C. J. Williams. 1995. Variability in automatic activation as an unobtrusive measure of racial attitudes: A bona fide pipeline? Journal of Personality and Social Psychology 69, 6 (1995), 1013–1027. https://doi.org/10.1037/0022-3514....
1995 doi
-
[25]
French and Helen A
Bryana H. French and Helen A. Neville. 2013. Sexual Co- ercion Among Black and White Teenagers: Sexual Stereotypes and Psychobehavioral Correlates. The Counseling Psychologist 41, 8 (2013), 1186–1212. https://doi.org/10.1177/0011000012461379 arXiv:https://doi.org/10.1177/00110...
2013 doi
-
[26]
EU Artificial Intelligence Act. 2024. Annex 3: Risk-Based Classification of AI Systems. https://artificialintelligenceact.eu/annex/3/ Accessed: 2024-12-13
2024
-
[27]
Galvan and B
Manuel J. Galvan and B. Keith Payne. 2024. Implicit Bias as a Cognitive Manifestation of Systemic Racism. Daedalus 153, 1 (March 2024), 106–122. https://doi.org/10.1162/daed_a_02051
2024 doi
-
[28]
PA Goff, JL Eberhardt, MJ Williams, and MC Jackson. 2008. Not yet human: implicit knowledge, historical dehumanization, and contemporary consequences. J Pers Soc Psychol 94, 2 (Feb 2008), 292–306. https://doi. org/10.1037/0022-3514.94.2.292
2008 doi
-
[29]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. arXiv:2309.00770 [cs.CL] https://arxiv.org/abs/2309.00770
2024 arXiv
-
[30]
Andrew Gounardes. 2023. Senate Bill S7694A: Establishes the Stop Addic- tive Feeds Exploitation (SAFE) for Kids Act prohibiting the provision of addictive feeds to minors. New York State Senate, 2023-2024 Legislative Ses- sion. https://www.nysenate.gov/legislation/bills/2023/S...
2023
-
[31]
Goyal, Katie L
Monika K. Goyal, Katie L. Hayes, and Cynthia J. Mollen. 2012. Racial disparities in testing for sexually transmitted infections in the emergency department. Acad Emerg Med 19, 5 (2012), 604–607. https://doi.org/10. 1111/j.1553-2712.2012.01338.x Free article
2012
-
[32]
Phillip Atiba Goff, Matthew Christian Jackson, Brooke Allison Lewis Di Leone, Carmen Marie Culotta, and Natalie Ann DiTomasso. 2014. The essence of innocence: Consequences of dehumanizing Black children. Journal of Personality and Social Psychology 106, 4 (April 2014), 526–545...
2014 doi
-
[33]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, An- thony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...
2025 arXiv
-
[34]
A. G. Greenwald and M. R. Banaji. 1995. Implicit social cognition: Attitudes, self-esteem, and stereotypes. Psychological Review 102, 1 (1995), 4–27. https://doi.org/10.1037/0033-295X.102.1.4
1995 doi
-
[35]
Nico Grant. 2024. Google Chatbot’s A.I. Images Put People of Color in Nazi- Era Uniforms. https://www.nytimes.com/2024/02/22/technology/google- gemini-german-uniforms.html
2024
-
[36]
Melissa Hall, Candace Ross, Adina Williams, Nicolas Carion, Michal Drozdzal, and Adriana Romero Soriano. 2024. DIG In: Evaluating Dis- parities in Image Generations with Indicators for Geographic Diversity. arXiv:2308.06198 [cs.CV] https://arxiv.org/abs/2308.06198
2024 arXiv
-
[37]
Susan Hao, Renee Shelby, Yuchi Liu, Hansa Srinivasan, Mukul Bhutani, Burcu Karagol Ayan, Ryan Poplin, Shivani Poddar, and Sarah Laszlo. 2024. Harm Amplification in Text-to-Image Models. arXiv:2402.01787 [cs.CY] https://arxiv.org/abs/2402.01787
2024 arXiv
-
[38]
Bell, Candace Ross, Adina Williams, Michal Drozdzal, and Adriana Romero Soriano
Melissa Hall, Samuel J. Bell, Candace Ross, Adina Williams, Michal Drozdzal, and Adriana Romero Soriano. 2024. Towards Geographic Inclu- sion in the Evaluation of Text-to-Image Models. arXiv:2405.04457 [cs.CV] https://arxiv.org/abs/2405.04457
2024 arXiv
-
[39]
Edward Herbert. 2023. https://www.childrenssociety.org.uk/what-we- do/blogs/artificial-intelligence-body-image-and-toxic-expectations
2023
-
[40]
Valentin Hofmann, Pratyusha Ria Kalluri, Dan Jurafsky, and Sharese King
-
[41]
Deborah Hellman. 2023. Big Data and Compounding Injustice. Journal of Moral Philosophy (2023), 1–22. https://doi.org/10.2139/ssrn.3840175 Virginia Public Law and Legal Theory Research Paper No. 2021-27
2023 doi
-
[42]
Ariba Khan, Stephen Casper, and Dylan Hadfield-Menell. 2025. Random- ness, Not Representation: The Unreliability of Evaluating Cultural Align- ment in LLMs. arXiv:2503.08688 [cs.CY] https://arxiv.org/abs/2503.08688
2025 arXiv
-
[43]
Jon Kleinberg, Sendhil Mullainathan, and Manish Raghavan. 2016. Inherent Trade-Offs in the Fair Determination of Risk Scores. arXiv:1609.05807 [cs.LG] https://arxiv.org/abs/1609.05807
2016 arXiv
-
[44]
Klaus Krippendorff. 2013. Computing Krippendorff’s Alpha- Reliability. https://www.asc.upenn.edu/sites/default/files/2021- 03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
2013
-
[45]
Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. 2023. Taxonomizing and measuring representational harms: a look at image tagging. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence ...
2023
-
[46]
Seitchik, Jordan R
Benedek Kurdi, Allison E. Seitchik, Jordan R. Axt, Timothy J. Carroll, Arpi Karapetyan, Neela Kaushik, Diana Tomezsko, Anthony G. Greenwald, and Mahzarin R. Banaji. 2019. Relationship between the Implicit Association Test and intergroup behavior: A meta-analysis. American Psyc...
2019 doi
-
[47]
Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2023. Stable Bias: Analyzing Societal Representations in Diffusion Models. arXiv:2303.11408 [cs.CY] https://arxiv.org/abs/2303. 11408
2023 arXiv
-
[48]
Vera López and Meda Chesney-Lind. 2014. Latina girls speak out: Stereo- types, gender and relationship dynamics. Latino Studies 12, 4 (Dec. 2014), 527–549. https://doi.org/10.1057/lst.2014.54
2014 doi
-
[50]
Meta AI. 2024. A Social ’Study Buddy’ Gets a Conversational Lift from Meta Llama. https://ai.meta.com/blog/foondamate-study-aid-education-llama/ 6-minute read. FoondaMate, a fast-growing study aid, uses Meta Llama to assist students in emerging markets via WhatsApp and Messeng...
2024
-
[51]
Brent Mittelstadt, Sandra Wachter, and Chris Russell. 2023. The Unfairness of Fair Machine Learning: Levelling down and strict egalitarianism by default. Michigan Technology Law Review 4331652 (Jan. 2023). https: //papers.ssrn.com/abstract=4331652
2023
-
[52]
Ranjita Naik and Besmira Nushi. 2023. Social Biases through the Text-to- Image Generation Lens. arXiv:2304.06034 [cs.CY] https://arxiv.org/abs/ 2304.06034
2023 arXiv
-
[53]
Meta. 2024. Meet Your New Assistant: Meta AI, Built With Llama 3. https:// about.fb.com/news/2024/04/meta-ai-assistant-built-with-llama-3/. https: //about.fb.com/news/2024/04/meta-ai-assistant-built-with-llama-3/
2024
-
[54]
Department of Education
U.S. Department of Education. 2016. A First Look: Key Data Highlights on Equity and Opportunity Gaps in Our Nation’s Public Schools. U.S. Department of Education, Office for Civil Rights. https://civilrightsdata. ed.gov/assets/downloads/2013-14-first-look.pdf Revised on Octobe...
2016
-
[55]
Official Journal of the European Union. [n. d.]. Digital Services Act. https: //eur-lex.europa.eu/eli/reg/2022/2065/oj
2022
-
[56]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[57]
OECD. [n. d.]. Recommendation of the Council on Children in the Digital Environment. https://legalinstruments.oecd.org/en/instruments/OECD- LEGAL-0389
-
[58]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022
-
[59]
Devah Pager and Hana Shepherd. 2008. The Sociology of Discrimination: Racial Discrimination in Employment, Housing, Credit, and Consumer Markets. Annual Review of Sociology 34 (2008), 181–209. https://doi.org/ 10.1146/annurev.soc.33.040406.131740
2008
-
[60]
Why don’t I look like her?
Alana Papageorgiou, Colleen Fisher, and Donna Cross. 2022. “Why don’t I look like her?” How adolescent girls view social media and its connection to body image. BMC Women’s Health 22, 1 (June 2022), 261. https://doi. org/10.1186/s12905-022-01845-4
2022 doi
-
[61]
Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, and Shin’ichi Satoh. 2023. Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition ...
2023
-
[62]
Sarah Perez. 2025. ChatGPT Doubled Its Weekly Active Users in Under 6 Months, Thanks to New Releases. https: //techcrunch.com/2025/03/06/chatgpt-doubled-its-weekly-active- users-in-under-6-months-thanks-to-new-releases/
2025
-
[63]
Prolific. 2024. Easily find vetted research participants and AI taskers at scale. https://www.prolific.com/
2024
-
[64]
M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, R. Comanescu, C. Akbulut, T. Stepleton, J. Mateos-Garcia, S. Bergman, J. Kay, C. Grif- fin, B. Bariach, I. Gabriel, V. Rieser, W. Isaac, and L. Weidinger. 2024. Gaps in the Safety Evaluation of Generative AI. Proceedings of the...
2024 doi
-
[65]
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. BBQ: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022 , Smaranda Mures...
2022
-
[66]
Philippe Remy. 2021. Name Dataset. https://github.com/philipperemy/ name-dataset
2021
-
[67]
Katherine-Marie Robinson. 2024. Tales from the Wild West: Craft- ing Scenarios to Audit Bias in LLMs. In HEAL@CHI’24. https://api. semanticscholar.org/CorpusID:271545975
2024
-
[68]
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Han- nah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. In Proceedings of the 62nd Annual ...
2024
-
[69]
Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Ramona Comanescu, Canfer Akbulut, Tom Stepleton, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, William Isaac, and Laura Weidinger. 2024. Gaps in the...
2024
-
[70]
Pushkar Shukla, Aditya Chinchure, Emily Diana, Alexander Tolbert, Kartik Hosanagar, Vineeth N Balasubramanian, Leonid Sigal, and Matthew Turk
-
[71]
Zara Siddique, Liam Turner, and Luis Espinosa-Anke. 2024. Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Nat- ural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and...
2024 doi
-
[72]
I’m sorry to hear that
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. "I’m sorry to hear that": Finding New Biases in Lan- guage Models with a Holistic Descriptor Dataset. arXiv:2205.09209 [cs.CL] https://arxiv.org/abs/2205.09209
2022 arXiv
-
[73]
Joshua Rovner. 2023. Black Disparities in Youth Incarceration. https://www.sentencingproject.org/fact-sheet/black-disparities-in- youth-incarceration/
2023
-
[74]
Harini Suresh and John Guttag. 2021. A Framework for Understand- ing Sources of Harm throughout the Machine Learning Life Cycle. In Proceedings of the 1st ACM Conference on Equity and Access in Algo- rithms, Mechanisms, and Optimization (–, NY, USA) (EAAMO ’21). Associa- tion ...
2021
-
[75]
arXiv:2505.17280 [cs.CV] https://arxiv.org/abs/2505
Mitigate One, Skew Another? Tackling Intersectional Biases in Text- to-Image Models. arXiv:2505.17280 [cs.CV] https://arxiv.org/abs/2505. 17280
-
[76]
Feder Cooper, An- gelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P
Hanna Wallach, Meera Desai, Nicholas Pangakis, A. Feder Cooper, An- gelina Wang, Solon Barocas, Alexandra Chouldechova, Chad Atalla, Su Lin Blodgett, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexan- dra Olteanu, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Va...
-
[77]
Yixin Wan and Kai-Wei Chang. 2024. The Male CEO and the Female Assistant: Evaluation and Mitigation of Gender Biases in Text-To-Image Generation of Dual Subjects. arXiv:2402.11089 [cs.CV] https://arxiv.org/ abs/2402.11089
2024 arXiv
-
[78]
Jay Stanley. 2024. Police Departments Shouldn’t Allow Officers to Use AI to Draft Police Reports. https://www.aclu.org/documents/aclu-on-police- departments-use-of-ai-to-draft-police-reports Senior Policy Analyst with the ACLU’s Speech, Privacy, and Technology Project
2024
-
[79]
Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. 2024. Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation. arXiv:2404.01030 [cs.CV] https://arxiv.org...
2024 arXiv
-
[80]
Tate, Jacob Steiss, Drew Bailey, Steve Graham, Youngsun Moon, Daniel Ritchie, Waverly Tseng, and Mark Warschauer
Tamara P. Tate, Jacob Steiss, Drew Bailey, Steve Graham, Youngsun Moon, Daniel Ritchie, Waverly Tseng, and Mark Warschauer. 2024. Can AI provide useful holistic essay scoring? Computers and Education: Artificial Intelli- gence 7 (Dec. 2024), 100255. https://doi.org/10.1016/j.c...
2024
-
[81]
Dickerson
Angelina Wang, Jamie Morgenstern, and John P. Dickerson. 2025. Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence 7, 3 (March 2025), 400–411. https://doi.org/10.1038/s42256-025-00986-z
2025 doi
-
[82]
arXiv:2411.10939 [cs.CY] https://arxiv.org/abs/2411.10939
Evaluating Generative AI Systems is a Social Science Measurement Challenge. arXiv:2411.10939 [cs.CY] https://arxiv.org/abs/2411.10939
-
[83]
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023. Jailbroken: how does LLM safety training fail?. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA,...
2023
-
[84]
Kelly is a Warm Person, Joseph is a Role Model
Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor,...
2023 doi
-
[85]
Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, and William Isaac. 2025. Toward an Evalua- tion Science for Generative AI Systems. arXiv:2503.05336 [cs.AI] https: //arxiv.o...
2025 arXiv
-
[86]
Yixin Wan, Di Wu, Haoran Wang, and Kai-Wei Chang. 2024. The Fac- tuality Tax of Diversity-Intervened Text-to-Image Generation: Bench- mark and Fact-Augmented Intervention. In Proceedings of the 2024 Con- ference on Empirical Methods in Natural Language Processing . Associ- ati...
2024 doi
-
[87]
Kyra Wilson and Aylin Caliskan. 2024. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society 7, 11 (Oct. 2024), 1578–1590. https://doi.org/10.1609/aies.v7i1.31748
2024 doi
-
[88]
Jialu Wang, Xinyue Gabby Liu, Zonglin Di, Yang Liu, and Xin Eric Wang
-
[89]
arXiv:2306.00905 [cs.CL] https://arxiv.org/abs/2306.00905
T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation. arXiv:2306.00905 [cs.CL] https://arxiv.org/abs/2306.00905
-
[90]
Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, Ruoxi Jia, Bo Li, Kai Li, Danqi Chen, Peter Henderson, and Prateek Mittal
-
[91]
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. 2024. Assessing the brittleness of safety alignment via pruning and low-rank modifications. In Proceedings of the 41st International Conference on ...
2024
-
[92]
What is adultification bias?
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InThirty-seventh Conference on Neural ...
2023
-
[93]
Carolyn M. West. 2009. Still on the auction block: The (s)exploitation of Black adolescent in rap(e) music and hip hop culture. In The Sexualization of Childhood, S. Olfman (Ed.). Praeger Press, Westport, CT, 89–102. https: //doi.org/10.5040/9798216013778.ch-007 Chapter 7
2009 doi
-
[96]
Robert Wolfe, Yiwei Yang, Bill Howe, and Aylin Caliskan. 2023. Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias. In Proceedings of the 2023 ACM Con- ference on Fairness, Accountability, and Transparency (Chicag...
2023
-
[98]
https://openreview.net/forum?id=YfKNaRktan
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal. https://openreview.net/forum?id=YfKNaRktan
-
[99]
Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De la Torre. 2023. ITI- GEN: Inclusive Text-to-Image Generation. In 2023 IEEE/CVF In- ternational Conference on Computer Vision (ICCV) . 3969–3980. https://openaccess.thecvf.com/conte...
2023
-
[2023]
arXiv:2307.00101 [cs.CL] https: //arxiv.org/abs/2307.00101
Queer People are People First: Deconstructing Sexual Identity Stereotypes in Large Language Models. arXiv:2307.00101 [cs.CL] https: //arxiv.org/abs/2307.00101
-
[2024]
Nature 633, 8028 (Sept
AI generates covertly racist decisions about people based on their dialect. Nature 633, 8028 (Sept. 2024), 147–154. https://doi.org/10.1038/ s41586-024-07856-5
2024
-
[2025]
Proceedings of the National Academy of Sciences 122, 8 (Feb
Explicitly unbiased large language models still form biased associa- tions. Proceedings of the National Academy of Sciences 122, 8 (Feb. 2025), e2416228122. https://doi.org/10.1073/pnas.2416228122
2025 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.