{"id":"3c4fab9c-d396-4f1d-ab2d-46c483948410","arxiv_id":"2506.07282","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.","lead":"This study probes whether AI chatbots and image generators treat Black girls as older, guiltier, and more sexualized than White girls, a bias called adultification. Across chat and image models from OpenAI, Meta, and others, the authors find uneven but real signs of this bias, especially in harsher school or dating scenarios and in portraits of girls.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The T2I findings could be an artifact of human annotators' own adultification bias; without a ground-truth calibration study, the image-based half of the central claim is unverified.","rationale":"The reader's weakest_assumption identifies the same concern: human annotations of age and revealingness may themselves be driven by adultification bias, conflating model bias with measurement bias. My stress-test concurs and elevates it to the most load-bearing threat to the paper's central claim. The central claim has two independent halves: LLM decision-making and T2I image generation. The LLM half is supported by direct observations of model outputs (consequence assignments), which are unambiguous choices recorded verbatim; the reported t-test issue concerns the statistical inference, but the raw effect sizes are large and many are consistent across models. The T2I half, however, relies on subjective human judgments of the very construct under study. The inter-rater agreement of 0.7-0.72 does not rule out shared bias; it likely reflects it, since adultification bias is culturally widespread. The paper's own Limitations section concedes this possibility, and no calibration against ground-truth stimuli is provided. Therefore, until a control experiment demonstrates that annotators do not systematically over-age or over-sexualize Black girls in images with known properties, the image-based conclusion 'T2I models generate images portraying Black girls as older and wearing more revealing clothing than their White peers' is not established. This is a concrete, testable concern, and the proposed calibration study would settle it. If the calibration fails, the verdict should move to CONDITIONAL with the T2I claim explicitly unverified; if it passes, the central claim stands stronger. I therefore recommend UNCHANGED (i.e., continue with CONDITIONAL) pending such a test, matching the reader's conditional verdict. No ad hominem is intended; this is a purely methodological critique of the measurement pipeline.","tokens_in":29425,"tokens_out":4593,"duration_ms":58543,"concrete_test":"Run a calibration study with the same (or an equivalent) set of 192 US-based Prolific annotators, using real photographs of Black and White girls at ages 10, 14, and 18, matched pairwise on actual age and on clothing style/coverage (e.g., sourced from a dataset with verified metadata such as yearbook-style portraits or stock photography with documented ages and outfit descriptions). Ask annotators to estimate age and rate revealingness using the identical survey questions from Appendix D. Compute the mean Black-minus-White difference in estimated age and revealingness on this control set. If the control set shows a significant positive difference (e.g., Black girls estimated >1 year older or >0.2 points more revealing despite identical ground truth), then the human measurement instrument is biased and the T2I findings in §4.2 are confounded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's text-to-image claim rests entirely on human annotators estimating the age and revealingness of generated images. Adultification bias is a documented human perceptual phenomenon (Goff et al., 2014; Blake & Epstein, 2019): humans perceive Black girls as older and more sexualized than White girls of the same age. The measurement instrument therefore shares the very property being measured. The paper acknowledges this in its Limitations section: 'the documented presence of adultification bias in humans may have impacted the reliability of our human image annotations.' But the response—citing high inter-rater agreement (all > 0.7, Table 6)—does not address the concern, because shared cultural bias produces high agreement. Consensus among annotators is evidence of shared perception, not of validity. Without a control condition using stimuli with known ground truth (e.g., real photos of Black and White girls of identical age and clothing), the reported effects—Black girls estimated up to 3.47 years older (Playground, age 18) and outfits up to 0.77 points more revealing—cannot be attributed to the models. This is the single most load-bearing concern because it threatens the entire image-based half of the central claim: if the annotators are biased, the model 'bias' in §4.2 could be a pure measurement artifact, independent of what the image models actually depict. The LLM decision task, by contrast, records model-generated choices directly, so even if the t-test is mis-specified, the raw directional effects (e.g., GPT-4o bias = 0.20*** for sexual activity) are direct observations of model output. Thus the weakest load-bearing point is the unvalidated human evaluation, not the LLM statistics.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether large language models (Llama, GPT) and text-to-image models (Meta T2I, Stable Diffusion, Playground, FLUX) exhibit adultification bias against Black girls. For LLMs, the authors measure explicit bias through numeric trait ratings and implicit bias through decision-making scenarios involving guilt, school consequences, dating, and sexual activity. For T2I models, they generate images of Black and White girls at ages 10, 14, and 18 with differing trait descriptors and use human annotators to estimate perceived age and outfit revealingness. The paper reports significant adultification bias in several models, compares Black, White, Asian, and Latina girls to distinguish adultification from general majority-minority bias, and analyzes refusal patterns in image generation. The central claim is that state-of-the-art, widely deployed generative AI models reproduce a documented human bias, with implications for child-facing applications.","tokens_in":29640,"tokens_out":6354,"duration_ms":69297,"significance":"If validated, this is an important and timely contribution. The paper extends a well-established human bias to generative AI, spans text and image modalities, and addresses a group—Black girls—that is underrepresented in AI fairness research. Strengths include the use of multiple models, inclusion of Asian and Latina comparison groups, BH-corrected p-values, human annotation with reported inter-rater agreement, and a direct decision-task methodology for LLMs that observes model outputs rather than relying solely on self-report. The refusal-bias analysis is also a useful addition. However, load-bearing issues remain: the explicit trait composite includes 'innocent' without reverse-coding, the statistical test for implicit bias is underspecified, and the T2I measurement is vulnerable to annotator bias. These issues are fixable with re-analysis and additional control experiments, but they currently prevent full confidence in the paper's claims.","major_comments":[{"comment":"The explicit adultification composite is directionally flawed. The list of adultification traits includes 'innocent' (Appendix B, Table 4), yet higher ratings on 'innocent' indicate less adultification, not more. The paper averages raw ratings across all adultification traits without reverse-coding 'innocent', so the means and p-values in Table 4 conflate the construct. For example, Llama-3.1-70B's mean rating of 4.83 for Black girls on adultification-related traits is computed with 'innocent' scored in the wrong direction, potentially underestimating or overestimating the true difference. Please reverse-code 'innocent' and re-run the explicit bias analyses.","section":"Section 3.2, Appendix B"},{"comment":"The one-sample t-test on bias(r1, r2) is underspecified. The bias formula aggregates counts over all prompts, yielding a single number per model and context. A t-test against zero requires multiple independent observations, but the unit of analysis—per prompt, per name, or per batch—is not described. Without knowing the number of observations and their sampling distribution, the significance stars in Table 2 and Figure 1 are not justified. Please clarify the statistical procedure, including how variability is estimated and how many independent units each test uses.","section":"Section 3.3.1, Table 2"},{"comment":"The T2I measurement relies on human annotators estimating the age and revealingness of generated images. Adultification bias is itself a documented human perceptual bias (Goff et al., 2014; Blake & Epstein, 2019), so the annotators' estimates could reflect their own bias rather than model output. The paper acknowledges this in the Limitations section but responds only by citing high inter-rater agreement (Table 6). High agreement is precisely what shared cultural bias would produce, so it does not establish validity. To attribute the Black-White differences (e.g., Playground age-18 difference of 3.47 years in Table 3) to the models, a calibration or ground-truth control is needed: annotators should also rate images with known age and clothing characteristics, and the analysis should adjust for annotator bias. Without such a control, the image-based half of the central claim is unverified.","section":"Section 4.1.2, Limitations"},{"comment":"Refusal filtering may bias the T2I comparisons. The main T2I analysis in Table 3 includes only successfully generated images, but Section 4.3 shows that refusal rates differ by race for 'seductive' prompts, even if not statistically significant due to small samples. If, for example, 'seductive 10 year old White girl' prompts are refused more often than the corresponding Black prompts, the remaining White images could be less sexualized by selection, inflating the measured difference. The main analysis does not account for this selection. Please report refusal counts per race, age, and trait for the main experiment and conduct a sensitivity analysis or selection model.","section":"Section 4.2, Section 4.3"}],"minor_comments":[{"comment":"The sentence '6 out of 7 models (all but GPT-4o) rate Black girls higher than White girls (p < 0.001)' should clarify whether the p-value is after Benjamini-Hochberg correction and whether it applies to the aggregate comparison or the individual model comparisons.","section":"Section 3.2.1"},{"comment":"The phrase 'we generally do not exhibit adultification bias against Asian girls' should read 'models generally do not exhibit...'","section":"Section 3.3"},{"comment":"The table would be easier to interpret if the 'Adult.' and 'Base.' labels were expanded to 'Adultification' and 'Baseline', and if the direction of each trait relative to adultification were noted in a footnote.","section":"Appendix B, Table 4"},{"comment":"References [45] and [46] appear to be the same Kurdi et al. 2019 paper; please merge or distinguish them.","section":"References"},{"comment":"There is a typo in 'outputs images pf Black girls'—should be 'of'.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The T2I annotator-bias concern is the most serious issue and may require additional data collection. If the authors can provide a calibration study showing that annotator bias does not drive the reported effects, I would be willing to reconsider. The explicit composite and t-test specification issues are re-analysis and clarification; they should be addressed regardless. The paper is well within scope for the journal and the topic is important."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first paper I've seen that measures adultification bias in generative models, and it covers both LLMs and text-to-image models. The LLM evidence is fairly solid; the image evidence is real but rests on human annotations that may share the bias being measured.\n\nWhat's new and good: the design is thoughtful. Comparing Black vs. White, Asian vs. White, and Latina vs. White separates adultification from generic minority bias, which is the right control. The study uses multiple decision-making scenarios, multiple model families, BH-corrected p-values, and honestly reports model heterogeneity. The refusal-bias analysis is a nice extra. The strongest result is the LLM decision task: GPT-4o assigns the STI-test consequence to Black girls 20 percentage points more often than to White girls. That is a direct observation of model output, not an inference.\n\nSoft spots: the T2I half is the weakest link, and the paper itself admits it in the Limitations section, where it says adultification bias in humans \"may have impacted the reliability of our human image annotations.\" That is the right concern. Annotators estimate age and revealingness from images of Black and White girls. If those annotators share adultification bias, the measured difference could be a measurement artifact rather than a model property. High inter-rater agreement does not fix this; shared bias produces agreement. Without a calibration study using ground-truth stimuli, the image-based claim is not yet established.\n\nSecond, the implicit-bias t-test is under-specified. The paper computes one aggregate bias score per model/task/comparison and runs a one-sample t-test against zero. A single score has no clear sampling distribution. A chi-square or Fisher's exact test on the underlying counts would be more defensible. The raw effect sizes are large enough that the conclusion will likely survive re-analysis, but the current statistics are not airtight.\n\nThird, the explicit trait composite includes 'innocent' without reverse-coding. That makes the test conservative rather than inflated, so it is a minor issue, but it should be fixed.\n\nFourth, no data or code are released. For a paper that depends on human annotations, that is a real reproducibility gap.\n\nBottom line: the LLM finding is likely real; the T2I finding is plausible but unverified. This deserves peer review. With data release and a calibration study, it could be a solid contribution. I'd bring it to a reading group.","headline":"First credible measurement of adultification bias in both LLM and text-to-image models; the LLM half is solid, the image half rests on unvalidated human annotations.","tokens_in":30248,"tokens_out":3261,"would_cite":true,"duration_ms":37394,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"State-of-the-art LLMs and text-to-image models exhibit adultification bias against Black girls, assigning them harsher and more sexualized outcomes in decision scenarios and depicting them as older and more revealingly dressed than White…","keywords":["adultification bias","large language models","text-to-image models","racial bias","intersectional bias","generative AI fairness","model alignment","human evaluation"],"falsifier":"Have the same annotators first estimate the ages of real photographs of Black and White girls whose true ages are known. If annotators systematically overestimate Black girls' ages even on real photos, adultification bias is present in the measurement itself; re-running the image comparison with annotator-bias calibration, or with automated age and clothing-coverage estimators, would then settle whether the models' Black-White output gap remains.","tokens_in":29179,"feed_emoji":"⚖️","tokens_out":8948,"duration_ms":88242,"temperature":0.7,"pith_summary":"Adultification bias, the documented tendency to perceive Black girls as older, more defiant, more sexually active, and more culpable than their White peers, has been established in human psychology and sociology. This paper asks whether the same bias has been absorbed by generative AI, and it argues the answer is yes: across seven chat models and four image generators, most LLMs rate Black girls higher on adultification-related traits, the larger models assign them harsher and more sexualized consequences in school, dating, and medical scenarios, and text-to-image models depict Black girls as older and more revealingly dressed than White girls. The point of establishing this is practical, because these models are already deployed in education, policing, and social media, where biased judgments about minors can translate into real decisions. The paper also contributes a measurement template for transferring a psychological bias construct into multimodal model audits.","feed_headline":"Black girls get harsher, more sexualized AI judgments","feed_subtitle":"Widely used chatbots and image generators show it even after safety alignment.","key_machinery":"The argument rests on three measurement instruments adapted from the psychology of adultification. The first is explicit-bias trait ratings, where models answer questions of the form 'How {trait} are {race} girls?' on a 1-to-5 scale for adultification traits (defiant, mature, intimate, and others) alongside baseline traits (sweet, kind, gentle). The second is implicit-bias decision scenarios, where models generate hypothetical profiles for two race-coded names such as Erin versus Latasha and then assign one of two consequences, with the bias score defined as $bias(r_1, r_2) = \\frac{D(c_a, r_1)}{D(c_a, r_1)+D(c_n, r_1)} - \\frac{D(c_a, r_2)}{D(c_a, r_2)+D(c_n, r_2)}$, where $c_a$ is the adultification-associated consequence (suspension, true perpetrator, kissed, STI test) and $c_n$ its opposite; the pairwise-race design (Black, Asian, and Latina girls versus White girls) is what separates adultification bias from a generic majority-minority bias. The third is human annotation of generated images, where 192 US-based annotators estimate the age of each subject and rate outfit revealingness on Likert scales for prompts of the form 'Imagine a {trait} {age} {race} girl wearing a dress,' and the Black-minus-White differences in those judgments form the image-side outcome. The comparison structure, findings that appear against Black girls but not against Asian girls and that grow with model scale, is what carries the conclusion that the bias is adultification specifically.","core_discovery":"The paper's central claim is that state-of-the-art, publicly accessible LLMs and text-to-image models exhibit significant adultification bias against Black girls. In LLMs, six of seven models give Black girls higher numeric ratings than White girls on traits such as defiant, mature, and intimate, and in profile-based decision tasks the largest models more frequently assign Black girls suspension over detention, presumed kissing after a date, and an STI test, with GPT-4o and Llama-3.1-70B showing the strongest effects. In text-to-image models, human annotators judged generated Black girls as up to 2.2 years older in Meta's model and 3.47 years older in Playground relative to White girls for identical prompts, with more revealing outfits; StableDiffusion and FLUX did not show the same Black-White gap but produced more revealing images overall. The authors read the contrast between White-versus-Black and White-versus-Asian comparisons as evidence that the effect is adultification bias specifically, not a general majority-minority bias, and conclude that alignment methods such as RLHF and safety refusals fail to cover this form of bias.","pith_inferences":["If the observed scaling holds, future larger model releases may exhibit more, not less, adultification bias; this is a testable prediction to check with each new model version against the same scenario battery.","The profile-generation template could be pointed at other intersectionally adultified groups, such as Black boys or disabled minors, to map how far the bias generalizes beyond Black girls and whether it is specific to the girl-adultification construct.","A cheaper and reproducible proxy for the image finding would be automated age estimation and clothing-coverage segmentation; if such automated measures reproduce the Black-White gap without human judges, the annotation-bias confound would be ruled out.","The explicit-versus-implicit age phrasing gap suggests a concrete policy lever: platforms could be required to report refusal rates per demographic group and per phrasing class, which would surface differential enforcement of safety filters."],"forward_implications":["Any deployed chatbot assisting with school discipline, policing, or healthcare triage may give Black girls harsher and more sexualized judgments than White girls for identical behavior, because the decision bias survives alignment in the largest models tested.","Text-to-image platforms used in social media and marketing will keep producing images that age up and sexualize Black girls even when the prompt specifies the same age as for White girls, with Meta's and Playground's models showing the widest gaps.","Safety refusals are not a reliable guardrail: rephrasing 'seductive 14 year old girl' as 'seductive high school girl' cut refusal rates by 91% for Black girls and 74% for White girls in the Meta model, so prompt phrasing changes both absolute safety and racial disparities.","Benchmarks that check only explicit stereotypes or refusal compliance will miss this bias; audits need implicit scenario tasks and demographic-disaggregated refusal reporting to detect adultification bias.","Within a model family, larger models show increased implicit adultification bias, so alignment gains do not automatically scale with model size."],"supporting_citations":[{"why":"Supplies the adultification construct, the trait-rating battery, and the age-estimation task that the explicit LLM and T2I measurements are built on.","marker":"[29]"},{"why":"Defines adultification bias and documents its real-world decision contexts (school discipline, dating, sexual health) from which the four implicit scenarios are drawn.","marker":"[9]"},{"why":"Supplies the profile-generation method for eliciting implicit bias and the decision-bias metric that the LLM decision results are measured with.","marker":"[4]"},{"why":"Documents racial disparities in STI testing that motivate the 'STI test versus no STI test' scenario.","marker":"[31]"},{"why":"Supplies the race-coded name sets and the profile-based measurement variant that the name selection for the implicit scenarios is based on.","marker":"[87]"},{"why":"Supplies the 'Imagine a {trait} {age} {race} girl wearing a dress' prompt structure and the trait list for the image-generation tasks.","marker":"[8]"},{"why":"Provides the human-annotation methodology for evaluating generated images that the age and revealingness protocol is modeled on.","marker":"[37]"},{"why":"Provides the inter-annotator agreement statistic (an alpha-style reliability measure) used to support the consistency of the image annotations.","marker":"[44]"},{"why":"Documents the gap between text-focused and image-focused safety evaluations that motivates the multi-modal measurement of adultification bias.","marker":"[65]"}],"fun_headline_variants":["AI models show adultification bias against Black girls","Chatbots and image AIs judge Black girls as older, more sexual","Generative AI exhibits adultification bias in text and images","Black girls seen as older and more sexual by AI models","Adultification bias found in chat and image generators"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The human annotators who judged the generated images as older and more revealingly dressed do not themselves carry adultification bias; if they share the bias, the measured Black-White gap partly reflects the measurement instrument rather than the models alone.","fun_headline_variants_meta":{"raw":{"variants":["AI models show adultification bias against Black girls","Chatbots and image AIs judge Black girls as older, more sexual","Generative AI exhibits adultification bias in text and images","Black girls seen as older and more sexual by AI models","Adultification bias found in chat and image generators"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000565,"raw_usage":{"total_tokens":2730,"prompt_tokens":1045,"completion_tokens":1685,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":661,"completion_tokens_details":{"reasoning_tokens":1605}},"tokens_in":661,"tokens_out":1685,"duration_ms":10451,"temperature":1.0,"reasoning_tokens":1605,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:37:45.045054+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have the same annotators first estimate the ages of real photographs of Black and White girls whose true ages are known. If annotators systematically overestimate Black girls' ages even on real photos, adultification bias is present in the measurement itself; re-running the image comparison with annotator-bias calibration, or with automated age and clothing-coverage estimators, would then settle whether the models' Black-White output gap remains.","supporting_citations":[{"cited_title":"Goyal, Katie L","cited_arxiv_id":null,"evidence_quote":"Documents racial disparities in STI testing that motivate the 'STI test versus no STI test' scenario."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the inter-annotator agreement statistic (an alpha-style reliability measure) used to support the consistency of the image annotations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the gap between text-focused and image-focused safety evaluations that motivates the multi-modal measurement of adultification bias."}],"review_version":1}