{"id":"ac394f97-dfb0-4723-bb6d-d9704516a87a","arxiv_id":"2412.09668","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"VLMs generate more homogeneous stories about Black individuals with higher perceived racial phenotypicality, with effects varying by gender and model.","lead":"This study finds that vision-language models write more similar stories about Black people with more stereotypically Black facial features than about Black people with fewer such features. The finding suggests AI image-to-text systems can reproduce within-group stereotyping similar to documented human phenotypicality bias.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported observation count and mixed-model structure are inconsistent with the described within-image pairwise design; the 499,000-observation LMM likely pools dependent pairwise similarities, so the headline p-values may be artifacts of non-independence.","rationale":"I focused on the statistical unit-of-analysis problem rather than the reader's visual-similarity confound because if the p-values are artifacts, the central claim is not established even if the image stimuli are perfectly controlled. The reader correctly notes that the observation count is inconsistent, but their named weakest assumption is the visual confound; my concern is more fundamental and upstream. The paper's own limitation section acknowledges the embedding black box, but not the non-independence of pairwise similarities. The GANFD design is a genuine strength, and the direction of the effect is consistent across three models, which is some evidence; however, all three models use the same, likely misspecified, analysis pipeline, so consistency across models does not mitigate the concern. If the image-level reanalysis confirms the effect, the paper would be substantially stronger; if not, the current statistical evidence cannot support the headline. I therefore keep the reader's CONDITIONAL verdict rather than moving to REJECT, because the needed reanalysis is well-defined and could settle the issue.","tokens_in":12560,"tokens_out":10552,"duration_ms":102747,"concrete_test":"Reanalyze the Phenotypicality Model at the image level: for each of the 40 images, compute the mean cosine similarity among its 50 stories (one number per image), then fit a paired test or mixed model comparing high vs low phenotypicality within the 20 GANFD pairs (random intercept for Pair ID, n=40). Also refit the original model using only within-image pairs (n=49,000) with random intercepts for Pair ID and image. If the phenotypicality coefficient is no longer consistently positive and significant, or changes direction, the headline result is not robust to the unit-of-analysis problem. The authors should also provide the analysis script and derived pairwise data to verify the observation count.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the mixed-effects models in Sec 3.1–3.3. Sec 2.3 says cosine similarity is computed between all pairs of stories generated for each image; with 50 stories per image that is C(50,2)=1,225 observations per image and 40 images, giving 49,000 observations. But Tables S2–S4 report 499,000 observations for every VLM. The only way to reach 499,000 is to pool all 500 stories in each gender×phenotypicality condition (10 images × 50 stories) and take C(500,2)=124,750 pairwise similarities per condition, times four conditions. Under that pooling, a pairwise observation can involve two stories from different GANFD sets, so the stated random intercept \"Pair ID\" (Sec 2.4) is undefined for most observations. More importantly, each story embedding is reused in 499 pair comparisons, so the 499,000 observations are not independent; the effective sample size is at most the number of stories (2,000) or images (40), not 499,000. The reported standard errors (e.g., 0.0026 for b=0.044) are therefore almost certainly far too small, and the p<.001 values in Sec 3.1 cannot be taken at face value. The same problem affects the gender and interaction models, including the \"primarily driven by Black women\" finding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether vision-language models (VLMs) generate more homogeneous stories about Black individuals with higher versus lower phenotypicality, using images from the GAN Face Database and measuring pairwise cosine similarity of Sentence-BERT embeddings of generated stories. Three VLMs are tested (GPT-4o mini, GPT-4 Turbo, Llama-3.2). The authors report a significant main effect of phenotypicality in all models, a robust gender effect (women more homogeneous), and an interaction in two models that they interpret as the phenotypicality effect being driven primarily by Black women. They frame these results as evidence that VLMs mirror human phenotypicality bias.","tokens_in":12818,"tokens_out":5323,"duration_ms":46531,"significance":"The research question is timely and important: extending AI bias auditing from between-group to within-group comparisons is a genuine gap, and the use of controlled GAN-generated stimuli is a methodological strength that could enable cleaner causal claims than real-world images. If the findings survive appropriate statistical treatment, they would make a useful contribution to the study of intersectional bias in multimodal models. However, the paper's current statistical analysis has a serious unit-of-analysis inconsistency that undermines the reported p-values and effect sizes, and there are internal contradictions about which models show the interaction. These issues need to be resolved before the empirical claims can be assessed.","major_comments":[{"comment":"The reported observation count is inconsistent with the described design. The text states that cosine similarity is calculated between all combinations of stories generated for each image (Sec. 2.3) and that 50 stories per image yield 2,500 measurements per Pair ID (Sec. 2.2). With 40 images (10 sets × 2 genders × 2 phenotypicality levels), a within-image pairwise design would produce 40 × C(50,2) = 49,000 observations. Tables S2–S4 instead report 499,000 observations for every model, which equals 4 × C(500,2) — i.e., all 500 stories from the 10 images in a condition pooled together. This implies that most pairwise observations compare stories generated from different images, so the stated random intercept \"Pair ID\" (defined as the image set in Sec. 2.4) is undefined for the majority of observations. Moreover, in the pooled design each story embedding is reused in 499 pairwise comparisons, so the 499,000 observations are not independent; the effective sample size is at most 2,000 stories (or 40 images). The reported standard errors (e.g., 0.0026 for a coefficient of 0.044) are therefore almost certainly grossly inflated in precision, and the p < .001 values cannot be taken at face value. The authors must clarify the actual unit of analysis, re-fit the models with appropriate random effects (e.g., random intercepts for image and story identity, or analyze image/condition-level summaries), and re-report all significance tests.","section":"Sec. 2.3–2.4, Tables S2–S4"},{"comment":"There is an internal contradiction about which models exhibit the gender × phenotypicality interaction. Sec. 3.3 states that the interaction is significantly positive for GPT-4o mini and Llama-3.2, and not significant for GPT-4 Turbo. Sec. 4.2, however, says 'In two of three VLMs—GPT-4 Turbo and Llama-3.2—the effect of phenotypicality ... was significantly greater for women than for men.' These statements are incompatible. In addition, the simple-slopes results in Table S6 show that for both GPT-4o mini and Llama-3.2, the effect of phenotypicality within women has a 95% confidence interval that includes zero (e.g., GPT-4o mini women: 0.040, [-0.014, 0.094]; Llama-3.2 women: 0.053, [-0.012, 0.12]), which does not support the claim that the main effect is 'primarily driven by Black women.' The authors should reconcile the narrative with the reported statistics and either temper the interpretation or provide appropriate supporting analyses.","section":"Sec. 3.3 vs. Sec. 4.2, Tables S4 and S6"},{"comment":"The interpretation of the phenotypicality effect as a social-stereotyping phenomenon is potentially confounded by low-level visual similarity. The GAN manipulation holds other facial characteristics constant within a pair (lower vs. higher phenotypicality from the same set), but it does not ensure that the ten higher-phenotypicality images resemble each other more than the ten lower-phenotypicality images do (e.g., systematically darker skin or coarser hair). If the VLM generates more similar stories for images that share low-level visual features, the observed homogeneity difference could be a visual-similarity artifact rather than a reflection of social phenotypicality bias. The paper provides no control or covariate for pairwise image similarity. At minimum, the authors should discuss this alternative explanation explicitly and, ideally, include an image-similarity covariate or a per-image random slope to demonstrate that the effect is not explained by low-level visual resemblance.","section":"Sec. 2.1, Sec. 5"}],"minor_comments":[{"comment":"The power-analysis numbers do not match the described design: C(50,2) = 1,225, but the text says 1,245 measurements per pair of stimuli, and then says 2,500 per Pair ID. The arithmetic should be corrected and clarified.","section":"Sec. 2.2"},{"comment":"For GPT-4 Turbo, the text reports the interaction as 'b = -0.0048, p = .038' but also calls it not significant. The SE (0.0055) gives z ≈ -0.87, p ≈ .38, consistent with the likelihood-ratio test but not with p = .038. This is likely a typographical error but should be fixed.","section":"Sec. 3.3, Table S4"},{"comment":"The captions refer to 'all four VLMs' and Tables S2–S4 say 'across all four VLMs,' but only three models are used in the study. This should be corrected to 'three.'","section":"Figures 3 and Tables S2–S4"}],"recommendation":"major_revision","confidential_remarks":"The unit-of-analysis discrepancy is the key issue. If the 499,000-observation model is indeed based on pooled cross-image pairwise similarities, the current p-values are not credible. I would encourage the editor to request that the authors provide the analysis code and data so that the actual structure can be verified during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper asks a genuinely interesting question: whether VLMs, like humans, show within-group phenotypicality bias in homogeneity. The use of GAN-controlled face pairs to manipulate phenotypicality while holding other features constant is a smart design, and the idea that higher-phenotypicality Black women get more uniform stories is exactly the kind of finding that would extend AI-bias work beyond between-group comparisons.\n\nBut the statistical core does not hold up. The text says 50 stories per image and 2,500 cosine similarities per Pair ID, which would yield roughly 25,000 observations. The tables report 499,000. That number only arises if you pool all 500 stories in each gender×phenotypicality condition and compute every pairwise similarity. That means many comparisons are across different images and different Pair IDs, so the random intercept “Pair ID” is undefined for those pairs, and each story embedding is reused in hundreds of comparisons. The observations are massively non-independent; the standard errors (e.g., 0.0026) are far too small, and the p<.001 headline is an artifact. The same problem infects the gender and interaction results, including the “Black women drive it” claim.\n\nThere are also signs of carelessness: Figure 3 says “all four VLMs” though only three were used; Section 4.2 lists GPT-4 Turbo as one of the models with a significant interaction when the results section says its interaction was not significant and the LR test gives p=.38; and the 2,500-per-Pair-ID claim is arithmetically wrong (it should be 2,450). The visual-similarity confound you raised is real too: if high-phenotypicality images across sets share low-level visual features, the model might generate similar stories for that reason, with no covariate to rule it out.\n\nThe question is worth pursuing and the GAN-stimulus approach is a step forward, but this version's evidence does not support the conclusion. The paper needs a reanalysis that respects the nesting (stories within images within pairs) or uses cluster-robust inference, plus a check on visual similarity. If the effect survives that, it will be a real contribution. I'd send it to referees in the hope of forcing that reanalysis, but I wouldn't accept anything close to the current numbers.\n\nBest.","headline":"The question is good, but the statistics are broken: the reported p-values rest on non-independent, pooled pairwise similarities and cannot be taken at face value.","tokens_in":13365,"tokens_out":3375,"would_cite":false,"duration_ms":30923,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vision-language models generate more homogeneous stories about Black individuals with higher perceived racial phenotypicality.","keywords":["racial phenotypicality","homogeneity bias","vision-language models","stereotyping","sentence embeddings","cosine similarity","intersectionality","GAN-generated faces"],"falsifier":"Re-run the story-generation and similarity pipeline but add a covariate that measures the pairwise visual similarity of the face images themselves, for instance cosine similarity of image embeddings. If the phenotypicality effect on story homogeneity disappears or drops sharply once visual similarity is controlled, the central claim of stereotyping-driven homogeneity is not supported.","tokens_in":12327,"feed_emoji":"🤖","tokens_out":5491,"duration_ms":46599,"temperature":0.7,"pith_summary":"This paper asks whether vision-language models (VLMs) treat members of the same racial group differently based on how typically Black their facial features look. The authors claim that when a VLM writes a short story about a computer-generated Black face, the stories become significantly more similar to each other for faces rated higher in phenotypicality. The effect appears in all three models tested (GPT-4o mini, GPT-4 Turbo, and Llama-3.2), with Black women driving the effect in two of the three models. This matters because it extends homogeneity bias from language models to multimodal AI and suggests that even within a single racial category, AI reproduces the human tendency to stereotype people with more phenotypically Black features more strongly.","feed_headline":"AI stories get more uniform for more typically Black faces","feed_subtitle":"GPT-4o mini, GPT-4 Turbo, and Llama-3.2 all show the within-group homogeneity effect.","key_machinery":"The machinery is a measurement pipeline that combines three components: GAN Face Database pairs that hold identity constant while manipulating facial features associated with perceived Blackness; a free-generation prompt that asks the VLM to write a 50-word story about the individual in the image; and a homogeneity metric built from Sentence-BERT embeddings (all-mpnet-base-v2) that computes pairwise cosine similarity between story embeddings. Linear mixed-effects models with Pair ID as random intercepts test whether phenotypicality, gender, and their interaction predict this similarity, with the load-bearing assumption that higher cosine similarity means less individualized representation.","core_discovery":"The central claim is that VLMs generate more homogeneous stories about Black individuals with higher perceived racial phenotypicality than about those with lower phenotypicality, where homogeneity is measured as pairwise cosine similarity of sentence embeddings of the generated stories. Using 10 pairs of GAN-generated faces matched for identity but manipulated in phenotypicality, the authors elicited 50-word stories from three VLMs and fit mixed-effects models. They found significant positive effects of phenotypicality in all models (bs = 0.044, 0.15, and 0.080), significant gender effects with Black women represented more homogeneously than Black men, and a positive phenotypicality-by-gender interaction in GPT-4o mini and Llama-3.2, meaning the phenotypicality effect is largely carried by Black women. The paper interprets this as evidence that VLMs mirror the human racial phenotypicality bias documented in social psychology.","pith_inferences":["A plausible alternative mechanism is that high-phenotypicality images share more low-level visual features with each other than low-phenotypicality images do; a direct test would add image-similarity covariates to the mixed models to see if the phenotypicality effect survives.","A testable extension is to apply the same pipeline to real face photos matched by human-rated phenotypicality; if the effect disappears with real photos, it may be an artifact of GAN-generated feature distributions rather than a general property of VLMs.","Another extension is to check whether the homogeneity effect appears in other tasks such as image captioning or question answering, which would indicate whether it is a general representational bias or specific to narrative generation.","The interaction with gender suggests that the phenotypicality manipulation may also alter gender-typical appearance; future work should verify that the manipulation affects male and female faces equally on dimensions other than race."],"forward_implications":["VLMs inherit not only between-group but also within-group stereotypical homogeneity, so audits should measure variation across phenotypicality rather than only across racial categories.","Because the effect is concentrated in Black women for two of three models, intersectional analysis is necessary to identify who is most affected by AI stereotyping.","Embedding-based cosine similarity of generated texts can serve as a scalable quantitative audit tool for homogeneity bias in generative models.","The pattern mirrors documented human racial phenotypicality bias, suggesting that debiasing methods must address feature-based stereotyping rather than simply making models colorblind.","The mixed evidence across models, with GPT-4 Turbo showing no interaction, indicates that homogeneity bias varies by model architecture and training data."],"supporting_citations":[{"why":"Supplies the GAN Face Database stimuli, the computer-generated face pairs with controlled phenotypicality manipulations that are the paper's key experimental material.","marker":"Marsden et al., 2024"},{"why":"Provides the homogeneity measure (pairwise cosine similarity of sentence embeddings) and the effect size used for power analysis, establishing the methodological and quantitative benchmark extended here.","marker":"Lee et al., 2024"},{"why":"Sentence-BERT (all-mpnet-base-v2) is the embedding model used to compute story similarity, so the entire metric depends on it.","marker":"Reimers and Gurevych, 2019"},{"why":"The lme4 package implements the mixed-effects models used for all statistical tests of phenotypicality, gender, and interaction effects.","marker":"Bates et al., 2014"},{"why":"Provides the theoretical review of racial phenotypicality bias in humans, which the paper predicts VLMs will mirror.","marker":"Maddox, 2004"},{"why":"Empirical evidence that phenotypicality affects perception and categorization, supporting the meaningfulness of the manipulated feature differences.","marker":"Stepanova and Strube, 2018"},{"why":"Prior demonstration of intersectional bias (gender by skin tone) in AI systems, which the paper extends to generative VLMs.","marker":"Buolamwini and Gebru, 2018"}],"fun_headline_variants":["AI stories get more uniform for more typically Black faces","VLM homogenization strongest for Black women with typical traits","Higher phenotypicality yields more similar AI stories, especially for Black women","Within-group bias: VLMs portray high-phenotypicality Black individuals as more alike","Phenotypicality boosts AI story similarity, most for Black women"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The high-phenotypicality face images may be more visually similar to one another than the low-phenotypicality ones are, so the more similar stories could simply reflect more similar input images rather than social stereotyping.","fun_headline_variants_meta":{"raw":{"variants":["AI stories get more uniform for more typically Black faces","VLM homogenization strongest for Black women with typical traits","Higher phenotypicality yields more similar AI stories, especially for Black women","Within-group bias: VLMs portray high-phenotypicality Black individuals as more alike","Phenotypicality boosts AI story similarity, most for Black women"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1278,"prompt_tokens":971,"completion_tokens":307,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":587,"completion_tokens_details":{"reasoning_tokens":213}},"tokens_in":587,"tokens_out":307,"duration_ms":3426,"temperature":1.0,"reasoning_tokens":213,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:54:38.471925+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the story-generation and similarity pipeline but add a covariate that measures the pairwise visual similarity of the face images themselves, for instance cosine similarity of image embeddings. If the phenotypicality effect on story homogeneity disappears or drops sharply once visual similarity is controlled, the central claim of stereotyping-driven homogeneity is not supported.","supporting_citations":[],"review_version":1}