Pith. sign in

REVIEW 4 major objections 6 minor 29 references

Hidden Bias in the Machine: Stereotypes in Text-to-Image Models

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Text-to-image models systematically reproduce societal stereotypes across race, gender, age, and body type, an audit of 16,000 generated images finds.

desk verdict A broad, useful prompt-expansion audit with a genuinely new model, but the headline percentages lack the statistical support the text implies. read the letter →

arxiv 2506.13780 v1 pith:DB4U3THO submitted 2025-06-09 cs.CV cs.AIcs.CYcs.LG

classification cs.CVcs.AIcs.CYcs.LG
keywords text-to-imagestereotypesdemographicbiasgenderracialStableDiffusionFluxpromptaudit
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that popular text-to-image models do not neutrally depict prompts; they systematically skew images of people and places toward demographic stereotypes. The authors generated 16,000 images from Stable Diffusion 1.5 and Flux-1 across 160 topics, labeled them by racial, gender, age, and body-type categories, and compared the ratios with 8,000 images from an internet image search. They find that positive prompts (e.g., 'attractive', 'rich', 'peaceful') skew toward White, Western, male, and thinner depictions, while negative prompts skew toward racialized groups, overweight bodies, and male anger or aggression. If correct, the biases are not confined to occupational prompts but run through emotions, ideologies, family roles, religions, and place descriptions.

What carries the argument

The audit machinery is a prompt battery of 160 manually curated topics with multiple wording variations, run through two model families (Stable Diffusion 1.5 and Flux-1) with fixed sampling settings, followed by human visual labeling into coarse demographic bins (White/Black/East Asian/Latino/Middle Eastern/Other; Male/Female/Other; Young/Adult/Senior; Underweight/Average/Overweight). The comparison set of 8,000 search-engine images provides a baseline against which model skews are measured. The load-bearing step is the labeling: each generated image is assigned a demographic category by visual inspection, and the label ratios across prompt groups become the reported bias statistics.

What would settle it

Take a random sample of, say, 500 of the generated images from the paper's prompt set, have a diverse panel of labelers independently assign the same demographic categories, and compute agreement (e.g., Fleiss' kappa). If agreement is low (for instance below 0.6) for race, age, or somatotype, then the specific percentages are not stable. A second test: regenerate the same 160 topics with different random seeds and a different model version and check whether the sign and size of the gaps (e.g., high-income White 70% vs low-income 43%) reproduces; if the gaps vanish under reseeding, the findings may be sampling artifacts.

Watch

Extended reading notes

Core claim

The paper's central discovery is that text-to-image models exhibit stereotype-reinforcing disparities across a wide set of human-centric dimensions. For example, high-income roles are depicted as White (70% vs 43% for low-income in SD1.5), negative attributes and actions are overwhelmingly associated with males (91% and 87% respectively), positive place descriptions are 99% Western while negative place descriptions shift toward Africa and the Middle East, and neutral 'place of worship' prompts produce Christian imagery 83% of the time. Flux-1 shows even stronger skews, generating almost exclusively White individuals across most prompt groups. These patterns hold in both a UNet-based and a DiT-based model, suggesting the bias does not depend on one architecture.

Load-bearing premise

The statistics depend on human labelers categorizing each generated image into fixed demographic buckets by looking at it; the paper reports no measure of how often labelers agree, so if different labelers put the same image in different buckets, the percentages in the tables would shift.

Editorial extensions

If this is right

  • If biases persist across architectures, future models trained on AI-generated web content could inherit and amplify these skews as training data.
  • Users relying on text-to-image models for professional or creative work would unknowingly reproduce stereotyped imagery, reinforcing representational harms.
  • The search-engine comparison suggests the models are more extreme than web image distributions in some categories, so the problem is not merely a reflection of general internet content.
  • Simple prompt engineering without explicit demographic specification does not remove the biases; neutral prompts still produce skewed outputs.
  • The observed biases intersect (e.g., race and gender combine, as in 'sushi maker' being Asian female), so mitigation must be intersectional.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A concrete next test would be to measure inter-annotator agreement on a sample of the generated images; if agreement is low for race or age categories, the reported percentages may be unstable.
  • If the prompt battery is released as claimed, others could run the same prompts on newer models (e.g., SDXL, DALL-E 3) to track whether biases recede or shift with scale and fine-tuning.
  • The 'a person eating watermelon' result hints at cultural-context effects; a follow-up could systematically vary the food item and nationality to map how culinary prompts trigger gendered and racialized defaults.
  • The study's reliance on binary gender and coarse race bins may obscure non-binary and multiracial representations; an extension could use open-ended descriptions or continuous skin-tone scales.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper reports an empirical audit of social biases in two text-to-image models, Stable Diffusion 1.5 and Flux-1, and compares them with Google Image Search results. The authors generated over 16,000 images from 160 prompt topics spanning occupations, attributes, actions, ideologies, emotions, family descriptions, place descriptions, religion, and life events, then manually labeled the images for gender, race, age, somatotype, and place/religion categories. They report percentage breakdowns for positive versus negative prompt groups and conclude that the models exhibit significant disparities that often reinforce harmful stereotypes.

Significance. If the quantitative results were fully substantiated, the paper would be a useful broadening of bias evaluation in T2I models beyond the usual occupation/gender/skin-tone focus, covering actions, ideologies, emotions, family structures, place descriptions, and life events. The study has notable strengths: two architecturally distinct model families (UNet-based SD1.5 and DiT-based Flux-1), consistent generation settings, a comparison corpus from Google Image Search, and qualitative examples that illustrate the phenomena. However, the central quantitative claim of 'significant disparities' is currently unsupported by the reported evidence because the annotation process, sample sizes, and statistical measures are not disclosed.

major comments (4)
  1. [Section 3.1/3.2, Tables 1-3] The central quantitative claim rests on percentage breakdowns that are not auditable. The paper does not report the number of images per prompt group or per demographic cell, the number of annotators, or any inter-annotator reliability statistic (e.g., Cohen's kappa). Section 3.1 states only that images 'have been labeled by multiple human operators,' and Section 3.2 acknowledges the subjectivity of somatotype labeling. Without per-cell counts and reliability data, the percentages in Tables 1-3 could be dominated by a small number of topics or by annotator disagreement. The authors should report sample sizes, an agreement metric, and ideally release the annotations and prompts to support the headline claim of significant disparities.
  2. [Section 3.1] The exclusion of 'distorted, unclear, abstract, or nonsensical' images is described with no criteria, no counts, and no per-model or per-prompt-group exclusion rates. Post hoc filtering can differentially remove images by demographic content, which would directly bias every percentage in Tables 1-3. The filtering protocol and exclusion statistics must be reported for the results to be interpretable.
  3. [Abstract and Section 4] The word 'significant' is used throughout without any statistical support. For example, Section 4 states that 'racial bias was insignificant' for attribute prompts and describes percentage gaps as 'striking' and 'significant,' but no confidence intervals, hypothesis tests, or multiple-comparison corrections are reported for any table. The authors should compute appropriate uncertainty measures (e.g., bootstrap confidence intervals or chi-square tests) for the differences they highlight; otherwise the abstract's claim of 'significant disparities' is not justified.
  4. [Section 3.2] The grouping of prompts into 'positive' and 'negative' categories is a normative assumption that drives many of the paper's headline comparisons (e.g., positive vs. negative attribute, action, family, and emotion prompts). No validation of this valence classification is provided, and there is no neutral-prompt baseline or sensitivity analysis. Since a misclassification of even a few prompts could alter the reported percentages, the authors should justify the classification (e.g., with independent valence ratings) or show that the results are robust to alternative groupings.
minor comments (6)
  1. [Section 4, 'Bias in Roles'] There is a typo: 'the rations were close' should read 'the ratios were close,' and later 'SD.15' should be 'SD1.5'.
  2. [Conclusion] The sentence 'The observed biases in political views, emotions, family structures, places, religious depictions, and life events' is a sentence fragment and should be completed or merged with the following sentence.
  3. [References] Reference [11] is misattributed: the U.S. Census Bureau is not authored by Black Forest Labs. Please correct the author/organization and URL.
  4. [Figures and Tables] The text refers to 'Figure 4' and 'Tables 1, 2, 3,' but the figures are not numbered in the captions, and Table 2's column headers do not align with the rows. Also, the table numbering should be checked for consistency with the in-text citations.
  5. [Section 3.1] The arithmetic is unclear: 160 topics at 'over 50 images per topic' would yield about 8,000 images, not 16,000. The relationship between the two models, the number of images per prompt, and the total count should be stated explicitly.
  6. [Footnote 1] The promised release of the prompt benchmark has no URL or timeline; for reproducibility, the prompts and annotations should be made available with the paper.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical audit whose reported disparities are measured outcomes, not consequences of fitted parameters or self-citations.

full rationale

This paper is an empirical measurement study, not a derivation. The pipeline is: manually curated prompts → generate images with fixed checkpoints → human annotators label demographic categories → compare percentage distributions. The central claim (that T2I models exhibit disparities across gender, race, age, and somatotype) is a reported observation about generated image sets, not a quantity derived from an equation or a fitted parameter. The authors never fit a model to a subset of data and then "predict" a closely related quantity; the percentages in Tables 1–3 are descriptive statistics of human-assigned labels. The prompt categories ("positive" vs. "negative", high-income vs. low-income occupations) are author-chosen inputs, but they do not by construction force the reported demographic outcomes: for example, labeling a prompt as "negative attribute" does not mathematically imply that 91% of SD1.5 images will be labeled Male, or that negative-place prompts will yield 37% African-coded images. Those are contingent empirical findings. The paper's demographic label definitions (Section 3.2) are operational definitions for annotation, not self-referential derivations of the conclusion. There are no load-bearing self-citations: the references are to prior external bias audits and technical model papers, and no uniqueness theorem or prior result by the same authors is invoked to forbid alternatives. The acknowledged subjectivity in somatotype labeling and the absence of inter-annotator reliability statistics are methodological validity/robustness concerns, not circularity. For this reason, the appropriate circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The study relies on several hand-chosen design elements but no mathematical free parameters. The main dependencies are the subjective prompt construction, image filtering, and label definitions, along with the assumption that visual inspection can read demographic attributes.

free parameters (3)
  • Prompt wording per topic
    The exact prompt templates (e.g., 'a photo of a democratic/republican person') are hand-crafted and not released; results depend entirely on these choices.
  • Image filtering threshold
    Criteria for excluding 'distorted, unclear, abstract, or nonsensical' images are subjective and applied post hoc, affecting which images enter the frequency counts.
  • Annotation label set
    The race, gender, age, and somatotype categories are chosen by the authors and applied by human annotators, with no inter-annotator reliability reported.
assumptions (4)
  • domain assumption Gender, race, age, and somatotype can be reliably inferred from visual cues in generated images
    Section 3.2 defines labels based on observable visual cues; the entire measurement relies on this premise, which is acknowledged as subjective.
  • domain assumption Google Image Search provides a valid external reference for comparison
    Section 3.1 uses 8,000 Google images as a comparison baseline, but search results are algorithmically curated and not an unbiased ground truth.
  • domain assumption The selected 160 prompt topics are representative of real-world T2I usage
    The topics are manually chosen and may not reflect actual user prompt distributions or the diversity of global contexts.
  • domain assumption Stable Diffusion 1.5 and Flux-1 are representative enough of current T2I models to support general conclusions
    Only two models are tested; the paper generalizes its conclusions to T2I models broadly, while noting newer versions differ.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hidden Bias in the Machine: Stereotypes in Text-to-Image Models." pith.science (2026). https://pith.science/paper/DB4U3THO

@misc{pith2026250613780,
  author       = {Pith},
  title        = {Pith review of: Hidden Bias in the Machine: Stereotypes in Text-to-Image Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DB4U3THO}},
  note         = {Machine review of arXiv:2506.13780}
}
read the original abstract

Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To investigate these issues, we curated a diverse set of prompts spanning thematic categories such as occupations, traits, actions, ideologies, emotions, family roles, place descriptions, spirituality, and life events. For each of the 160 unique topics, we crafted multiple prompt variations to reflect a wide range of meanings and perspectives. Using Stable Diffusion 1.5 (UNet-based) and Flux-1 (DiT-based) models with original checkpoints, we generated over 16,000 images under consistent settings. Additionally, we collected 8,000 comparison images from Google Image Search. All outputs were filtered to exclude abstract, distorted, or nonsensical results. Our analysis reveals significant disparities in the representation of gender, race, age, somatotype, and other human-centric factors across generated images. These disparities often mirror and reinforce harmful stereotypes embedded in societal narratives. We discuss the implications of these findings and emphasize the need for more inclusive datasets and development practices to foster fairness in generative visual systems.

Figures

Figures reproduced from arXiv: 2506.13780 by the authors.

Figure 1
Figure 1. Random generated images from Flux-1 for prompts from actions category. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Random generated images from Flux-1 for prompts from actions category. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Random generated images from Flux-1 for prompts from actions category. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: SD1.5 percentages for positive and negative prompt groups (attributes, actions, roles, family descriptions, and emotions) across [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 21 canonical work pages

  1. [1]

    DALL-E-3: Improving Image Genera- tion with Better Captions.Computer Science, 2023

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, and etal. DALL-E-3: Improving Image Genera- tion with Better Captions.Computer Science, 2023. 1

  2. [2]

    Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets

    Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. In IJCNLP, 2021. 2

  3. [3]

    Generating In- terior Design from Text: A New Diffusion Model-Based Method for Efficient Creative Design.Buildings, 2023

    Junming Chen, Zichun Shao, and Bin Hu. Generating In- terior Design from Text: A New Diffusion Model-Based Method for Efficient Creative Design.Buildings, 2023. 1

  4. [4]

    PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Syn- thesis.arXiv:2310.00426, 2023

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Syn- thesis.arXiv:2310.00426, 2023. 1

  5. [5]

    Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, and Matthew Turk. Tibet: Identifying and Evaluating Biases in Text-to-Image Genera- (rounded %) White Black E.Asian Latino Middle E Other Male Female Other Underweight Average Overweight Young Adult Senior Attributes Positive 98 0 0 2 0 0 68 32 0 20 66 14 6 70 24 A...

  6. [6]

    Dall-eval: Probing the Reasoning Skills and Social Biases of Text-to- Image Generation Models

    Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the Reasoning Skills and Social Biases of Text-to- Image Generation Models. InCVPR, pages 3043–3054,

  7. [7]

    Publications Office of the European Union, 2024

    Europol Innovation Lab.Facing Reality? – Law Enforce- ment and the Challenge of Deepfakes – An observatory re- port. Publications Office of the European Union, 2024. 1

  8. [8]

    Beat Biden.https://www.youtube.com/ watch?v=kLMMxgtxQ1Y&t=32s, 2023

    GOP. Beat Biden.https://www.youtube.com/ watch?v=kLMMxgtxQ1Y&t=32s, 2023. 1

Show all 29 references
  1. [9]

    Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models.CVPR,

    Jiuxiang Gu, Jianfei Cai, Shafiq Joty, Li Niu, and Gang Wang. Look, Imagine and Match: Improving Textual-Visual Cross-Modal Retrieval with Generative Models.CVPR,

  2. [10]

    Exploring text- to-image generation models: Applications and cloud re- source utilization.Elsevier Computers and Electrical En- gineering, 123, 2025

    Sahani Jaiprakash and Choudhary Prakash. Exploring text- to-image generation models: Applications and cloud re- source utilization.Elsevier Computers and Electrical En- gineering, 123, 2025. 1

  3. [11]

    Black Forest Labs. U.s. census bureau.https://www. census.gov, 2020. 3

  4. [12]

    Flux.https://github.com/ black-forest-labs/flux, 2024

    Black Forest Labs. Flux.https://github.com/ black-forest-labs/flux, 2024. 1

  5. [13]

    StoryGAN: A Sequential Conditional GAN for Story Visualization .CVPR, 2019

    Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. StoryGAN: A Sequential Conditional GAN for Story Visualization .CVPR, 2019. 1

  6. [14]

    Stable Bias: Analyzing So- cietal Representations in Diffusion Models.NeurIPS, 2023

    Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. Stable Bias: Analyzing So- cietal Representations in Diffusion Models.NeurIPS, 2023. 1, 2

  7. [15]

    Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models

    Nila Masrourisaadat, Nazanin Sedaghatkish, Fatemeh Sar- shartehrani, and Edward A Fox. Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models. arXiv:2407.00138, 2024. 1, 2

  8. [16]

    Social Biases Through the Text-to-Image Generation Lens

    Ranjita Naik and Besmira Nushi. Social Biases Through the Text-to-Image Generation Lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 786–808, 2023. 1, 2

  9. [17]

    Multilingual Diversity Improves Vision- Language Representations.arXiv:2405.16915, 2024

    Thao Nguyen, Matthew Wallingford, Sebastin Santy, Wei- Chiu Ma, Sewoong Oh, Ludwig Schmidt, Pang Wei Koh, and Ranjay Krishna. Multilingual Diversity Improves Vision- Language Representations.arXiv:2405.16915, 2024. 2

  10. [18]

    Modelling agency — deep agency.https: //www.deepagency.com/, 2024

    Danny Postma. Modelling agency — deep agency.https: //www.deepagency.com/, 2024. 1

  11. [19]

    Hierarchical Text-Conditional Image Gen- eration with CLIP Latents.arXiv:2204.06125, 2022

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical Text-Conditional Image Gen- eration with CLIP Latents.arXiv:2204.06125, 2022. 1

  12. [20]

    High-Resolution Image Synthesis with Latent Diffusion Models .CVPR, 2022

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-Resolution Image Synthesis with Latent Diffusion Models .CVPR, 2022. 1

  13. [21]

    Fast High- Resolution Image Synthesis with Latent Adversarial Diffu- sion Distillation.arXiv:2403.12015, 2024

    Axel Sauer, Frederic Boesel, Tim Dockhorn, Andreas Blattmann, Patrick Esser, and Robin Rombach. Fast High- Resolution Image Synthesis with Latent Adversarial Diffu- sion Distillation.arXiv:2403.12015, 2024. 1

  14. [22]

    LAION- 400M: Open Dataset of CLIP-filtered 400 Million Image- Text Pairs.arXiv:2111.02114, 2021

    Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION- 400M: Open Dataset of CLIP-filtered 400 Million Image- Text Pairs.arXiv:2111.02114, 2021. 1

  15. [23]

    Andrew Shaw, Andre Ye, Ranjay Krishna, and Amy X. Zhang. Unsettling the Hegemony of Intention: Agonistic Image Generation.arXiv:2502.15242, 2025. 1, 2

  16. [24]

    ReStGAN: A step towards visually guided shopper experi- ence via text-to-image synthesis.WACV, 2020

    Shiv Surya, Amrith Setlur, Arijit Biswas, and Sumit Negi. ReStGAN: A step towards visually guided shopper experi- ence via text-to-image synthesis.WACV, 2020. 1

  17. [25]

    Survey of Bias In Text- to-Image Generation: Definition, Evaluation, and Mitiga- tion.arXiv:2404.01030, 2024

    Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Re- becca Pattichis, and Kai-Wei Chang. Survey of Bias In Text- to-Image Generation: Definition, Evaluation, and Mitiga- tion.arXiv:2404.01030, 2024. 1, 2

  18. [26]

    Fleet, Radu Soricut, Jason Baldridge, Mo- hammad Norouzi, Peter Anderson, and William Chan

    Su Wang, Chitwan Saharia, Ceslee Montgomery, Jordi Pont- Tuset, Shai Noy, Stefano Pellegrini, Yasumasa Onoe, Sarah Laszlo, David J. Fleet, Radu Soricut, Jason Baldridge, Mo- hammad Norouzi, Peter Anderson, and William Chan. Ima- gen Editor and EditBench: Advancing and Evaluati...

  19. [27]

    Hwang, Amy X

    Andre Ye, Sebastin Santy, Jena D. Hwang, Amy X. Zhang, and Ranjay Krishna. Cultural and Linguistic Diversity Im- proves Visual Representations.arXiv:2310.14356, 2024. 2

  20. [28]

    Text-to-Image Synthesis: A Decade Survey.arXiv:2411.16164, 2024

    Nonghai Zhang and Hao Tang. Text-to-Image Synthesis: A Decade Survey.arXiv:2411.16164, 2024. 1

  21. [29]

    SINE: SINgle Image Editing with Text-to-Image Diffusion Models.CVPR, 2022

    Zhixing Zhang, Ligong Han, Arnab Ghosh, Dimitris Metaxas, and Jian Ren. SINE: SINgle Image Editing with Text-to-Image Diffusion Models.CVPR, 2022. 1

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.