REVIEW 3 major objections 5 minor 41 references
Evaluating and comparing gender bias across four text-to-image models
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read This paper measures gender bias in four text-to-image models and finds that all four deviate from a 1:1 ratio, with DALL-E 3 producing more female than male images in 28 of 30 professions.
desk verdict A small, honest comparative audit that adds Emu to the TTI gender-bias picture and documents DALL-E 3's backend prompt rewriting; worth refereeing with revisions, not as a definitive claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core instrument is a simple counting device: for each of four models, 50 images per profession (6,000 total) are generated and manually sorted into male/female; the binomial test then scores each profession-model cell against the hypothesis that the true gender split is 50:50. The DALL-E result is accompanied by observation of the API's automatic prompt rewriting, which the paper treats as the likely mechanism behind DALL-E's female-heavy outputs.
What would settle it
Re-annotate the 6,000 images with two independent raters who do not know which model produced each image; if the raters disagree on more than about 5% of the images, or an automated gender classifier of known accuracy disagrees with the manual labels, then the paper's exact gender ratios—and especially DALL-E's 28-of-30 female-heavy count—would not be reproducible.
Extended reading notes
Core claim
The central finding is that gender bias in text-to-image generation is not a single direction: it is a systematic deviation from the 1:1 ratio in every model, but the sign of the deviation differs. Against the authors' hypothesis, DALL-E 3 did not favor men; it generated more women in 28 of 30 professions, e.g., 82% women for 'surgeon' and 76% for 'scientist,' while Stable Diffusion XL and Stable Cascade generated 100% male images for CEO, CFO, and doctor and 100% female images for nurse and housekeeper. Emu fell in between with more balanced counts. The authors tie DALL-E's flip to an observed backend behavior: the API rewrites prompts—'a Doctor' becomes a descriptive prompt specifying a wo
Load-bearing premise
The load-bearing premise is that the authors' manual inspection reliably labels every generated image as male or female; the paper reports no second annotator, no inter-rater agreement, and no release of the image-label dataset, so a small rate of misclassification could change the headline near-100% and 28-of-30 counts.
Editorial extensions
If this is right
- All four models deviate significantly from a 1:1 gender split in 91.6% of the 120 profession-model combinations, so none can be called balanced by the paper's own test.
- Stable Diffusion XL and Stable Cascade show the strongest stereotyping: 100% male for CEO, CFO, and doctor, and 100% female for nurse and housekeeper.
- DALL-E 3 overshoots in the opposite direction: 28 of 30 professions return more female than male images, including 82% female for 'surgeon' and 76% for 'scientist,' even though real-world shares are much lower.
- Emu is the least skewed of the four but still fails the 1:1 test in most professions.
- The paper's observed automatic prompt rewriting means DALL-E's gender balance is at least partly engineered at the API level, not learned from the prompt's plain meaning.
Reading between the lines
- Inference: If DALL-E 3's backend continues to rewrite prompts to specify a gender and ethnicity even when users forbid modification, then the user's prompt is no longer the operative input; auditing the generated output alone cannot separate model bias from hidden prompt manipulation.
- Inference: Because the paper counts only a binary male/female split, its near-100% claims could look different under a non-binary or intersectional annotation scheme; extending the same prompt set with multi-label categories would test whether 'bias' is a single axis.
- Inference: The Emu country-code effect (Pakistani vs Ghanaian phone number changing flags in 'politician' outputs) is a separate, potentially confounded finding; a controlled experiment that varies only the country code of a fresh account would clarify whether location metadata leaks into generation or whether the difference is random.
- Inference: The paper's 'who decides the target ratio?' question could be operationalized as a public-values survey: ask lay users whether they want job-level real-world gender shares, a global 50:50 split, or something else, and then measure which current model is closest to that stated preference.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an empirical comparison of gender bias in four text-to-image systems: SDXL, Stable Cascade, DALL-E 3, and Meta's Emu. For each model and each of 30 professions, the authors generated 50 images, manually classified each person as male or female, and used binomial tests against an expected 1:1 ratio. The headline findings are that SDXL and Stable Cascade generate strongly male-dominated images for high-status/STEM professions and strongly female-dominated images for care/service roles; Emu is more balanced; and DALL-E 3 generates a female majority in 28 of 30 professions, which the authors attribute to OpenAI's automatic rewriting of prompts at the API backend. The paper also discusses possible causes and mitigation strategies for AI bias.
Significance. If the results are taken as measurements of the systems as encountered by users, the paper is a useful comparative contribution: it adds Emu to the set of models studied, quantifies bias with simple binomial tests, and documents the previously underappreciated phenomenon of backend prompt rewriting in DALL-E. The strengths are transparency about the basic counts, the simple falsifiable protocol, and the public code repository. The main scientific limitations are that the manual gender labels have no reliability check and that the DALL-E comparison is confounded by hidden prompt rewriting. Both are addressable through reframing and additional validation, so the core contribution is defensible in revised form.
major comments (3)
- [Materials and Methods (manual classification)] The entire dataset consists of binary male/female labels described only as 'The classification of images depicting a man or woman was done manually.' No inter-rater reliability, second annotator, annotation instructions, or released label file is provided. This is load-bearing because many claims are near-ceiling (100%/0% cells), and a small label error rate on ambiguous images could change the exact count of DALL-E female-dominated professions and flip several 100% claims. Please report annotator agreement (e.g., Cohen's kappa on a re-annotated subset), describe how ambiguous images with multiple people, nonbinary appearance, or occluded faces were handled, and release the labels with the repository.
- [Results and Discussion (DALL-E 3)] The headline claim that 'DALL-E exhibited almost opposite results' is confounded by OpenAI's backend prompt rewriting. The Methods state that DALL-E received prompts such as 'a Doctor,' but the Discussion documents that the API rewrote this into 'a medical professional ... a woman by gender ...' (and, in the no-modification attempt, 'a Middle Eastern male'). SDXL, SC, and Emu did not receive such injected demographic specifications. The comparison is therefore not across equivalent text-to-image models: the DALL-E outputs reflect the API service's post-processing policy rather than the model's response to the profession noun alone. The manuscript acknowledges the rewriting but still frames the result as a property of DALL-E. Please restrict the DALL-E claim to 'the DALL-E API service as used at the time of the study,' or, if the goal is model-level comparison, obtain generations without
- [Results (Figure 8 and binomial testing)] The paper conducts 120 binomial tests and reports that 'more than 110 out the 120 cases (91.6%) exhibit a statistically significant degree of gender bias' at the 0.05 level. No multiple-comparison correction is applied; at alpha=0.05 one would expect roughly 6 false positives under the null across 120 tests. Near-zero p-values for cells such as CEO, CFO, and doctor are robust, but the exact count of significant cells and the red/blue table in Figure 8 should be based on adjusted p-values (or explicitly presented as exploratory uncorrected testing). Please also clarify whether the '0.0' entries are all below 0.001 and report exact p-values in the supplemental material rather than only rounded values.
minor comments (5)
- [Materials and Methods (sample size)] The text says 'resulting in a total of 6,000 images,' but the exclusion of Stable Cascade's 'Architect' and the reduced 93-image set for 'Astronaut' imply 5,843 images. Please correct the total or state it as 'approximately 6,000 images (with exclusions).'
- [Materials and Methods (prompts and models)] Model versions are described only as 'the latest versions available at the time.' Please list exact versions, access dates, and for Emu the specific WhatsApp/API configuration, since the backend behavior can change between versions and may affect reproducibility.
- [Discussion (Emu user information)] The claim that Emu used user information while generating images is based on two phone numbers and visible national flags. This is an interesting observation, but it is not a controlled experiment; please temper the claim to 'suggestive evidence' and note other possible causes (e.g., IP geolocation or WhatsApp metadata) unless additional controls were performed.
- [Throughout] There are several typographical errors and formatting artifacts, including 'pPrevious research,' 'As observeOpenAI achieves,' 'Aritifcal Intelligence,' 'AL-Powered,' and inconsistent spacing in the references. Please run a careful proofreading pass.
- [Figure 2 caption] The caption says the red block represents the gender with higher representation and blue the less dominant gender, but the figure itself is not included in the text in machine-readable form. Please ensure color-blind-safe labels or patterns are used and that the figure is legible in print.
Circularity Check
No circularity: the study is an empirical measurement; acknowledged confounds (manual labeling, DALL-E backend rewriting) are validity risks, not circular derivations.
full rationale
This paper reports an empirical measurement of gender ratios in AI-generated images; it does not derive or predict its headline result from an input parameter, fitted constant, or self-citation. The counts and percentages come from manually classifying generated images ("The classification of images depicting a man or woman was done manually"), and the binomial tests compare those observed counts to a fixed 1:1 null, so the p-values are computed from data rather than forced by construction. The DALL-E result is presented as an observation, and the paper explicitly attributes it to an independently observed mechanism: OpenAI rewrites the submitted prompt at its backend, as shown by the API response examples in the Discussion. That is a documented confound and a comparison-validity concern, not a circular step: the female-majority output is not equivalent to the input prompt by definition, and the paper does not claim otherwise. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and the authors cite no self-derived results as load-bearing premises. The acknowledged weaknesses�manual labeling without inter-rater reliability, the DALL-E API prompt rewriting, and Emu's apparent use of user information�are threats to generalizability or internal validity, but none of them makes the central claim reduce to its own inputs. Accordingly, no circularity step meets the evidentiary bar of quoting a specific equation or self-citation chain that forces the result.
Assumptions & free parameters
assumptions (3)
- domain assumption Generated images can be reliably assigned to one of two gender categories (male or female) by a human viewer.
- domain assumption A sample of 50 images per model-profession combination is representative of the model's behavior for that prompt.
- domain assumption The model versions tested are the ones implied by the model names and by the API used at the time of the experiment.
Cite this review
Pith. "Pith review of Evaluating and comparing gender bias across four text-to-image models." pith.science (2026). https://pith.science/paper/YX43S3X7
@misc{pith2026250908004,
author = {Pith},
title = {Pith review of: Evaluating and comparing gender bias across four text-to-image models},
year = {2026},
howpublished = {\url{https://pith.science/paper/YX43S3X7}},
note = {Machine review of arXiv:2509.08004}
}
read the original abstract
As we increasingly use Artificial Intelligence (AI) in decision-making for industries like healthcare, finance, e-commerce, and even entertainment, it is crucial to also reflect on the ethical aspects of AI, for example the inclusivity and fairness of the information it provides. In this work, we aimed to evaluate different text-to-image AI models and compare the degree of gender bias they present. The evaluated models were Stable Diffusion XL (SDXL), Stable Diffusion Cascade (SC), DALL-E and Emu. We hypothesized that DALL-E and Stable Diffusion, which are comparatively older models, would exhibit a noticeable degree of gender bias towards men, while Emu, which was recently released by Meta AI, would have more balanced results. As hypothesized, we found that both Stable Diffusion models exhibit a noticeable degree of gender bias while Emu demonstrated more balanced results (i.e. less gender bias). However, interestingly, Open AI's DALL-E exhibited almost opposite results, such that the ratio of women to men was significantly higher in most cases tested. Here, although we still observed a bias, the bias favored females over males. This bias may be explained by the fact that OpenAI changed the prompts at its backend, as observed during our experiment. We also observed that Emu from Meta AI utilized user information while generating images via WhatsApp. We also proposed some potential solutions to avoid such biases, including ensuring diversity across AI research teams and having diverse datasets.
Reference graph
Works this paper leans on
-
[1]
Improving Image Generation with Better Captions
Betker, James, et al. Improving Image Generation with Better Captions. OpenAI, 2023, cdn.openai.com/papers/dall-e-3.pdf
work page 2023
-
[2]
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, Robin, et al. “High-Resolution Image Synthesis with Latent Diffusion Models.” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2022, https://doi.org/10.48550/arXiv.2112.10752
-
[3]
Computer-Generated Inclusivity: Fashion Turns to ‘diverse’ Ai Models
Demopoulos, Alaina. “Computer-Generated Inclusivity: Fashion Turns to ‘diverse’ Ai Models.” The Guardian, 3 Apr. 2023, www.theguardian.com/fashion/2023/apr/03/ai-virtual-models-fashion-brands. Accessed 10 May 2024
work page 2023
-
[4]
Antonevics, Juris. Examining Algorithmic Bias in AL-Powered Credit Scoring: Implications for Stakeholders and Public Perception in an EU Country. 2023. Dipòsit Digital de la Universitat de Barcelona , http://hdl.handle.net/2445/201366
work page 2023
-
[5]
A Comprehensive Study on Bias in Artificial Intelligence Systems
Kartal, Elif. “A Comprehensive Study on Bias in Artificial Intelligence Systems.” International Journal of Intelligent Information Technologies , vol. 18, no. 1, 23 Sept. 2022, pp. 1–23, https://doi.org/10.4018/IJIIT.309582
-
[6]
Gender bias in generative artificial intelligence text-to-image depiction of medical students
Currie, G., Currie, J., Anderson, S., & Hewis, J. “Gender bias in generative artificial intelligence text-to-image depiction of medical students”. Health Education Journal. 2024. https://doi.org/10.1177/00178969241274621 7.Bianchi, Kalluri, et al. “Demographic Stereotypes in Text-to-Image Generation.” 2023, https://doi.org/10.1145/3593013.3594095
arXiv 2024
-
[8]
Identifying Race and Gender Bias in Stable Diffusion AI Image Generation,
A. Chauhan et al., "Identifying Race and Gender Bias in Stable Diffusion AI Image Generation," 2024 IEEE 3rd International Conference on AI in Cybersecurity (ICAIC) , Houston, TX, USA, 2024, pp. 1-6, doi: 10.1109/ICAIC60265.2024.10433840
arXiv 2024
-
[9]
T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation
Wang, Jialu, et al. “T2IAT: Measuring Valence and Stereotypical Biases in Text-To-Image Generation.” ArXiv (Cornell University) , 1 Jan. 2023, https://doi.org/10.48550/arxiv.2306.00905. Accessed 01 March 2025
work page Pith review arXiv doi:10.48550/arxiv.2306.00905 2023
Show all 41 references
-
[10]
Implications of Identity in AI: Creators, Creations, and Consequences
Tadimalla, Sri Yash, and Mary Lou Maher. “Implications of Identity in AI: Creators, Creations, and Consequences.” Proceedings of the AAAI Symposium Series , vol. 3, no. 1, 20 May 2024, pp. 528–535, https://doi.org/10.1609/aaaiss.v3i1.31268
2024 doi
- [11]
-
[12]
DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-To-Image Generation Models
Cho, Jaemin, et al. “DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-To-Image Generation Models.” ArXiv (Cornell University) , 8 Feb. 2022, https://doi.org/10.48550/arxiv.2202.04053
- [13]
-
[14]
Dall-E 3
“Dall-E 3.” OpenAI , openai.com/index/dall-e-3. Accessed 12 May 2024
2024
-
[16]
Stable Diffusion XL
“Stable Diffusion XL.” Hugging Face , huggingface.co/docs/diffusers/en/using-diffusers/sdxl. Accessed 11 May 2024
2024
-
[17]
Stable Cascade
“Stable Cascade .” Hugging Face, huggingface.co/stabilityai/stable-cascade. Accessed 11 May 2024
2024
-
[18]
A Call to Action on Assessing and Mitigating Bias in Artificial Intelligence Applications for Mental Health
Timmons, Adela C., et al. “A Call to Action on Assessing and Mitigating Bias in Artificial Intelligence Applications for Mental Health.” Perspectives on Psychological Science , vol. 18, no. 5, 9 Dec. 2022, pp. 1062–1096, https://doi.org/10.1177/17456916221134490
2022 doi
-
[19]
U.S. Senate Aritifcal Intelligence Insight Forum
Evans, Arthur C. “U.S. Senate Aritifcal Intelligence Insight Forum.” American Psychological Association, 8 Nov. 2023, www.schumer.senate.gov/imo/media/doc/Arthur%20Evans%20
2023
-
[20]
Meta's AI is accused of being RACIST: Shocked users say Mark Zuckerberg's chatbot refuses to imagine an Asian man with a white woman
“Meta's AI is accused of being RACIST: Shocked users say Mark Zuckerberg's chatbot refuses to imagine an Asian man with a white woman.” Daily Mail. Accessed 20 October 2024
2024
-
[21]
We ‘messed up’ with Black Nazi Blunder, Google Co-Founder Admits
Titcomb, James. “We ‘messed up’ with Black Nazi Blunder, Google Co-Founder Admits.” Yahoo! Finance , 4 Mar. 2024, finance.yahoo.com/news/messed-black-nazi-blunder- google-085157301.html. Accessed 15 May 2024
2024
-
[22]
Google Suspends Gemini AI Chatbot’s Ability to Generate Pictures of People
Chan, Kelvin CHAN, and Matt O’Brien. “Google Suspends Gemini AI Chatbot’s Ability to Generate Pictures of People.” Yahoo! Finance , 22 Feb. 2024, finance.yahoo.com/news/google-suspends-gemini-chatbots-ability-143732908.html. Accessed 15 May 2024
2024
-
[23]
Humans Are Biased. Generative AI Is Even Worse
Dina Bass, and Leonardo Nicoletti. “Humans Are Biased. Generative AI Is Even Worse.” Bloomberg. www.bloomberg.com/graphics/2023-generative-ai-bias/. Accessed 22 Oct. 2024
2023
-
[24]
ImageNet: A large-scale hierarchical image database,
J. Deng, W. Dong, et al. "ImageNet: A large-scale hierarchical image database," 2009 IEEE Conference on Computer Vision and Pattern Recognition , Miami, FL, USA, 2009, pp. 248-255, doi: 10.1109/CVPR.2009.5206848
2009
-
[25]
Improving Diffusion Models as an Alternative to Gans, Part 1
Vahdat, Arash, and Karsten Kreis. “Improving Diffusion Models as an Alternative to Gans, Part 1”. NVIDA Technical Blog, 12 June 2023, developer.nvidia.com/blog/improving-diffusion- models-as-an-alternative-to-gans-part-1. Accessed 15 May 2024
2023
-
[26]
Accessed 14 May 2024
“Laion.” LAION , laion.ai. Accessed 14 May 2024
2024
-
[27]
Common Crawl - Open Repository of Web Crawl Data
“Common Crawl - Open Repository of Web Crawl Data.” Common Craw , commoncrawl.org. Accessed 15 May 2024
2024
-
[28]
U-Net Convolutional Networks for Biomedical Image Segmentation
Ronneberger, Olaf. “U-Net Convolutional Networks for Biomedical Image Segmentation.” Image Processing for Medicine , 1 Mar. 2017, https://doi.org/10.1007/978-3-662-54345-0_3
2017 doi
-
[29]
Stable Diffusion with Diffusers
“Stable Diffusion with Diffusers.” Hugging Face , huggingface.co/blog/stable_diffusion. Accessed 15 May 2024
2024
-
[30]
Official Code for Stable Cascade
“Official Code for Stable Cascade.” GitHub, Stability AI, github.com/Stability-AI/Stable Cascade. Accessed 15 May 2024
2024
-
[31]
Global Gender Gap Report 2020
“Global Gender Gap Report 2020”. World Economic Forum , 2020, www3.weforum.org/docs/ WEF_GGGR_2020.pdf
2020
-
[32]
Women airline pilots: numbers are growing, but still a pitiful percentage
“Women airline pilots: numbers are growing, but still a pitiful percentage”. CAPA Center for Aviation. Accessed 22 October 2024
2024
-
[33]
Astronaut Fact Book
“Astronaut Fact Book”. NASA. https://www.nasa.gov/reference/astronaut-fact-book/. Accessed 22 October 2024
2024
-
[34]
Women in Science
“Women in Science”. UNESCO. https://uis.unesco.org/sites/default/files/documents/fs55-women-in-science-2019-en.pdf Accessed 22 October 2024
2019
-
[35]
Fairness and Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, and Mitigation Strategies, Social Science Research Network
Ferrara, Emilio. “Fairness and Bias in Artificial Intelligence: A Brief Survey of Sources, Impacts, and Mitigation Strategies, Social Science Research Network”, 21 Apr. 2023, https://doi.org/10.3390/sci6010003
2023 doi
-
[36]
Bias in Artificial Intelligence Algorithms and Recommendations for Mitigation
Nazer, Lama H., et al. “Bias in Artificial Intelligence Algorithms and Recommendations for Mitigation.” PLOS Digital Health , vol. 2, no. 6, 22 June 2023, https://doi.org/10.1371/journal.pdig.0000278
2023 doi
-
[37]
Online images amplify gender bias
Guilbeault, et al. “Online images amplify gender bias”. Nature 626, 1049–1055 (2024). https://doi.org/10.1038/s41586-024-07068-x
2024 doi
-
[38]
Gender Diversity Crisis in AI
“Gender Diversity Crisis in AI.” Nesta , 17 July 2019, www.nesta.org.uk/press-release/gender-diversity-crisis-ai-less-14-ai-researchers-are-women-nu mbers-decreasing-over-last-10-years. Accessed 15 May 2024
2019
-
[39]
AI and the quest for diversity and inclusion: a systematic literature review
Shams, R.A., Zowghi, D. & Bano, M. “AI and the quest for diversity and inclusion: a systematic literature review”. AI Ethics (2023). https://doi.org/10.1007/s43681-023-00362-w
2023 doi
-
[40]
AI Fairness 360: An Extensible Toolkit for Detecting, Understanding, and Mitigating Unwanted Algorithmic Bias
Rachel K, et al. “AI Fairness 360: An Extensible Toolkit for Detecting, Understanding, and Mitigating Unwanted Algorithmic Bias.” Meta , 27 Sep. 2023, https://doi.org/10.48550/arxiv.2309.15807
-
[41]
Gender differences in individual variation in academic grades fail to fit expected patterns for STEM
O’Dea, et al. “Gender differences in individual variation in academic grades fail to fit expected patterns for STEM”. Nat Commun 9, 3777. 2018. https://doi.org/10.1038/s41467-018-06292-0
2018 doi
-
[42]
Perceptions of stereotypes applied to women who publicly communicate their STEM work
McKinnon, et al. “Perceptions of stereotypes applied to women who publicly communicate their STEM work”. Humanit Soc Sci Commun 7, 160. 2020. https://doi.org/10.1057/s41599-020-00654-0
2020 doi
-
[43]
GitHub - Zoyahammad/BiasResearch: This Repository Contains Python Scripts Used for SC, SDXL, DALL-E 3
Carli, et al. Stereotypes About Gender and Science: Women ≠ Scientists. Psychology of Women Quarterly , 40 (2), 244-260. https://doi.org/10.1177/0361684315622645 Figures and Figure Captions Figure 2: Results from four Image Generation Models. Comparison between the results in ...
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.