{"id":"967a7fa8-5e76-46e9-a306-4e5546621bbd","arxiv_id":"2509.08004","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 30 professions and 6,000 images, DALL-E 3 over-represented women, Stable Diffusion XL and Cascade over-represented men in high-status roles, and Emu was more balanced.","lead":"This study generated 50 images per profession with four text-to-image AI models and manually counted whether each image showed a man or a woman. It found that both Stable Diffusion models leaned male for high-status jobs, DALL-E 3 leaned female for most jobs, and Emu was closer to balanced.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"DALL-E's 'opposite' gender bias is confounded by OpenAI's backend prompt rewriting, which injects explicit gender into the actual prompt; the comparison may not be across equivalent text-to-image models.","rationale":"The reader's weakest assumption is the reliability of manual classification. That is a legitimate concern, but it is not the single most load-bearing one for this paper's central comparative claim. Even if a second annotator disagreed on 5% of images, the large reported margins for DALL-E (e.g., 82% female for surgeon, 78% female for CEO) would likely persist, so the qualitative pattern would survive. The prompt-rewriting discovery, however, affects every DALL-E profession and directly explains the paper's most surprising result. The paper's own examples show that 'A Doctor' was transformed into a prompt specifying 'a woman by gender'; thus the image generator did not receive the same input as the other models. Comparing DALL-E's output to SDXL/SC/Emu under these conditions is like comparing one system with an extra, unrequested demographic filter. The reader mentioned the rewriting in the rationale but did not elevate it to the weakest assumption. I therefore partially agree. The right verdict remains CONDITIONAL: the end-to-end observation is plausible and worth publishing, but the model-level attribution requires controlling for prompt rewriting. No change to the reader's verdict is needed.","tokens_in":16992,"tokens_out":7319,"duration_ms":93649,"concrete_test":"Take the exact rewritten prompt examples in the Discussion (e.g., the 'A Doctor' rewrite specifying a South Asian woman) and feed them to SDXL, SC, and Emu under the same generation settings. If these models also produce predominantly female images for gender-specified prompts, then DALL-E's female bias is fully explained by explicit gender injection. Additionally, if the DALL-E API permits any control over rewriting, query it with 'do not modify the prompt' while logging the returned revised prompt, and compare gender distributions between rewritten and non-rewritten conditions. If DALL-E still yields female-majority output for neutral prompts even when the prompt is not rewritten, the confound is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Materials and Methods section states that the prompts were simply profession names for SDXL, SC, and DALL-E (e.g., 'a Doctor'), while Emu used 'Create an image of ...'. However, the Discussion reports that the DALL-E API automatically rewrote 'A Doctor' into a fully specified prompt that explicitly chose 'a woman by gender' (and, in another attempt, 'a Middle Eastern male'). This means DALL-E's image generator did not receive the same neutral profession prompt as the other models; it received a prompt already containing a gender specification. The headline result—DALL-E returning female-majority images for 28 of 30 professions—may therefore be a property of OpenAI's prompt-rewriting policy, not of the underlying text-to-image model's learned association between profession and gender. The comparison is inequivalent: SDXL/SC/Emu outputs reflect the model's response to the profession word alone, whereas DALL-E outputs reflect a hidden demographic-injection layer. The paper openly acknowledges the rewriting and frames it as an explanation, but the central claim 'DALL-E exhibited almost opposite results' is then about the API service, not about the model. This confound is more load-bearing than label noise: even perfect manual labels would leave the cross-model comparison ambiguous.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical comparison of gender bias in four text-to-image systems: SDXL, Stable Cascade, DALL-E 3, and Meta's Emu. For each model and each of 30 professions, the authors generated 50 images, manually classified each person as male or female, and used binomial tests against an expected 1:1 ratio. The headline findings are that SDXL and Stable Cascade generate strongly male-dominated images for high-status/STEM professions and strongly female-dominated images for care/service roles; Emu is more balanced; and DALL-E 3 generates a female majority in 28 of 30 professions, which the authors attribute to OpenAI's automatic rewriting of prompts at the API backend. The paper also discusses possible causes and mitigation strategies for AI bias.","tokens_in":17271,"tokens_out":4227,"duration_ms":53364,"significance":"If the results are taken as measurements of the systems as encountered by users, the paper is a useful comparative contribution: it adds Emu to the set of models studied, quantifies bias with simple binomial tests, and documents the previously underappreciated phenomenon of backend prompt rewriting in DALL-E. The strengths are transparency about the basic counts, the simple falsifiable protocol, and the public code repository. The main scientific limitations are that the manual gender labels have no reliability check and that the DALL-E comparison is confounded by hidden prompt rewriting. Both are addressable through reframing and additional validation, so the core contribution is defensible in revised form.","major_comments":[{"comment":"The entire dataset consists of binary male/female labels described only as 'The classification of images depicting a man or woman was done manually.' No inter-rater reliability, second annotator, annotation instructions, or released label file is provided. This is load-bearing because many claims are near-ceiling (100%/0% cells), and a small label error rate on ambiguous images could change the exact count of DALL-E female-dominated professions and flip several 100% claims. Please report annotator agreement (e.g., Cohen's kappa on a re-annotated subset), describe how ambiguous images with multiple people, nonbinary appearance, or occluded faces were handled, and release the labels with the repository.","section":"Materials and Methods (manual classification)"},{"comment":"The headline claim that 'DALL-E exhibited almost opposite results' is confounded by OpenAI's backend prompt rewriting. The Methods state that DALL-E received prompts such as 'a Doctor,' but the Discussion documents that the API rewrote this into 'a medical professional ... a woman by gender ...' (and, in the no-modification attempt, 'a Middle Eastern male'). SDXL, SC, and Emu did not receive such injected demographic specifications. The comparison is therefore not across equivalent text-to-image models: the DALL-E outputs reflect the API service's post-processing policy rather than the model's response to the profession noun alone. The manuscript acknowledges the rewriting but still frames the result as a property of DALL-E. Please restrict the DALL-E claim to 'the DALL-E API service as used at the time of the study,' or, if the goal is model-level comparison, obtain generations without","section":"Results and Discussion (DALL-E 3)"},{"comment":"The paper conducts 120 binomial tests and reports that 'more than 110 out the 120 cases (91.6%) exhibit a statistically significant degree of gender bias' at the 0.05 level. No multiple-comparison correction is applied; at alpha=0.05 one would expect roughly 6 false positives under the null across 120 tests. Near-zero p-values for cells such as CEO, CFO, and doctor are robust, but the exact count of significant cells and the red/blue table in Figure 8 should be based on adjusted p-values (or explicitly presented as exploratory uncorrected testing). Please also clarify whether the '0.0' entries are all below 0.001 and report exact p-values in the supplemental material rather than only rounded values.","section":"Results (Figure 8 and binomial testing)"}],"minor_comments":[{"comment":"The text says 'resulting in a total of 6,000 images,' but the exclusion of Stable Cascade's 'Architect' and the reduced 93-image set for 'Astronaut' imply 5,843 images. Please correct the total or state it as 'approximately 6,000 images (with exclusions).'","section":"Materials and Methods (sample size)"},{"comment":"Model versions are described only as 'the latest versions available at the time.' Please list exact versions, access dates, and for Emu the specific WhatsApp/API configuration, since the backend behavior can change between versions and may affect reproducibility.","section":"Materials and Methods (prompts and models)"},{"comment":"The claim that Emu used user information while generating images is based on two phone numbers and visible national flags. This is an interesting observation, but it is not a controlled experiment; please temper the claim to 'suggestive evidence' and note other possible causes (e.g., IP geolocation or WhatsApp metadata) unless additional controls were performed.","section":"Discussion (Emu user information)"},{"comment":"There are several typographical errors and formatting artifacts, including 'pPrevious research,' 'As observeOpenAI achieves,' 'Aritifcal Intelligence,' 'AL-Powered,' and inconsistent spacing in the references. Please run a careful proofreading pass.","section":"Throughout"},{"comment":"The caption says the red block represents the gender with higher representation and blue the less dominant gender, but the figure itself is not included in the text in machine-readable form. Please ensure color-blind-safe labels or patterns are used and that the figure is legible in print.","section":"Figure 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is an empirical audit rather than a methodological advance, but it fits the journal's applied AI/fairness scope if the revision addresses the DALL-E API confound and label reliability. I would not reject on novelty grounds; the Emu comparison and documented prompt rewriting are of interest. The main risk is that the headline result is interpreted as a model-level property when the evidence supports a service-level property."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things worth knowing about this paper. It's the first comparison I've seen that puts Emu next to SDXL, Stable Cascade, and DALL-E 3 on occupation-based gender output, with 50 images per cell across 30 professions. And it documents, with concrete API responses, that DALL-E 3 rewrites neutral profession prompts into fully specified gender and ethnicity descriptions. That second observation is the most valuable part of the paper.\n\nCredit where due: the experimental setup is transparent about prompts, exclusions (Stable Cascade 'architect' produced buildings; astronaut images lacked faces), and the manual labeling. The GitHub repo has the generation scripts. The Emu-via-WhatsApp finding that the country code of the phone number changed the flag in generated politician images is a nice uncontrolled variable, and the authors flag it.\n\nThe soft spots are the usual ones for this kind of audit. The manual male/female labels have no inter-rater reliability check and the image-label dataset isn't released. For extreme cells like 100% or 98% this doesn't matter, but for DALL-E's 28-of-30 claim some of those professions are not extreme; a 5% labeling shift could change a few counts. They also run 120 binomial tests with no multiple-comparison correction, though most p-values are tiny. Model versions are described as 'latest at time of experiment' but not pinned to hashes or release dates.\n\nThe stress-test note is on target, but the paper is more self-aware than it gives it credit for. The authors explicitly report that 'A Doctor' got rewritten to specify 'a woman by gender' and even after asking not to modify, it still specified 'a Middle Eastern male.' So the DALL-E numbers describe the behavior of the API service, including its prompt-rewriting policy, not the bare model's learned associations. That said, the headline phrase 'DALL-E exhibited almost opposite results' is too clean; it should be framed as 'the DALL-E API returned...' throughout, because that is what they measured.\n\nBottom line: this is a useful time-stamped snapshot for people tracking text-to-image bias. It's not a definitive ranking of models. With better labeling protocol, multiple comparisons, pinned versions, and a softer causal framing, it would be a solid conference or journal contribution. I would not desk reject it; I'd send it to reviewers expecting revision. I probably won't cite it in the next year, but I might bring it to a reading group as an example of how API-side rewriting complicates black-box audits.","headline":"A small, honest comparative audit that adds Emu to the TTI gender-bias picture and documents DALL-E 3's backend prompt rewriting; worth refereeing with revisions, not as a definitive claim.","tokens_in":17751,"tokens_out":2689,"would_cite":false,"duration_ms":30726,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper measures gender bias in four text-to-image models and finds that all four deviate from a 1:1 ratio, with DALL-E 3 producing more female than male images in 28 of 30 professions.","keywords":["gender bias","text-to-image generation","occupational stereotypes","DALL-E 3","Stable Diffusion XL","Stable Cascade","Emu","binomial test"],"falsifier":"Re-annotate the 6,000 images with two independent raters who do not know which model produced each image; if the raters disagree on more than about 5% of the images, or an automated gender classifier of known accuracy disagrees with the manual labels, then the paper's exact gender ratios—and especially DALL-E's 28-of-30 female-heavy count—would not be reproducible.","tokens_in":16899,"feed_emoji":"⚖️","tokens_out":10261,"duration_ms":95360,"temperature":0.7,"pith_summary":"The paper asks whether four current text-to-image models—Stable Diffusion XL, Stable Cascade, DALL-E 3, and Emu—reproduce occupational gender stereotypes when asked to picture people in 30 jobs. It generates 50 images per profession per model, sorts each image by hand into male or female, and tests each cell against a 50:50 expectation with a binomial test. The answer is that none of the models is balanced: 91.6% of the 120 model–profession combinations deviate significantly from 1:1. Stable Diffusion XL and Cascade return 100% male CEOs, CFOs, and doctors and 100% female nurses and housekeepers; Emu is the least skewed; DALL-E 3 swings the other way, with women in 28 of 30 professions, which the authors attribute to the model's automatic prompt rewriting. The paper matters because these tools are already used in advertising and career-facing settings, and the authors leave open who should decide the target ratio.","feed_headline":"DALL-E flips gender bias, favoring women in 28 of 30 jobs","feed_subtitle":"None of four leading models hits a 1:1 gender balance across 30 professions","key_machinery":"The core instrument is a simple counting device: for each of four models, 50 images per profession (6,000 total) are generated and manually sorted into male/female; the binomial test then scores each profession-model cell against the hypothesis that the true gender split is 50:50. The DALL-E result is accompanied by observation of the API's automatic prompt rewriting, which the paper treats as the likely mechanism behind DALL-E's female-heavy outputs.","core_discovery":"The central finding is that gender bias in text-to-image generation is not a single direction: it is a systematic deviation from the 1:1 ratio in every model, but the sign of the deviation differs. Against the authors' hypothesis, DALL-E 3 did not favor men; it generated more women in 28 of 30 professions, e.g., 82% women for 'surgeon' and 76% for 'scientist,' while Stable Diffusion XL and Stable Cascade generated 100% male images for CEO, CFO, and doctor and 100% female images for nurse and housekeeper. Emu fell in between with more balanced counts. The authors tie DALL-E's flip to an observed backend behavior: the API rewrites prompts—'a Doctor' becomes a descriptive prompt specifying a wo","pith_inferences":["Inference: If DALL-E 3's backend continues to rewrite prompts to specify a gender and ethnicity even when users forbid modification, then the user's prompt is no longer the operative input; auditing the generated output alone cannot separate model bias from hidden prompt manipulation.","Inference: Because the paper counts only a binary male/female split, its near-100% claims could look different under a non-binary or intersectional annotation scheme; extending the same prompt set with multi-label categories would test whether 'bias' is a single axis.","Inference: The Emu country-code effect (Pakistani vs Ghanaian phone number changing flags in 'politician' outputs) is a separate, potentially confounded finding; a controlled experiment that varies only the country code of a fresh account would clarify whether location metadata leaks into generation or whether the difference is random.","Inference: The paper's 'who decides the target ratio?' question could be operationalized as a public-values survey: ask lay users whether they want job-level real-world gender shares, a global 50:50 split, or something else, and then measure which current model is closest to that stated preference."],"forward_implications":["All four models deviate significantly from a 1:1 gender split in 91.6% of the 120 profession-model combinations, so none can be called balanced by the paper's own test.","Stable Diffusion XL and Stable Cascade show the strongest stereotyping: 100% male for CEO, CFO, and doctor, and 100% female for nurse and housekeeper.","DALL-E 3 overshoots in the opposite direction: 28 of 30 professions return more female than male images, including 82% female for 'surgeon' and 76% for 'scientist,' even though real-world shares are much lower.","Emu is the least skewed of the four but still fails the 1:1 test in most professions.","The paper's observed automatic prompt rewriting means DALL-E's gender balance is at least partly engineered at the API level, not learned from the prompt's plain meaning."],"supporting_citations":[{"why":"Prior study of DALL-E's gender depiction of medical professions; provides the baseline (29.7% women) that the paper's DALL-E result reverses.","marker":"[6]"},{"why":"Study documenting demographic stereotypes in text-to-image generation, notably male software engineers and female housekeepers, which the paper compares against.","marker":"[7]"},{"why":"IEEE study identifying race and gender bias in Stable Diffusion, used to frame expected male bias.","marker":"[8]"},{"why":"T2IAT measurement of stereotypical biases in text-to-image models, used as prior evidence of Stable Diffusion gender stereotypes.","marker":"[9]"},{"why":"Bloomberg analysis reporting Stable Diffusion gives only 7% female doctors; the paper's 100% male doctor result extends it.","marker":"[23]"},{"why":"Technical report for Emu describing its two-stage training on 1.1 billion pairs and fine-tuning on high-quality images.","marker":"[15]"},{"why":"World Economic Forum statistic that 26% of data and AI roles are held by women, used to show DALL-E's 65% female AI researcher output overshoots reality.","marker":"[31]"},{"why":"Statistic that only 5.8% of airline pilots are women, used to contrast DALL-E's 73% female pilot images.","marker":"[32]"}],"fun_headline_variants":["DALL-E skews female, Stable Diffusion skews male in AI image test","No text-to-image model achieves gender balance, study finds","DALL-E's gender bias flips: more women, not men","Study reveals why DALL-E favors women over men","AI image bias is two-sided: DALL-E and SD opposite"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the authors' manual inspection reliably labels every generated image as male or female; the paper reports no second annotator, no inter-rater agreement, and no release of the image-label dataset, so a small rate of misclassification could change the headline near-100% and 28-of-30 counts.","fun_headline_variants_meta":{"raw":{"variants":["DALL-E skews female, Stable Diffusion skews male in AI image test","No text-to-image model achieves gender balance, study finds","DALL-E's gender bias flips: more women, not men","Study reveals why DALL-E favors women over men","AI image bias is two-sided: DALL-E and SD opposite"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000318,"raw_usage":{"total_tokens":1665,"prompt_tokens":811,"completion_tokens":854,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":764}},"tokens_in":555,"tokens_out":854,"duration_ms":9269,"temperature":1.0,"reasoning_tokens":764,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T23:49:48.535015+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the 6,000 images with two independent raters who do not know which model produced each image; if the raters disagree on more than about 5% of the images, or an automated gender classifier of known accuracy disagrees with the manual labels, then the paper's exact gender ratios—and especially DALL-E's 28-of-30 female-heavy count—would not be reproducible.","supporting_citations":[{"cited_title":"T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation","cited_arxiv_id":"2306.00905","evidence_quote":"T2IAT measurement of stereotypical biases in text-to-image models, used as prior evidence of Stable Diffusion gender stereotypes."},{"cited_title":"Humans Are Biased. Generative AI Is Even Worse","cited_arxiv_id":null,"evidence_quote":"Bloomberg analysis reporting Stable Diffusion gives only 7% female doctors; the paper's 100% male doctor result extends it."},{"cited_title":"Global Gender Gap Report 2020","cited_arxiv_id":null,"evidence_quote":"World Economic Forum statistic that 26% of data and AI roles are held by women, used to show DALL-E's 65% female AI researcher output overshoots reality."},{"cited_title":"Women airline pilots: numbers are growing, but still a pitiful percentage","cited_arxiv_id":null,"evidence_quote":"Statistic that only 5.8% of airline pilots are women, used to contrast DALL-E's 73% female pilot images."}],"review_version":1}