{"id":"fb29d504-029b-40a7-b277-c991b51ba905","arxiv_id":"2501.12433","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"DALL-E 3 disproportionately generates animals matching cultural stereotypes when prompted with trait adjectives, and a simple anti-stereotyping instruction only partially mitigates this.","lead":"The authors asked DALL-E 3 to generate images of loyal, wise, gentle, unfaithful, mischievous, and violent animals, then counted which species appeared. They find the model heavily favors cultural archetypes, such as owls for wisdom and foxes for unfaithfulness, and that adding 'do not stereotype animals' partially broadens the output.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The decisive gap is not annotation but the missing neutral control: trait-prompt frequencies are never compared to a no-adjective 'animal' baseline, so '100% dogs' and 'predominant owls' could reflect DALL-E's generic animal distribution rather than stereotype-driven association.","rationale":"I agree with the reader that the unvalidated manual annotation is a real weakness; without raw data or inter-rater reliability, the exact percentages are uncertain. However, the more fundamental threat to the central claim is construct validity: the study asks whether trait adjectives evoke stereotyped animals, but it never measures the model's animal distribution in the absence of an adjective. Section 3's 'dogs appeared exclusively in all 100 generations' is the paper's strongest result, yet if a neutral animal prompt also produces dogs in all 100 runs, the result says nothing about loyalty stereotypes. The same logic applies to all six traits. A control condition is standard in stereotype-elicitation experiments and is cheap to run. The paper's debiasing experiment (Section 4) does not fill this gap because it only shows that adding 'Do not stereotype animals' changes the distribution for two traits; it does not establish that the original trait prompt deviates from a neutral baseline. This is an internal-validity issue, not a disagreement with consensus. I am not saying the phenomenon is absent; the qualitative examples in Figure 2 suggest real stereotype propagation. But the publishable claim needs the neutral control to be decisive. Since the missing control is addressable by a simple additional experiment, conditional acceptance remains appropriate; I would keep the reader's verdict unchanged. The manuscript would also benefit from releasing annotated data and computing inter-rater reliability, but that is secondary.","tokens_in":7457,"tokens_out":6211,"duration_ms":65797,"concrete_test":"Run the same DALL-E 3 configuration used in Section 2 with the neutral control prompt 'Generate an image of an animal' 100 times. Label the resulting species distribution using the same protocol (ideally with two independent raters and Cohen's kappa). Compare the neutral distribution to each trait condition, e.g., by a chi-square test on the top-species counts or by computing P(species | trait) versus P(species | neutral). If the dog frequency under 'loyal' is not significantly higher than under the neutral prompt, or if owl/deer/fox frequencies are not significantly elevated relative to neutral, then the stereotype-specific claim is unsupported. Report raw counts and the comparison in the same style as Figure 1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim in Section 3 is that DALL-E 3 'predominantly' selects stereotyped animals for six trait prompts. All counts come from the authors' manual categorization (Section 2: 'We categorized the animals depicted in the images and counted the occurrences'), which the reader correctly flags as under-validated. But even if every label is perfect, the counts do not establish a stereotype-specific association because no neutral control condition is reported. The prompt template is 'Generate an image of a/an adj animal', and no equivalent 'Generate an image of an animal' run is provided. Consequently, 'dogs appeared exclusively in all 100 generations' for 'loyal' is uninterpretable without knowing how often DALL-E 3 outputs dogs for an unmodified animal prompt. If DALL-E's default animal distribution already favors dogs, owls, and predators, the apparent 'stereotyping' is just generic output bias, not trait-conditioned stereotyping. The paper's Limitations section lists only a small prompt set, a single model, and limited debiasing traits; it does not mention the missing baseline. The debiasing comparison in Section 4 partially addresses distribution shift for 'wise' and 'mischievous', but it compares modified prompts to original prompts, not to a neutral prompt, and covers only two traits. Thus the core inference that the adjective steers animal choice is missing a critical control.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether DALL-E 3 reproduces culturally familiar animal stereotypes in text-to-image generation. For six trait adjectives (loyal, wise, gentle, unfaithful, mischievous, violent), the authors prompted DALL-E 3 with \"Generate an image of a/an ADJ animal\" 100 times each, manually categorized the animal depicted in each of the 600 resulting images, and report frequency counts such as dogs appearing in all 100 \"loyal\" generations, owls dominating \"wise,\" deer leading \"gentle,\" foxes leading \"unfaithful,\" raccoons/foxes leading \"mischievous,\" and large predators leading \"violent.\" The paper also claims a dual-layered reinforcement, where the model not only selects the stereotyped animal but also depicts the trait visually (e.g., a fox sneaking near a henhouse), and reports a prompt-modification debiasing experiment for two traits (\"wise\" and \"mischievous\") using the suffix \"Do not stereotype animals.\" The authors conclude that VLMs inherit and propagate animal stereotypes and that prompt engineering can partially mitigate them.","tokens_in":7748,"tokens_out":4369,"duration_ms":41666,"significance":"If the central claim is confirmed, this is a useful extension of bias research from human-centric categories to non-human animal representations in generative models, an area the paper says is underexplored. The observation of dual-layered reinforcement (animal choice plus visual portrayal) is interesting and goes beyond simple label counting. The study is simple and transparent in design, and the use of repeated generations (100 per condition) is a reasonable starting point. However, the evidence as presented is too weak to support the paper's conclusions: there is no neutral control condition, no statistical testing or uncertainty quantification, no validation of the manual animal labeling, and no data or code release. The debiasing results are reported only qualitatively for two of six traits. Thus the significance is conditional on substantial additional methodological work.","major_comments":[{"comment":"The experiment never includes a neutral control condition. The prompt template is \"Generate an image of a/an ADJ animal\" and no unmodified \"Generate an image of an animal\" baseline is reported. Consequently, the counts in Figure 1, such as \"dogs appeared exclusively in all 100 generations\" for \"loyal,\" are uninterpretable as evidence of stereotype-driven association unless we know DALL-E 3's base-rate distribution of animals for an unqualified prompt. If the model's generic animal distribution already favors dogs, owls, deer, foxes, and predators, the trait-conditioned frequencies would not demonstrate stereotyping. Please add a neutral baseline run with the same sample size and report a formal comparison (e.g., a chi-square test, permutation test, or confidence intervals for the difference in proportions) between each trait condition and the baseline.","section":"Section 2, Prompt Formation; Section 3, Figure 1"},{"comment":"All quantitative results in Section 3 depend on the authors' manual categorization of the animal depicted in each of the 600 generated images, but the paper provides no annotation protocol, no inter-rater reliability, no automated validation, and no release of the categorized labels. A systematic misclassification of even a small fraction of images could change the reported \"exclusive\" (100% dogs) and \"predominant\" (owls, foxes, deer) statements. Please provide a detailed annotation rubric, use at least two independent annotators with an agreement metric (e.g., Cohen's kappa), and make the annotations, or the images with their labels, available for verification.","section":"Section 2, Image Generation"},{"comment":"The results are reported as raw frequency counts from 100 generations per prompt, with no confidence intervals, error bars, or significance tests. The abstract's claim of \"significant stereotyped instances\" is not supported by any inferential statistic. Please report uncertainties on the proportions (e.g., Wilson intervals) and, where relevant, tests for differences across prompts or against a baseline. Without such quantification, the strength and reliability of the reported associations cannot be assessed.","section":"Section 3, Figure 1"},{"comment":"The debiasing experiment is limited to two of the six traits and provides no quantitative evaluation. The text states that modified prompts \"resulted in a broader representation of animals\" and gives examples (kangaroos, gorillas, octopuses; monkeys, koalas, hamsters), but no counts, percentages, effect sizes, or statistical comparison to the original prompts are reported. The conclusion that prompt engineering is a \"lightweight and effective\" mitigation strategy is therefore not substantiated. Please provide the full frequency distributions for the modified prompts and quantify the change in diversity (e.g., number of distinct animal categories, Shannon entropy) with appropriate uncertainty.","section":"Section 4, Debiasing"},{"comment":"The six adjectives were explicitly hand-picked because they correspond to known animal stereotypes (the paper says so in Section 2). The experiment therefore shows that DALL-E 3 reproduces these particular stereotypes, but it does not support the broader claim in the Abstract and Conclusion that VLMs \"perpetuate animal stereotypes\" as a general phenomenon. No control adjectives (e.g., non-stereotyped, positive, or abstract traits) are included, and the Limitations section does not acknowledge that the trait set is selected from the stereotypes under study. Please soften the generalizing language or add control adjective conditions and a more systematic trait sample.","section":"Section 2, Prompt Formation; Section 5, Limitations; Abstract"}],"minor_comments":[{"comment":"The panels are referenced as (a)–(f) in the text, but the caption does not map each panel to its prompt or provide the numerical counts in the figure itself; adding explicit panel titles and a counts table would greatly improve readability.","section":"Figure 1"},{"comment":"The caption of Figure 3 appears to contain a copy-paste error: after the debiasing paragraph it lists \"(c) loyal (d) Wise (c) Gentle (d) Unfaithful (e) Mischievous (f) Violent,\" which does not match the figure's content or the surrounding text; please correct the caption.","section":"Figure 3"},{"comment":"The sentence on dogs appearing in the unfaithful context cites references [13,14] about negation in VLMs, but the cited works do not obviously support a claim about \"inconsistencies or overgeneralization in the model's understanding of traits\"; the link should be clarified or a more directly relevant citation provided.","section":"Section 3, Unfaithful animals"},{"comment":"The limitations list omits several issues that are central for an empirical measurement study: the missing neutral baseline, the lack of inter-rater reliability for the manual labeling, the absence of statistical uncertainty, and the unavailability of data/code for reproducibility.","section":"Section 5, Limitations"},{"comment":"The paper does not report the exact DALL-E 3 version, the access date, the API or interface used, or generation parameters (e.g., temperature, seed, default settings); such details are necessary for reproducibility of the frequency counts.","section":"Section 2, VLM (DALL-E 3)"},{"comment":"The claim that this is \"the first of its kind\" systematic examination of animal stereotyping in VLMs may be too strong given the existing bias evaluations in DALL-Eval [10] and implicit social bias work [11]; please qualify the novelty claim to specifically cover animal stereotypes.","section":"Abstract and Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be an accepted author manuscript with a placeholder DOI, but my review treats the provided text as the submitted version. The main risk is that the central quantitative claim is not yet supported by the evidence as presented: the missing neutral control alone invalidates the interpretation of the headline percentages. The other issues (annotation validation, no statistics, weak debiasing evidence) reinforce the need for revision. I believe the authors can address these within the manuscript's scope by adding a baseline condition, reporting uncertainty, releasing annotations, and softening the general claims. If the journal's format prevents adding new experiments, the claims should be substantially reduced to match what the current data can support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper shows something plausible but doesn't prove it. The cleanest problem is not the hand-labeling, though that is a problem too; it is the absence of a no-adjective baseline. Without \"Generate an image of an animal,\" we do not know if DALL-E 3 already defaults to dogs, owls, and foxes for generic prompts. The 100% dogs for \"loyal\" could be generic output bias rather than stereotype-driven selection.\n\nWhat is genuinely new: I am not aware of prior work on animal stereotypes in text-to-image models. The prompt-and-count audit is straightforward, and running 100 images per trait is a reasonable sample size for a first pass. The debiasing experiment, even if limited to two traits, suggests prompt engineering can shift the distribution, which is a useful practical note. The Limitations section is honest about the small prompt set, single model, and partial debiasing, though it omits the baseline and the annotation validation.\n\nSoft spots: First, the missing control, as above. Second, the manual categorization is unvalidated and the data is not released; without inter-rater reliability or at least a sample of labeled images, the frequencies are unverifiable. Third, the trait set is hand-picked from the stereotypes the study reports, so the design partially presupposes the outcome. That is acceptable for an existence proof but not for prevalence claims. Fourth, there are no statistical tests or effect sizes; \"predominantly\" is doing a lot of work. The debiasing results for two traits are qualitative. None of these are fatal, but together they make the quantitative claims weaker than the text implies.\n\nBottom line: this is a useful pilot for the bias-audit community, best read as a prompt for better-controlled follow-ups. A serious referee should see it, but the authors should add a neutral baseline, release the data and labels, validate the annotation, and run basic stats. I would engage with it if I worked on VLM bias, and I would cite it as the first step in this direction, with caveats.","headline":"Plausible and novel but under-controlled: animal stereotype claim needs a neutral baseline and released data before the frequencies mean anything.","tokens_in":8265,"tokens_out":2782,"would_cite":true,"duration_ms":28877,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that DALL-E 3, prompted with a trait adjective like 'loyal' or 'wise,' reliably chooses the culturally stereotyped animal — dogs for loyalty in all 100 runs, owls for wisdom, foxes for unfaithfulness — and then stages…","keywords":["generative AI","vision-language models","DALL-E 3","animal stereotypes","bias","text-to-image generation","prompt engineering","debiasing"],"falsifier":"Re-running the same six prompts on DALL-E 3 with two independent annotators labeling each generated image would settle the central claim: if their labels disagree substantially, or if the counts do not reproduce 'dogs in all 100 loyal images' and the other predominant associations, the reported stereotype frequencies would not hold.","tokens_in":7288,"feed_emoji":"🦉","tokens_out":7254,"duration_ms":63690,"temperature":0.7,"pith_summary":"The paper investigates whether a vision-language model reproduces cultural animal stereotypes when asked to generate images from simple trait prompts, and frames the study as the first systematic examination of animal stereotyping in such models. Using the fixed template 'Generate an image of a/an [trait] animal,' the authors ran six traits—loyal, wise, gentle, unfaithful, mischievous, violent—through DALL-E 3 one hundred times each, producing 600 images that they manually categorized by species. They report stark associations: dogs appear in all 100 'loyal' images, owls dominate 'wise,' deer lead 'gentle,' foxes lead 'unfaithful,' raccoons and foxes lead 'mischievous,' and large predators lead 'violent.' The authors argue this matters because AI image generators can quietly reinforce one-dimensional, culturally specific views of animals, and they show that appending the instruction 'Do not stereotype animals' visibly widens the species diversity for the two traits tested.","feed_headline":"DALL-E 3 turns 'loyal' into dogs 100 percent of the time","feed_subtitle":"Owls, foxes, deer, raccoons and big predators get the same cultural shorthand in 600 test images; a warning prompt only partly fixes it.","key_machinery":"The mechanism that carries the argument is a minimal trait-prompt probe: the fixed template 'Generate an image of a/an adj animal,' varying only the trait adjective, run 100 times per trait, followed by the authors' manual categorization of the depicted species and a qualitative reading of the scene. This design isolates the model's adjective-to-species association while keeping prompt wording neutral. The companion mechanism is the debiasing probe, which appends 'Do not stereotype animals' to the same template and compares the resulting species distribution.","core_discovery":"The paper's central finding is that DALL-E 3 does not merely reflect a loose cultural tendency; it produces near-monoculture outputs. For 'loyal animal,' 100 of 100 images show dogs. For 'wise,' owls are the predominant choice, with elephants second. For 'gentle,' deer dominate, with rabbits secondary. For 'unfaithful,' foxes lead, though dogs and cats also appear. For 'mischievous,' raccoons and foxes dominate. For 'violent,' bears, lions, and tigers are the main outputs. The authors also observe a second layer: the model visually dramatizes the trait—a fox sneaking near a henhouse with a cunning expression, a bear roaring in a forest—so the stereotype is reinforced both in species selection and in the depicted scene. They further report that a modified prompt, 'Generate an image of a/an [trait] animal. Do not stereotype animals,' increases diversity for 'wise' and 'mischievous' but does not eliminate the bias.","pith_inferences":["Editorial inference: the same six prompts could be run on other text-to-image models (e.g., Stable Diffusion) to test whether the stereotype pattern is unique to DALL-E 3's training distribution or common to internet-scale image-text data.","Editorial inference: because the six trait words are loaded with Western cultural meanings, the observed mappings are likely culture-specific; repeating the probe in other languages or regions could reveal different stereotype sets (e.g., different 'wisdom' animals).","Editorial inference: the authors' manual counts have no reported inter-rater reliability, so the precise percentages (including 100% dogs) should be treated as indicative; an automated species classifier or a second annotator could either confirm or weaken the headline numbers.","Editorial inference: a direct testable extension is to measure whether the debiasing prompt reduces not only species diversity but also the visual reinforcement (cunning expressions, aggressive poses), since the paper reports the latter only qualitatively."],"forward_implications":["If DALL-E 3's behavior is representative, text-to-image systems systematically propagate culturally specific animal stereotypes rather than producing neutral or diverse depictions of nature.","The near-exclusive 'dogs for loyalty' result means a user asking for a 'loyal animal' will get essentially one species, so the model's output space is far narrower than the prompt implies.","Because a simple appended instruction changes the species distribution for two traits, some stereotyping is addressable at inference time without retraining.","The persistence of bias after the debiasing prompt suggests the associations live partly in the model's weights and training data, so prompt-level fixes are partial at best.","The dual-layered reinforcement (choosing the stereotyped species and staging the trait visually) implies that even diverse prompts may inherit behavioral clichés in the image composition."],"supporting_citations":[{"why":"Supplies the text-to-image model family under test; the paper treats it as the state-of-the-art VLM for image generation.","marker":"[2]"},{"why":"Precedent that large language models persistently associate a social group with a trait (Muslims with violence), the pattern this paper transfers to animals.","marker":"[8]"},{"why":"Introduces a measure of stereotypical bias in pretrained language models with controlled prompts, the template for this paper's prompt probing.","marker":"[9]"},{"why":"Prior work probing reasoning skills and social biases (gender, skin tone) in text-to-image models, which this study extends from human categories to animals.","marker":"[10]"},{"why":"Evidence that vision-language models contain implicit social biases, motivating the check for animal stereotypes in a VLM.","marker":"[11]"},{"why":"Used to explain why dogs also appear for 'unfaithful': the model's overgeneralization or inconsistency with trait understanding.","marker":"[13,14]"}],"fun_headline_variants":["DALL-E 3 says 'loyal' equals dog in all 100 test images","DALL-E 3 stereotypes animals: 'wise' owls, 'unfaithful' foxes","Warning prompt only partly fixes DALL-E 3's animal bias","100% of DALL-E 3's 'loyal' images are dogs","DALL-E 3 reinforces animal stereotypes, even with anti-bias prompt"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every reported frequency depends on the authors' manual categorization of the animals in the 600 generated images, which is presented without an annotation protocol or independent verification.","fun_headline_variants_meta":{"raw":{"variants":["DALL-E 3 says 'loyal' equals dog in all 100 test images","DALL-E 3 stereotypes animals: 'wise' owls, 'unfaithful' foxes","Warning prompt only partly fixes DALL-E 3's animal bias","100% of DALL-E 3's 'loyal' images are dogs","DALL-E 3 reinforces animal stereotypes, even with anti-bias prompt"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001264,"raw_usage":{"total_tokens":5141,"prompt_tokens":874,"completion_tokens":4267,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":4160}},"tokens_in":490,"tokens_out":4267,"duration_ms":31051,"temperature":1.0,"reasoning_tokens":4160,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:13:37.641422+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the same six prompts on DALL-E 3 with two independent annotators labeling each generated image would settle the central claim: if their labels disagree substantially, or if the counts do not reproduce 'dogs in all 100 loyal images' and the other predominant associations, the reported stereotype frequencies would not hold.","supporting_citations":[{"cited_title":"Discovering the Cognitive Bias of Toxic Language through Metaphorical Concept Mappings","cited_arxiv_id":null,"evidence_quote":"Evidence that vision-language models contain implicit social biases, motivating the check for animal stereotypes in a VLM."},{"cited_title":"Generate an image of a/an adj animal","cited_arxiv_id":null,"evidence_quote":"Supplies the text-to-image model family under test; the paper treats it as the state-of-the-art VLM for image generation."},{"cited_title":"and Sutskever, I., 2021, July","cited_arxiv_id":null,"evidence_quote":"Precedent that large language models persistently associate a social group with a trait (Muslims with violence), the pattern this paper transfers to animals."},{"cited_title":"Gender Bias in Text-to-Video Generation Models: A case study of Sora","cited_arxiv_id":"2501.01987","evidence_quote":"Prior work probing reasoning skills and social biases (gender, skin tone) in text-to-image models, which this study extends from human categories to animals."}],"review_version":1}