{"id":"66b93a7e-2103-4d61-9ae4-b08e1a265559","arxiv_id":"2509.07050","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A benchmark of 13 image-generation models finds that most amplify occupational gender stereotypes, producing men in 93% of male-stereotyped prompts and 22.5% of female-stereotyped prompts, while one model approached parity.","lead":"Researchers tested 13 commercial text-to-image models with 75 gender-neutral prompts and found they usually depict men, especially for male-stereotyped jobs, amplifying real-world gender gaps. One model, Amazon's Nova Canvas, came much closer to gender parity, which the paper argues shows bias is a design choice, not inevitable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The amplification claim for female-stereotyped professions is contradicted by the paper's own data: generated images show 22.5% men vs 17.0% in labor data, yet Table 11 reports -5.61 percentage points and text says 'fewer male images.'","rationale":"The reader's weakest assumption (LLM-retrieved labor statistics) is legitimate and should be addressed, but the most load-bearing problem is internal: Table 11's female-category amplification value is inconsistent with the paper's own aggregate numbers. Section 3.2 reports 17.0% male in U.S. labor for female-stereotyped professions; Section 4.1 reports 22.5% male in generated images. The model is therefore 5.5pp more male than the baseline, not 5.6pp less. The text and tables treat the two categories with different metrics (male share for male professions, female share for female professions) or contain a sign error, invalidating the ANOVA in Table 12 and the associated t-tests. The later relative-amplification metric actually shows negative amplification for female professions (-16.7%), so the paper's own analyses conflict. If the corrected analysis shows no significant difference in amplification by category, the headline claim 'LMMs systematically amplify stereotypes' must be narrowed. This is independently checkable from raw image-level data and labor stats, so the current version should not be accepted without correction. Hence REJECT for the present manuscript; a corrected reanalysis could support a conditional accept.","tokens_in":13231,"tokens_out":11949,"duration_ms":126077,"concrete_test":"Recover the per-profession image-level judgments and labor statistics (the authors should release these, including the withheld prompts) and recompute for each of the 25 female-stereotyped professions: mean(generated_male_share - labor_male_share). Also compute mean(generated_female_share - labor_female_share) separately. Then re-run the one-way ANOVA in Table 12 and the one-sample t-tests using a single, pre-specified metric (male share or distance from parity) for both categories. If the female-category mean is +5.5 rather than -5.6 and the ANOVA no longer reaches significance, the abstract and §4.4 must be revised to say amplification is specific to male-stereotyped professions (or to the gender gap), not that LMMs systematically amplify stereotypes generally.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is that LMMs amplify real-world occupational gender stereotypes, which requires comparing generated male share to labor-force male share. The paper's own aggregate numbers contradict the reported female-category amplification. §3.2 Table 2 gives U.S. female-stereotyped professions a mean labor male share of 17.03%. §4.1 Table 6 gives 22.5% male in generated images for the same category. Thus generated images contain more men than the labor baseline (+5.5 pp), i.e., they move toward parity, not away from it. Yet Table 11 lists -5.61 percentage points for stereotypically-female professions and §4.4 says the generators produced 'fewer male images... indicating an underrepresentation of men relative to labor statistics.' This is internally inconsistent: either Table 11 uses the female share (77.5% - 83.0% = -5.5) while the text says male share, or the sign is reversed, or Tables 2 and 6 are incompatible. In any case, the one-way ANOVA in Table 12 (F=14.46, p=.0004) and the one-sample t-test for female professions (p=.15) are computed on an inconsistent variable. With the correct +5.5 pp female value, the female-category mean is positive and the difference from the male category (+11.9 pp) is far smaller (rough t≈1.4, p≈0.17 under reported SDs), so the claimed significant difference in amplification by stereotype may vanish. The headline finding 'systematically amplify stereotypes' is therefore not supported as stated; at most, male-stereotyped professions are amplified while female-stereotyped professions move toward parity. The later relative metric (dM/dL) in §4.4 confirms negative amplification for female professions (-16.7% aggregate), contradicting the percentage-point table's prose.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Aymara Image Fairness Evaluation, a benchmark that uses 75 gender-neutral profession prompts to test 13 commercial text-to-image LMMs. Prompts are generated programmatically through the Aymara AI SDK; 965 generated images are scored by an unnamed LLM-as-a-judge from the same SDK; and the resulting male/female rates are compared with U.S. and global labor statistics. The paper claims three headline findings: (1) LMMs systematically reproduce and amplify occupational gender stereotypes, moving male-stereotyped professions to 93.0% male and female-stereotyped professions to 22.5% male; (2) models show a default-male bias, producing 68.3% male for non-stereotypical professions; and (3) bias varies strongly across models, with Gen-4 near parity and Recraft V3 most biased. The paper also reports a significant interaction between model and prompt stereotype and proposes fairness scores. The overarching contribution is a scalable, automated, cross-model benchmark; the central quantitative claim is the amplification effect relative to real-world labor data.","tokens_in":13635,"tokens_out":7680,"duration_ms":87866,"significance":"If the claims survive scrutiny, the paper would provide one of the largest cross-model comparisons of gender bias in text-to-image generation, with an automated pipeline that can be re-run as models change. The strongest assets are the identical-prompt design across 13 models, the explicit handling of refusals, and the finding that one model (Nova Canvas) approaches parity while others do not, which is informative regardless of the amplification analysis. However, the paper's headline 'systematic amplification' claim is currently contradicted by its own aggregate numbers, the labor-force baselines are obtained via LLM retrieval rather than direct official database access, and the evaluator model is not disclosed. These issues are central and require re-analysis before the contribution can be assessed.","major_comments":[{"comment":"The paper's own data contradict the reported female-category amplification. U.S. labor statistics for stereotypically female professions give 17.03% men (Table 2), while generated images show 22.5% men (Table 6). The difference is +5.5 percentage points, not -5.61. The Table 11 value of -5.61 is consistent with the female share (77.5% - 83.0%), whereas the text interprets it as 'fewer male images' and the male category is computed as male share (93.0 - 81.1 = +11.94). The ANOVA in Table 12 therefore compares a male-share difference for male-coded professions with a female-share difference for female-coded professions, so the F(1,48)=14.46, p=.0004 result is not interpretable. With the sign corrected, the female mean is about +5.5 pp; the gap from the male category shrinks from about 17.6 pp to about 6.4 pp, and under the reported SDs a two-sample comparison is not significant (approximat","section":"§4.4, Tables 2, 6, 11"},{"comment":"The amplification analysis rests on labor-force baselines obtained by asking four LLMs (GPT-4o, Gemini 1.5 Pro, Claude 3 Sonnet, Perplexity) to retrieve BLS/ILO statistics, rather than by direct queries to official databases. The authors then take the median of these estimates as ground truth. Low inter-model variance (U.S. average SD=4.13) only shows agreement among LLMs, not accuracy. A systematic bias in the retrieval step would directly change the sign, magnitude, and statistical significance of every amplification result in Tables 10-13. The authors should query the official databases directly or release the per-profession retrieved values and sources so readers can verify them. This is load-bearing, not a presentation detail.","section":"§3.2"},{"comment":"The LLM-as-a-judge model is never named. The text says 'We used the Aymara Python SDK to score all 965 generated images' and reports 96.4% agreement with one human rater (kappa=0.92), but it does not state which LLM made the judgments, which version, what judgment prompt was used, or whether the judge is one of the 13 evaluated models. If the judge is a similar proprietary model, it may share the latent biases under test. Moreover, a single human rater cannot establish inter-rater reliability, and it is unclear whether this rater is one of the authors. Please disclose the judge model and version, the exact scoring instruction, and provide independent multi-rater validation or at least a second rater.","section":"§3.4"},{"comment":"The image-level ANOVAs (F(2,962)=268.01 and the two-way ANOVA with residual df=926) treat all 965 images as independent. However, images are nested within 75 prompts and 13 models: the same prompt is sent to every model, and each model generates 75 (sometimes 74/66) images. This clustering means the effective sample size is far smaller than 965, and the reported p-values, especially for the interaction in §4.3, are likely overstated. The authors should use cluster-robust standard errors or a mixed-effects model with random intercepts for prompt and model, and should report the corresponding p-values.","section":"§4.1, §4.3"}],"minor_comments":[{"comment":"Table 2 is titled 'Men in Generated Images (%)' but it reports labor-force statistics. This mislabeling likely contributed to the sign inconsistency in §4.4; it should read 'Men in Labor Force (%).'","section":"Table 2 title"},{"comment":"Reference [9] is incomplete: 'arXiv:2303.XXXXX' is a placeholder, not a citable identifier.","section":"References"},{"comment":"The model is called 'Titan G1 V2' in Table 9 but 'Titan Image Generator v2' in Table 5 and 'Titan G1 v2' in §3.3. Pick one consistent name.","section":"Table 9"},{"comment":"There are spacing artifacts in the PDF: 'F airness', 'F uture W ork', 'AUTOMA TED'. These are presumably rendering issues but should be corrected in the source.","section":"Throughout"},{"comment":"The full 75-prompt set is not released, and only sample prompts are shown. For a benchmark meant to be reusable, the full prompt set (or a controlled access procedure) should be provided; otherwise the category validation in §3.2 cannot be independently reproduced.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The author is affiliated with Aymara, and the paper uses the company's proprietary SDK both to generate prompts and to score images. This is a potential conflict of interest and a fit concern for a peer-reviewed venue; the manuscript should include a competing-interests statement and should disclose any funding from Aymara. The incomplete reference [9] and the mislabeled Table 2 also look like signs of rushed preparation. My recommendation is based on the correctable but central issues outlined above; if the corrected amplification analysis no longer supports the headline claim, the authors will need to substantially reframe the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe one thing you need to know: the paper's headline amplification claim contradicts its own data. For stereotypically-female professions, generated images are 22.5% male (Table 6) against a U.S. labor baseline of 17.0% (Table 2) — that's a +5.5 point shift towards parity, not amplification. Yet Table 11 reports -5.61 points and the text claims 'fewer male images.' The sign is wrong. Once you correct it, the male-stereotyped professions are still amplified (+11.9 points), but female-stereotyped professions move toward parity. The ANOVA on amplification (F=14.46, p=.0004) is computed on the inconsistent negative values; with the corrected positive female mean, the difference between categories is no longer significant (roughly t=1.4, p=0.17). So the abstract's 'systematically not only reproduce but actually amplify' is not supported as written.\n\nThat said, the paper has real value. It's the largest cross-model snapshot I know of: 13 current commercial LMMs, 75 standardized gender-neutral prompts, 965 images, and a single validated LLM judge with 96.4% agreement to a human rater. The finding that Nova Canvas approaches parity while Recraft V3 sits at 73% male is an actionable, concrete result. The method of procedurally generated prompts and LLM-as-a-judge is not new, but applying it at this scale is a legitimate extension.\n\nThe soft spots beyond the sign error are the ones the reader flagged: the labor statistics come from four LLMs retrieving BLS/ILO data rather than direct queries; the judge model is unnamed; only one human rater validated the judge; and the full prompt set is withheld. All are fixable with more transparency. Also, the repeated emphasis on the author's own Aymara SDK creates an appearance of a commercial pitch, but the underlying evaluation is reproducible in principle.\n\nThe sign error is the load-bearing flaw. It doesn't sink the entire paper—the cross-model variation and the parity result stand—but it invalidates the central amplification claim as currently reported. I'd send this to referees with a clear request to fix the sign, re-run the amplification analysis, and release the prompts and judge identity. The benchmark is worth having; the conclusion needs to be rewritten to say that male-stereotyped professions are amplified while female-stereotyped ones are not.\n\nRecommendation: major revision with re-analysis, not desk reject. But don't cite it until the numbers are corrected.\n\nBest","headline":"Useful cross-model benchmark, but the amplification claim is reversed for female-stereotyped professions and the key ANOVA likely doesn't survive the correction.","tokens_in":14151,"tokens_out":4634,"would_cite":false,"duration_ms":48374,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T01","68T50","62P35"],"pacs":[],"model":"deepseek-v4-flash","headline":"Commercial text-to-image models reproduce and amplify occupational gender stereotypes, generating men in 93% of male-stereotyped professions and only 22.5% of female-stereotyped ones.","keywords":["gender bias","text-to-image generation","large multimodal models","LLM-as-a-judge","occupational stereotypes","bias amplification","fairness benchmark","automated evaluation"],"falsifier":"Pull the same 75 professions directly from Bureau of Labor Statistics and International Labour Organization tables and recompute the bias-amplification analysis; if the official medians differ from the LLM-retrieved medians by more than a few percentage points in the stereotyped categories, the paper's amplification finding would not replicate.","tokens_in":13119,"feed_emoji":"🖼️","tokens_out":7226,"duration_ms":73281,"temperature":0.7,"pith_summary":"This paper argues that modern text-to-image models do not merely reflect occupational gender stereotypes; they amplify them relative to real-world labor statistics. Across 13 commercial models and 75 gender-neutral prompts, male-stereotyped professions produced images of men 93.0% of the time, female-stereotyped professions only 22.5%, and non-stereotyped professions 68.3%—a default-male bias. The paper reports large cross-model variation, with overall male representation from 46.7% to 73.3%, and one model (Nova Canvas) not differing from parity on stereotyped categories. The paper reads that variation as evidence that high bias is a design choice, not an inevitability, and argues for standardized automated benchmarks to track fairness.","feed_headline":"AI image generators show men for 93% of male-stereotyped jobs","feed_subtitle":"Across 13 text-to-image models, neutral prompts default to men 68% of the time; one model reaches near parity.","key_machinery":"The load-bearing mechanism is the Aymara Image Fairness Evaluation pipeline: procedural prompt generation (75 gender-neutral prompts in three stereotype categories), zero-shot image generation from 13 commercial LMMs (965 images), LLM-as-a-judge scoring of each image as man or not-man, and statistical validation of the prompt categories against LLM-retrieved US and global labor data. The scoring was validated against human ratings at 96.4% agreement (kappa 0.92). The quantitative engine is the bias-amplification metric, which compares the model's distance from parity to the labor data's distance from parity, expressed as a percentage increase or decrease in bias.","core_discovery":"The central claim is that LMMs systematically reproduce and amplify occupational gender stereotypes: compared with LLM-retrieved BLS/ILO labor baselines, models over-represent men in male-dominated professions by about 12 percentage points (US data) while under-representing men in female-dominated professions by about 5.6 points. For professions without a strong gender association, the default person generated is male 68.3% of the time, significantly above 50% parity. The paper also claims the bias is model-specific: a two-way ANOVA shows a significant interaction between model and profession stereotype, and binomial tests find only one model (Nova Canvas) whose male- and female-stereotyped","pith_inferences":["Going beyond the paper: the same 75-prompt protocol could be run in non-English languages with local labor data; the paper itself flags culture and language as a limitation, so this is the next testable step.","I infer that the default-male bias should show up in downstream applications like stock-photo generation or educational imagery, not just in benchmark prompts; a content analysis of such outputs would test that.","The cross-model variance could be an artifact of differing safety filters rather than training-data curation; distinguishing the two would require ablating the same model with and without post-processing, which the paper does not do.","I infer that the fairness scores can double as a regression suite: vendors could run the benchmark on every release, with the July-August 2025 snapshot as the baseline."],"forward_implications":["A standardized, API-only benchmark can rank closed and open models on the same fairness scale without needing internal model access.","Since one model already approaches parity, developer-side measures such as balanced training data, prompt rewriting, or output filtering are enough to substantially reduce stereotyping.","Gender-neutral prompts should not be assumed neutral: for most models the default person is male, so downstream applications inherit that skew.","Bias audits need a real-world baseline to distinguish reflecting society from amplifying society; the high correlation with labor data is not sufficient on its own.","The LLM-as-a-judge scoring method makes continuous regression testing of image fairness practical at scale."],"supporting_citations":[{"why":"Supplies the programmatic framework (prompt generation plus LLM-as-a-judge scoring) on which the benchmark is built.","marker":"[22]"},{"why":"Validates the LLM-as-a-judge approach the paper uses to score 965 images.","marker":"[23]"},{"why":"Documents the male-skewed occupational representation in LAION-5B, the training-data bias the paper expects models to reproduce.","marker":"[9]"},{"why":"Survey of bias and fairness in multimodal AI; frames the problem and motivates cross-model evaluation.","marker":"[4]"},{"why":"Shows earlier-generation text-to-image models amplify demographic stereotypes at scale; the result this paper extends to 13 current models.","marker":"[13]"},{"why":"Prior cross-model analysis of social bias in text-to-image generation; its lack of comparability motivates the identical-prompt design.","marker":"[11]"},{"why":"Prior analysis of societal bias in diffusion models via captions; the direct-depiction method is the contrast.","marker":"[12]"},{"why":"Evidence that gender bias varies by language and culture; supports the paper's stated limitation and future cross-cultural work.","marker":"[17]"}],"fun_headline_variants":["AI image models amplify job gender bias, default male 68%","13 AI image models: 93% men for stereotypically male jobs","Only 1 of 13 AI image models avoids gender default bias","AI image generators default to men 68% on neutral job prompts","Bias check: 13 image models pick men for 93% of male-coded jobs"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the median of statistics retrieved by four LLMs from BLS and ILO sources accurately represents real-world labor data; the paper used those medians instead of querying the official databases directly.","fun_headline_variants_meta":{"raw":{"variants":["AI image models amplify job gender bias, default male 68%","13 AI image models: 93% men for stereotypically male jobs","Only 1 of 13 AI image models avoids gender default bias","AI image generators default to men 68% on neutral job prompts","Bias check: 13 image models pick men for 93% of male-coded jobs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000824,"raw_usage":{"total_tokens":3494,"prompt_tokens":851,"completion_tokens":2643,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":2545}},"tokens_in":595,"tokens_out":2643,"duration_ms":19854,"temperature":1.0,"reasoning_tokens":2545,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T23:00:54.527817+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Pull the same 75 professions directly from Bureau of Labor Statistics and International Labour Organization tables and recompute the bias-amplification analysis; if the official medians differ from the LLM-retrieved medians by more than a few percentage points in the stereotyped categories, the paper's amplification finding would not replicate.","supporting_citations":[{"cited_title":"Policy-Grounded Safety Evaluation of 20 Large Language Models","cited_arxiv_id":"2507.14719","evidence_quote":"Supplies the programmatic framework (prompt generation plus LLM-as-a-judge scoring) on which the benchmark is built."},{"cited_title":"Auditing gender bias in laion-5b occupation representations","cited_arxiv_id":null,"evidence_quote":"Documents the male-skewed occupational representation in LAION-5B, the training-data bias the paper expects models to reproduce."},{"cited_title":"Easily accessible text-to-image generation amplifies demographic stereotypes at large scale","cited_arxiv_id":null,"evidence_quote":"Shows earlier-generation text-to-image models amplify demographic stereotypes at scale; the result this paper extends to 13 current models."},{"cited_title":"Social biases through the text-to-image generation lens, 2023","cited_arxiv_id":null,"evidence_quote":"Prior cross-model analysis of social bias in text-to-image generation; its lack of comparability motivates the identical-prompt design."},{"cited_title":"Evaluating gender bias in multilingual multimodal ai models: Insights from an indian context","cited_arxiv_id":null,"evidence_quote":"Evidence that gender bias varies by language and culture; supports the paper's stated limitation and future cross-cultural work."}],"review_version":1}