Pith. sign in

REVIEW 3 major objections 5 minor 95 references

Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Existing automated gender-bias detectors for text-to-image models do not reproduce human-labeled bias, with one overestimating it by 26.95%; a combined face-detection and CLIP detector closes the gap.

desk verdict An overdue head-to-head validation of gender-bias detectors for T2I models; the recall insight is real and useful, but the quantitative claims need error bars and the proposed enhancement is evaluated too optimistically. read the letter →

arxiv 2501.15775 v2 pith:XBJN6FBP submitted 2025-01-27 cs.CV cs.SE

classification cs.CVcs.SE
keywords AItestingtext-to-imagegenerationgenderbiasfairnessdetectorsCLIPlow-qualityimagefilteringhuman-labeleddataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests whether existing automated tools for measuring gender bias in text-to-image models can be trusted. It builds a human-labeled dataset of 6,000 images from Stable Diffusion XL, Stable Diffusion 3, and Dreamlike Photoreal 2.0, and finds that all three models lean male, with SDXL the most biased. It then runs seven published detectors against these human labels. None of the detectors reproduces the actual bias level: CLIP-Prob overestimates bias by as much as 26.95% on Dreamlike, and several detectors with high classification accuracy fail because they cannot properly filter low-quality images (no face, multiple people, or no person). The paper concludes that prior automated bias measurements for text-to-image models are unreliable, and that a detector combining face detection with CLIP, called CLIP-Enhance, comes within 0.47% to 1.23% of the human-labeled bias while filtering out 82.91% of low-quality images.

What carries the argument

The paper's analytic machinery decomposes a gender-bias detector into three stages, filtering low-quality images, classifying gender, and computing bias scores, and then isolates detector error to the filtering and classification stages. The load-bearing metrics are the model bias score, the average per-prompt absolute male-female difference divided by the total, and the prompt bias score; the proposed CLIP-Enhance detector combines a face-detection model, YOLOv8-based multi-person filtering and cropping, and CLIP zero-shot gender classification. The face-detection stage is what lets CLIP-Enhance filter out 82.91% of low-quality images while keeping a recall of 97.54% on clear images, which is the property that existing vision-language-only detectors lack.

What would settle it

A re-run of the seven detectors on a fresh human-labeled sample from the same three text-to-image models, using the original authors' code or exact configurations, that shows deviations far smaller than 26.95% for CLIP-Prob would falsify the claim that the published detectors mis-measure bias, pointing instead to an implementation artifact in this study's setup.

Watch

Extended reading notes

Core claim

The central discovery is that widely used automated gender-bias detectors do not accurately capture the bias that human annotators see in text-to-image model outputs, and the mismatch is driven mostly by the filtering step rather than the gender-classification step. On a manually labeled set of 6,000 images, the ground-truth model bias scores are 0.752 for SDXL, 0.730 for SD3, and 0.631 for Dreamlike; the seven evaluated detectors deviate from these scores by up to 26.95%, and detectors with classification accuracy above 95% (CLIP-Prob, BLIP-2) still mis-measure bias because they discard clear images or keep low-quality ones. The paper identifies the cause and a remedy: a face-detection model filters images without clear faces, YOLOv8 removes and crops multi-person images, and CLIP then classifies gender. This CLIP-Enhance pipeline reports model-bias scores within 0.47% to 1.23% of the human labels and the lowest prompt-level error across all three models.

Load-bearing premise

The study assumes that the seven detector implementations it builds, for example CLIP-Prob with a face detector and a 90% confidence cutoff, faithfully match the detectors as originally proposed, so that the measured deviations reflect the original tools rather than artifacts of this study's reimplementation.

Editorial extensions

If this is right

  • Published bias numbers for text-to-image models that rely on CLIP, CLIP-Prob, CLIP-Uncertain, or BLIP-2 are likely to overestimate or underestimate the true bias, so model rankings based on those detectors may be wrong; the paper shows CLIP-Prob would rank Dreamlike as the second most biased model when it is actually the least biased.
  • A detector with very high gender-classification accuracy can still fail at bias measurement if its filtering stage has low recall, meaning future bias studies should report filtering performance, not just classification accuracy.
  • Face-detection-based filtering combined with a vision-language classifier is a reliable recipe: MiVOLO, FairFace, and CLIP-Enhance all stay within a few percentage points of the human-labeled bias, while detectors lacking face-based filtering do not.
  • The proposed CLIP-Enhance detector provides a concrete measuring stick for future text-to-image fairness claims, with model-bias scores within 0.47% to 1.23% and prompt-level errors of 0.065, 0.048, and 0.073 on the three tested models.
  • The findings imply that fairness testing for generative models should treat low-quality-image filtering as a first-class evaluation step, since a 12.48% average rate of low-quality images is large enough to distort any downstream gender distribution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the detector discrepancy extends beyond the three open-source models tested, earlier conclusions about which text-to-image model is most biased may need re-examination, because the bias ranking itself can flip depending on the detector used.
  • The same filtering failure likely affects bias measurements for race, age, and other demographic attributes, since any low-quality image corrupts the downstream classifier regardless of the attribute being measured.
  • A natural testable extension is to run CLIP-Enhance on closed models such as DALL-E 3 or Imagen with a human-labeled sample, to see whether the 0.47% to 1.23% accuracy holds outside the open-source model family.
  • The 12.48% low-quality-image rate may itself carry bias-relevant information: a model that generates more unreadable images for certain prompts could be hiding a stereotype rather than correcting it, so filtering these images out might understate the bias a user actually experiences.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper asks whether existing automated detectors can accurately measure gender bias in text-to-image models. It constructs a dataset of 6,000 images from SDXL, SD3, and Dreamlike Photoreal 2.0 using 100 gender-neutral prompts, manually labels each image as male, female, or low-quality (Cohen's Kappa 0.86), and compares the model bias scores and prompt bias scores computed from seven detectors from prior work with the human-labeled ground truth. The paper reports that all three models generate more male than female images, that profession prompts are the most biased category, that none of the seven detectors reproduces the human bias scores (CLIP-Prob overestimates Dreamlike's model bias score by 26.95%), and that vision-language model-based detectors are poor at filtering low-quality images. Based on these findings, the authors propose CLIP-Enhance, which combines dlib face detection, YOLOv8 multi-person filtering, and CLIP classification, and claim it achieves model-bias-score differences of only 0.47%-1.23% and filters 82.91% of low-quality images. The dataset and code are publicly available.

Significance. If the findings hold, this paper provides a useful, sobering benchmark for T2I bias testing: it would show that common automated detectors can deviate considerably from human judgments and that the filtering step is a major source of error. The strengths are the release of a human-labeled dataset with high inter-annotator agreement, the decomposition of detector errors into filtering and classification, and the reproducible study design. However, the bold quantitative claims are currently not backed by statistical uncertainty analysis, and the fidelity of the seven detector implementations to the originally proposed methods is not established; both are fixable and are addressed in the major comments.

major comments (3)
  1. [§3.1 and Appendix A] The RQ2 headline deviations, including the 26.95% overestimate for CLIP-Prob on Dreamlike (Table 1), are meaningful only if the seven detector implementations faithfully correspond to the original proposals. The manuscript provides no fidelity check: it does not compare against the original authors' code or published outputs, and for CLIP-Prob it is not demonstrated that Seshadri et al. [73] use MediaPipe face detection with a 90% CLIP-similarity cutoff as implemented here (Appendix A). If the face detector or threshold differs, the reported 80.9% filtering of clear images and the resulting bias deviation could be artifacts of this study's setup rather than properties of the original detectors. Please add a fidelity analysis (e.g., reproducing a sample of the original papers' reported results or a sensitivity analysis over face detectors and confidence thresholds) and report the provenance of each implementation.
  2. [§4.2, Tables 1 and 2] All detector comparisons are point estimates with no confidence intervals or significance tests. Because each prompt-model cell contains only 20 images and T2I generation is stochastic, the differences among the better detectors (e.g., CLIP-Enhance 0.53% vs. MiVOLO 0.93% for SDXL in Table 1) may lie within sampling noise, and the same may hold for the prompt bias score differences in Table 2. Please provide standard errors, bootstrap confidence intervals, or a per-prompt paired test to support the ranking 'CLIP-Enhance is the most accurate detector' and the claim that existing detectors are inaccurate.
  3. [§5.1] CLIP-Enhance is designed by inspecting the failure modes of the same dataset on which it is evaluated, and its key free parameter—the 50% second-person bounding-box-area ratio for multi-person filtering—is hand-chosen without any held-out data or sensitivity analysis. Measuring the detector on the data used for its design can overstate the reported 0.47%-1.23% model-bias-score accuracy and 82.91% filter rate. Please evaluate CLIP-Enhance on a held-out set of prompts or models, or at minimum present a sensitivity analysis over the 50% threshold.
minor comments (5)
  1. [Appendix A and B] The CLIP-Uncertain prompt contains a typo: 'a phot of a person' should be 'a photo of a person'.
  2. [§3.3 and Table 5] The text reports average male/female percentages of 63.57%/24.18%, while Table 5 gives 63.50%/24.02%; please reconcile the numbers.
  3. [§4.2] The text gives FairFace prompt bias score differences as 0.093, 0.055, and 0.109, but Table 2 reports 0.092, 0.053, and 0.108; the values are inconsistent.
  4. [§1 and §4.2] The statement that CLIP's deviation is 'seven times more' than FairFace holds only for SD3 (3.97%/0.55% = 7.2); on Dreamlike CLIP is actually closer than FairFace (4.91% vs 5.55%). Please qualify the claim per model.
  5. [§3.3] The image generation step does not report random seeds or sampling parameters; adding these details would improve reproducibility.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: the RQ2 comparison is anchored to independent human labels, and the paper's self-citations are non-load-bearing; the in-sample design of CLIP-Enhance and the unvalidated CLIP-Prob reimplementation are validity caveats, not circular reductions.

full rationale

The paper's derivation chain is: (i) build a 6,000-image dataset from three T2I models; (ii) obtain ground-truth gender labels from two human annotators (Cohen's kappa 0.86, Section 3.3); (iii) run seven detectors on the same images (Appendix A); (iv) compute model and prompt bias scores (Eqs. 1 and 2) from each labeling source; and (v) measure the discrepancy via Eq. 3. At no point is any detector's output defined in terms of the human labels, nor is the human ground truth derived from a detector. The central RQ2 finding — 'None of the detectors can accurately capture the gender bias in T2I models, with some overestimating bias by as much as 26.95%' — is therefore a comparison against an independent external benchmark, not a self-fulfilling construction. Eqs. 1-3 are definitional metrics but they do not encode the result; the 26.95% figure is an empirical consequence of which images CLIP-Prob's 90%-confidence filter discards. The self-citations present (e.g., Lyu et al. 2024 and 2023 for kappa-threshold and metric conventions; Yang et al./Lo-group fairness papers in Related Work) are non-load-bearing: none is invoked to justify a detector's design or to rule out alternatives. The only circularity-adjacent step is Section 5.1's CLIP-Enhance, whose components (dlib face detection, YOLOv8 multi-person filter with a 50% area threshold, cropping) are 'based on the empirical findings' obtained on the very same 6,000 images, and which is then evaluated on that same set with claimed deviations of 0.47%-1.23%. This is in-sample model selection rather than a derivation that reduces to its inputs: the 50% threshold is an un-fitted hand choice, and the reported accuracy is a measurement that could plausibly have been worse, so the result is not forced by construction. It does, however, mean the 0.47%-1.23% figure is not an out-of-sample validation. A separate caveat, outside circularity: the paper does not verify that its Appendix A reimplementation of CLIP-Prob (MediaPipe face detector plus 90% threshold) matches Seshadri et al.'s original, so the 26.95% overestimate may be an artifact of this study's configuration; that is a construct-validity threat, not circularity. The paper's own stated limitations (Section 7: binary gender framing, prompt/T2I model coverage) likewise bound generality without introducing a circular step.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No new physical or conceptual entities are postulated. The CLIP-Enhance detector is a combination of existing components (dlib, YOLOv8, CLIP) with one hand-picked threshold.

free parameters (1)
  • Multi-person filtering ratio (CLIP-Enhance) = second-largest YOLOv8 bounding box area > 50% of largest triggers filtering
    Introduced by the authors in Section 5.1; no cross-validation or justification; it directly affects CLIP-Enhance's reported filter rate (82.91%) and bias accuracy (0.47%-1.23%).
assumptions (4)
  • domain assumption Gender is treated as binary (male/female); non-binary identities are excluded from measurement
    Stated in Section 3.2 and Ethical Considerations; the detectors and bias scores only support binary categories, so the measured 'actual bias' does not cover non-binary people.
  • domain assumption A fair T2I model should generate equal numbers of male and female images for gender-neutral prompts
    Stated in Section 7 (Construct Threats) and used in Eq. 1 and Eq. 2; the bias score interprets any deviation from 50/50 as bias.
  • domain assumption Human labels from two annotators are the ground truth for both gender and image quality
    Section 3.3; Cohen's Kappa of 0.86 supports reliability, but the entire comparison treats annotator judgment as error-free.
  • domain assumption Gender bias is computed only on 'clear' images; low-quality images are excluded from the denominator
    Eq. 2 defines Nclear; Section 3.3 excludes low-quality images. This choice can understate bias if a model generates many low-quality images for one gender.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?." pith.science (2026). https://pith.science/paper/XBJN6FBP

@misc{pith2026250115775,
  author       = {Pith},
  title        = {Pith review of: Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Image Models?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XBJN6FBP}},
  note         = {Machine review of arXiv:2501.15775}
}
read the original abstract

Text-to-Image (T2I) models have recently gained significant attention due to their ability to generate high-quality images and are consequently used in a wide range of applications. However, there are concerns about the gender bias of these models. Previous studies have shown that T2I models can perpetuate or even amplify gender stereotypes when provided with neutral text prompts. Researchers have proposed automated gender bias uncovering detectors for T2I models, but a crucial gap exists: no existing work comprehensively compares the various detectors and understands how the gender bias detected by them deviates from the actual situation. This study addresses this gap by validating previous gender bias detectors using a manually labeled dataset and comparing how the bias identified by various detectors deviates from the actual bias in T2I models, as verified by manual confirmation. We create a dataset consisting of 6,000 images generated from three cutting-edge T2I models: Stable Diffusion XL, Stable Diffusion 3, and Dreamlike Photoreal 2.0. During the human-labeling process, we find that all three T2I models generate a portion (12.48% on average) of low-quality images (e.g., generate images with no face present), where human annotators cannot determine the gender of the person. Our analysis reveals that all three T2I models show a preference for generating male images, with SDXL being the most biased. Additionally, images generated using prompts containing professional descriptions (e.g., lawyer or doctor) show the most bias. We evaluate seven gender bias detectors and find that none fully capture the actual level of bias in T2I models, with some detectors overestimating bias by up to 26.95%. We further investigate the causes of inaccurate estimations, highlighting the limitations of detectors in dealing with low-quality images. Based on our findings, we propose an enhanced detector...

Figures

Figures reproduced from arXiv: 2501.15775 by the authors.

Figure 1
Figure 1. Examples of Low-Quality Images. Prompts Used [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Gender Bias Evaluation Process. Starting with the generated images from T2I models, the evaluation process involves [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Distribution of 300 prompt bias score outputs. The x-axis represents the prompt bias score. The y-axis represents the [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 43 canonical work pages

  1. [73]

    Preethi Seshadri, Sameer Singh, and Yanai Elazar. 2023. The bias amplification paradox in text-to-image generation. arXiv preprint arXiv:2308.00755 (2023)

  2. [1]

    Stability AI. 2024. Stable Diffusion 3 Medium. https://huggingface.co/stabilityai/ stable-diffusion-3-medium

  3. [2]

    Stability AI. 2024. Stable Diffusion XL Base 1.0. https://huggingface.co/stabilityai/ stable-diffusion-xl-base-1.0

  4. [3]

    Ana Kessler. 2023. Breakdown of Coca-Cola Commercial Made with Stable Diffusion Revealed. https://80.lv/articles/breakdown-of-coca-cola-commercial- made-with-stable-diffusion-revealed/

  5. [4]

    T2IReplication Anonymous. 2024. T2IReplication-ISSTA25. (10 2024). https: //doi.org/10.6084/m9.figshare.27377649.v1

  6. [5]

    Muhammad Hilmi Asyrofi, Zhou Yang, Imam Nur Bani Yusuf, Hong Jin Kang, Ferdian Thung, and David Lo. 2021. Biasfinder: Metamorphic test generation to uncover bias for sentiment analysis systems. IEEE Transactions on Software Engineering 48, 12 (2021), 5087–5101

  7. [6]

    Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. 2022. How well can Text-to-Image Generative Models understand Ethical Natural Language Interventions?. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 1358–1370

  8. [7]

    Solon Barocas and Andrew D Selbst. 2016. Big data’s disparate impact. Calif. L. Rev. 104 (2016), 671

Show all 95 references
  1. [8]

    Richard Berk, Hoda Heidari, Shahin Jabbari, Michael Kearns, and Aaron Roth

  2. [9]

    James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al . 2023. Improving im- age generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf 2, 3 (2023), 8

  3. [10]

    Yuriy Brun and Alexandra Meliou. 2018. Software fairness. In Proceedings of the 2018 26th ACM joint meeting on european software engineering conference and symposium on the foundations of software engineering . 754–759

  4. [11]

    Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accu- racy disparities in commercial gender classification. In Conference on fairness, accountability and transparency. PMLR, 77–91

  5. [12]

    Alessandro Castelnovo, Riccardo Crupi, Greta Greco, Daniele Regoli, Ilaria Giuseppina Penco, and Andrea Claudio Cosentini. 2022. A clarification of the nuances in the fairness metrics landscape. Scientific Reports 12, 1 (2022), 4209

  6. [13]

    Zhang, Max Hort, Mark Harman, and Federica Sarro

    Zhenpeng Chen, Jie M. Zhang, Max Hort, Mark Harman, and Federica Sarro

  7. [14]

    Zhenpeng Chen, Jie M Zhang, Federica Sarro, and Mark Harman. 2023. A com- prehensive empirical study of bias mitigation methods for machine learning classifiers. ACM transactions on software engineering and methodology 32, 4 (2023), 1–30

  8. [15]

    Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3043–3054

  9. [16]

    Tom Davenport. 2023. Cuebric: Generative AI Comes to Hollywood. (2023). https://www.forbes.com/sites/tomdavenport/2023/03/13/cuebric- generative-ai-comes-to-hollywood/?sh=19b07abb174b

  10. [17]

    Mark Díaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle

  11. [18]

    Dlib. 2017. High quality face recognition. http://dlib.net/dnn_face_recognition_ ex.cpp.html. Accessed: 2024-07-02

  12. [19]

    dreamlike art. 2023. Dreamlike Photoreal 2.0. https://huggingface.co/dreamlike- art/dreamlike-photoreal-2.0

  13. [20]

    Eran Eidinger, Roee Enbar, and Tal Hassner. 2014. Age and gender estimation of unfiltered faces. IEEE Transactions on information forensics and security 9, 12 (2014), 2170–2179

  14. [21]

    Piero Esposito, Parmida Atighehchian, Anastasis Germanidis, and Deepti Ghadi- yaram. 2023. Mitigating stereotypical biases in text to image generative systems. arXiv preprint arXiv:2310.06904 (2023)

  15. [22]

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. 2024. Scaling rectified flow transformers for high-resolution image synthesis. arXiv preprint arXiv:2403.03206 (2024)

  16. [23]

    Fairness analysis

    Anthony Finkelstein, Mark Harman, S Afshin Mansouri, Jian Ren, and Yuanyuan Zhang. 2008. “Fairness analysis” in requirements assignments. In 2008 16th IEEE International Requirements Engineering Conference . IEEE, 115–124

  17. [24]

    Kathleen C Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023. Diversity is not a one-way street: Pilot study on ethical interventions for racial bias in text-to-image systems. ICCV, accepted (2023)

  18. [25]

    Kathleen C Fraser, Svetlana Kiritchenko, and Isar Nejadgholi. 2023. A friendly face: Do text-to-image systems rely on stereotypes when the input is under- specified? arXiv preprint arXiv:2302.07159 (2023)

  19. [26]

    Felix Friedrich, Manuel Brack, Lukas Struppek, Dominik Hintersdorf, Patrick Schramowski, Sasha Luccioni, and Kristian Kersting. 2023. Fair diffusion: Instruct- ing text-to-image generation models on fairness. arXiv preprint arXiv:2302.10893 (2023)

  20. [27]

    Felix Friedrich, Katharina Hämmerl, Patrick Schramowski, Jindrich Libovicky, Kristian Kersting, and Alexander Fraser. 2024. Multilingual Text-to-Image Gen- eration Magnifies Gender Stereotypes and Prompt Engineering May Not Help You. arXiv preprint arXiv:2401.16092 (2024)

  21. [28]

    gofundme. 2023. GoFundMe | Help Changes Everything. https://www.youtube. com/watch?v=NqdC0WX-f6o

  22. [29]

    Google. 2024. MediaPipe Face Detector - Python API. https://ai.google.dev/ edge/mediapipe/solutions/vision/face_detector/python Accessed: April 10, 2025

  23. [30]

    GOP. 2023. Beat Biden. https://www.youtube.com/watch?v=kLMMxgtxQ1Y&t= 32s

  24. [31]

    Nina Grgic-Hlaca, Muhammad Bilal Zafar, Krishna P Gummadi, and Adrian Weller. 2016. The case for process fairness in learning: Feature selection for fair decision making. In NIPS symposium on machine learning and the law , Vol. 1. Barcelona, Spain, 11

  25. [32]

    Huizhong Guo, Jinfeng Li, Jingyi Wang, Xiangyu Liu, Dongxia Wang, Zehong Hu, Rong Zhang, and Hui Xue. 2023. FairRec: Fairness testing for deep recommender systems. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis. 310–321

  26. [33]

    Moritz Hardt, Eric Price, and Nati Srebro. 2016. Equality of opportunity in supervised learning. Advances in neural information processing systems 29 (2016)

  27. [34]

    Deborah Hellman. 2020. Measuring algorithmic fairness. Virginia Law Review 106, 4 (2020), 811–866

  28. [35]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840–6851

  29. [36]

    Glenn Jocher, Ayush Chaurasia, and Jing Qiu. 2023. Ultralytics YOLO. https: //github.com/ultralytics/ultralytics

  30. [37]

    Kimmo Karkkainen and Jungseock Joo. 2021. Fairface: Face attribute dataset for balanced race, gender, and age for bias measurement and mitigation. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1548–1558

  31. [38]

    Os Keyes, Chandler May, and Annabelle Carrell. 2021. You keep using that word: Ways of thinking about gender in computing research. Proceedings of the ACM on human-computer interaction 5, CSCW1 (2021), 1–23

  32. [39]

    Eunji Kim, Siwon Kim, Chaehun Shin, and Sungroh Yoon. 2023. De-stereotyping text-to-image models through prompt tuning. (2023)

  33. [40]

    Maksim Kuprashevich and Irina Tolstykh. 2023. Mivolo: Multi-input transformer for age and gender estimation. arXiv preprint arXiv:2307.04616 (2023)

  34. [41]

    Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017. Counterfac- tual fairness. Advances in neural information processing systems 30 (2017)

  35. [42]

    J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159–174

  36. [43]

    Julia Kaiwen Lau, Kelvin Kai Wen Kong, Julian Hao Yong, Per Hoong Tan, Zhou Yang, Zi Qian Yong, Joshua Chern Wey Low, Chun Yong Chong, Mei Kuan Lim, and David Lo. 2023. Synthesizing Speech Test Cases with Text-to-Speech? An Empirical Study on the False Alarms in Automated Spee...

  37. [44]

    Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Teufel, Marco Bellagente, et al. 2024. Holistic evaluation of text-to-image models. Advances in Neural Information Processing Systems 36 (2024)

  38. [45]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning . PMLR, 19730–19742

  39. [46]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900

  40. [47]

    Alexander Lin, Lucas Monteiro Paes, Sree Harsha Tanneru, Suraj Srinivas, and Himabindu Lakkaraju. 2023. Word-Level Explanations for Analyzing Bias in Text-to-Image Models. arXiv preprint arXiv:2306.05500 (2023)

  41. [48]

    Yiming Lin, Jie Shen, Yujiang Wang, and Maja Pantic. 2022. Fp-age: Leveraging face parsing attention for facial age estimation in the wild. IEEE Transactions on Image Processing (2022). 9 MM ’25, October 27–31, 2025, Dublin, Ireland Lyu et al

  42. [49]

    Haochen Liu, Jamell Dacon, Wenqi Fan, Hui Liu, Zitao Liu, and Jiliang Tang

  43. [50]

    Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2024. Stable bias: Evaluating societal representations in diffusion models. Advances in Neural Information Processing Systems 36 (2024)

  44. [51]

    Yunbo Lyu, Hong Jin Kang, Ratnadira Widyasari, Julia Lawall, and David Lo

  45. [52]

    Le, Ming Li, and David Lo

    Yunbo Lyu, Thanh Le-Cong, Hong Jin Kang, Ratnadira Widyasari, Zhipeng Zhao, Xuan-Bach D. Le, Ming Li, and David Lo. 2023. Chronos: Time-Aware Zero-Shot Identification of Libraries from Vulnerability Reports. In Proceedings of the 45th International Conference on Software Engin...

  46. [53]

    Elman Mansimov, Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov. 2015. Generating images from captions with attention. arXiv preprint arXiv:1511.02793 (2015)

  47. [54]

    Mary L McHugh. 2012. Interrater reliability: the kappa statistic. Biochemia medica 22, 3 (2012), 276–282

  48. [55]

    Megvii. 2024. Face++ Cognitive Services. https://www.faceplusplus.com/

  49. [56]

    IEEE Trans

    Evaluating SZZ Implementations: An Empirical Study on the Linux Kernel. IEEE Trans. Softw. Eng. 50, 9 (Sept. 2024), 2219–2239. https://doi.org/10.1109/ TSE.2024.3406718

  50. [57]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. A survey on bias and fairness in machine learning. ACM com- puting surveys (CSUR) 54, 6 (2021), 1–35

  51. [58]

    Shira Mitchell, Eric Potash, Solon Barocas, Alexander D’Amour, and Kristian Lum. 2021. Algorithmic fairness: Choices, assumptions, and definitions. Annual review of statistics and its application 8, 1 (2021), 141–163

  52. [59]

    Ranjita Naik and Besmira Nushi. 2023. Social biases through the text-to-image generation lens. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. 786–808

  53. [60]

    Author’s Full Name. 2024. 90% of Online Content Could Be Gen- erated by AI by 2025, Expert Says. Yahoo Finance (2024). https: //finance.yahoo.com/news/90-of-online-content-could-be-generated-by- ai-by-2025-expert-says-201023872.html Accessed: 2024-08-22

  54. [61]

    Ninareh Mehrabi, Thamme Gowda, Fred Morstatter, Nanyun Peng, and Aram Galstyan. 2020. Man is to person as woman is to location: Measuring gender bias in named entity recognition. In Proceedings of the 31st ACM conference on Hypertext and Social Media . 231–232

  55. [62]

    Evgeny Obedkov. 2023. How AI-Assisted RPG Tales of Syn Utilizes Stable Diffu- sion and ChatGPT to Create Assets and Dialogues. https://gameworldobserver. com/2023/03/06/tales-of-syn-ai-rpg-stable-diffusion-chatgpt-game Accessed: 2024-08-22

  56. [63]

    OpenAI. 2022. DALL-E 2. https://openai.com/index/dall-e-2/

  57. [64]

    OpenAI. 2023. Dall ·E 3 System Card. https://openai.com/research/dall-e-3- system-card

  58. [65]

    Hadas Orgad, Bahjat Kawar, and Yonatan Belinkov. 2023. Editing implicit as- sumptions in text-to-image diffusion models. arXiv preprint arXiv:2303.08084 (2023)

  59. [66]

    Leonardo Nicoletti and Dina Bass. 2023. Humans are biased. Generative AI is even worse. https://www.bloomberg.com/graphics/2023-generative-ai-bias/

  60. [67]

    Danny Postma. [n. d.]. AI Modelling Agency — Deep Agency. https://www. deepagency.com/

  61. [68]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  62. [69]

    Sai Sathiesh Rajan, Sakshi Udeshi, and Sudipta Chattopadhyay. 2022. Aequevox: Automated fairness testing of speech recognition systems. InInternational Confer- ence on Fundamental Approaches to Software Engineering . Springer International Publishing Cham, 245–267

  63. [70]

    Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. 2016. Generative adversarial text to image synthesis. In Inter- national conference on machine learning . PMLR, 1060–1069

  64. [71]

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023)

  65. [72]

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Infor...

  66. [74]

    Ezekiel Soremekun, Sakshi Udeshi, and Sudipta Chattopadhyay. 2022. Astraea: Grammar-based fairness testing. IEEE Transactions on Software Engineering 48, 12 (2022), 5188–5211

  67. [75]

    Lin Sze Khoo, Jia Qi Bay, Ming Lee Kimberly Yap, Mei Kuan Lim, Chun Yong Chong, Zhou Yang, and David Lo. 2023. Exploring and Repairing Gen- der Fairness Violations in Word Embedding-based Sentiment Analysis Model through Adversarial Patches. In 2023 IEEE International Conferen...

  68. [76]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695

  69. [77]

    Yixin Wan, Arjun Subramonian, Anaelia Ovalle, Zongyu Lin, Ashima Suvarna, Christina Chance, Hritik Bansal, Rebecca Pattichis, and Kai-Wei Chang. 2024. Sur- vey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation. arXiv preprint arXiv:2404.01030 (2024)

  70. [78]

    Wenxuan Wang, Haonan Bai, Jen-tse Huang, Yuxuan Wan, Youliang Yuan, Haoyi Qiu, Nanyun Peng, and Michael R Lyu. 2024. New Job, New Gender? Measuring the Social Bias in Image Generation Models. arXiv preprint arXiv:2401.00763 (2024)

  71. [79]

    Zichong Wang, Yang Zhou, Meikang Qiu, Israat Haque, Laura Brown, Yi He, Jianwu Wang, David Lo, and Wenbin Zhang. 2023. Towards fair machine learn- ing software: Understanding and addressing model bias through counterfactual thinking. arXiv preprint arXiv:2302.08018 (2023)

  72. [80]

    Yisong Xiao, Aishan Liu, Tianlin Li, and Xianglong Liu. 2023. Latent imitator: Generating natural individual discriminatory instances for black-box fairness testing. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis . 829–841

  73. [81]

    Yixin Wan and Kai-Wei Chang. 2024. The Male CEO and the Female Assistant: Probing Gender Biases in Text-To-Image Models Through Paired Stereotype Test. arXiv preprint arXiv:2402.11089 (2024)

  74. [82]

    Zhou Yang, Muhammad Hilmi Asyrofi, and David Lo. 2021. Biasrv: Uncovering biased sentiment predictions at runtime. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1540–1544

  75. [83]

    Zhou Yang, Harshit Jain, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2021. Biasheal: On-the-fly black-box healing of bias in sentiment analysis systems. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 644–648

  76. [84]

    Zhou Yang, Zhensu Sun, Terry Zhuo Yue, Premkumar Devanbu, and David Lo

  77. [85]

    Cheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, and Fernando De la Torre. 2023. Iti-gen: Inclusive text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 3969– 3980

  78. [86]

    Junjie Yang, Jiajun Jiang, Zeyu Sun, and Junjie Chen. 2024. A Large-Scale Empir- ical Study on Improving the Fairness of Image Classification Models. In Proceed- ings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 210–222

  79. [87]

    Lingfeng Zhang, Yueling Zhang, and Min Zhang. 2021. Efficient white-box fairness testing through gradient search. In Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis . 103–114

  80. [88]

    Peixin Zhang, Jingyi Wang, Jun Sun, and Xinyu Wang. 2021. Fairness testing of deep image classification with adequacy metrics. arXiv preprint arXiv:2111.08856 (2021)

  81. [89]

    a photo of a cat

    Zhifei Zhang, Yang Song, and Hairong Qi. 2017. Age progression/regression by conditional adversarial autoencoder. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5810–5818. 10 Do Existing Testing Tools Really Uncover Gender Bias in Text-to-Im...

  82. [90]

    arXiv preprint arXiv:2403.07506 (2024)

    Robustness, security, privacy, explainability, efficiency, and usability of large language models for code. arXiv preprint arXiv:2403.07506 (2024)

  83. [92]

    Jie M Zhang, Mark Harman, Lei Ma, and Yang Liu. 2020. Machine learning testing: Survey, landscapes and horizons. IEEE Transactions on Software Engineering 48, 1 (2020), 1–36

  84. [2018]

    In Proceedings of the 2018 chi conference on human factors in computing systems

    Addressing age-related bias in sentiment analysis. In Proceedings of the 2018 chi conference on human factors in computing systems . 1–14

  85. [2019]

    arXiv preprint arXiv:1910.10486 (2019)

    Does gender matter? towards fairness in dialogue systems. arXiv preprint arXiv:1910.10486 (2019)

  86. [2021]

    Fairness in criminal justice risk assessments: The state of the art.Sociological Methods & Research 50, 1 (2021), 3–44

  87. [2024]

    ACM Trans

    Fairness Testing: A Comprehensive Survey and Analysis of Trends. ACM Trans. Softw. Eng. Methodol. 33, 5, Article 137 (June 2024), 59 pages. https: //doi.org/10.1145/3652155

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.