Pith. sign in

REVIEW 15 cited by

Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01030 v3 pith:GLP42B6R submitted 2024-04-01 cs.CV cs.AIcs.CY

Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation

classification cs.CV cs.AIcs.CY
keywords biasbiasesgenderworkscurrentmitigationmodelsskintone
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The recent advancement of large and powerful models with Text-to-Image (T2I) generation abilities -- such as OpenAI's DALLE-3 and Google's Gemini -- enables users to generate high-quality images from textual prompts. However, it has become increasingly evident that even simple prompts could cause T2I models to exhibit conspicuous social bias in generated images. Such bias might lead to both allocational and representational harms in society, further marginalizing minority groups. Noting this problem, a large body of recent works has been dedicated to investigating different dimensions of bias in T2I systems. However, an extensive review of these studies is lacking, hindering a systematic understanding of current progress and research gaps. We present the first extensive survey on bias in T2I generative models. In this survey, we review prior studies on dimensions of bias: Gender, Skintone, and Geo-Culture. Specifically, we discuss how these works define, evaluate, and mitigate different aspects of bias. We found that: (1) while gender and skintone biases are widely studied, geo-cultural bias remains under-explored; (2) most works on gender and skintone bias investigated occupational association, while other aspects are less frequently studied; (3) almost all gender bias works overlook non-binary identities in their studies; (4) evaluation datasets and metrics are scattered, with no unified framework for measuring biases; and (5) current mitigation methods fail to resolve biases comprehensively. Based on current limitations, we point out future research directions that contribute to human-centric definitions, evaluations, and mitigation of biases. We hope to highlight the importance of studying biases in T2I systems, as well as encourage future efforts to holistically understand and tackle biases, building fair and trustworthy T2I technologies for everyone.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers

    cs.CV 2026-07 conditional novelty 7.0

    Bias in MM-DiTs is mediated by sparse stage-wise semantic binding hubs, and sparse inference-time steering at those hubs mitigates gender, race, and intersectional stereotypes with low overhead.

  2. Gender Artifacts from Art History to Text-to-Image Generation

    cs.CV 2026-06 unverdicted novelty 7.0

    Introduces the StyleGender dataset and PixelSGA/MaskSGA metrics showing that text-to-image models amplify gender artifacts present in artistic styles beyond historical baselines.

  3. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  4. BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models

    cs.CV 2026-06 unverdicted novelty 6.0

    BAFIS supplies a new dataset and human-feedback framework demonstrating systematic gender and ethnicity biases in occupational image generation by five text-to-image models, with partial alignment between automated me...

  5. Embedding Arithmetic: A Lightweight, Tuning-Free Framework for Post-hoc Bias Mitigation in Text-to-Image Models

    cs.CV 2026-04 unverdicted novelty 6.0

    Embedding Arithmetic performs vector operations in the embedding space of T2I models to mitigate bias at inference time, outperforming baselines on diversity while preserving coherence via a new Concept Coherence Score.

  6. BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models

    cs.CY 2026-04 conditional novelty 6.0

    BiasIG is a multi-dimensional benchmark for social biases in T2I models that shows debiasing interventions frequently cause confounding discrimination effects.

  7. Controllable Image Generation with Composed Parallel Token Prediction

    cs.LG 2026-04 unverdicted novelty 6.0

    A new formulation for composing discrete generative processes enables precise control over novel condition combinations in image generation, cutting error rates by 63% and speeding up inference.

  8. GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

    cs.CV 2026-02 conditional novelty 6.0

    GASS enhances fixed-prompt diversity in T2I models by expanding CLIP embedding spread along the text direction and a computed orthogonal background direction.

  9. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

  10. Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

    cs.HC 2025-10 conditional novelty 6.0

    Four major LLMs generate occupational personas whose race and gender distributions deviate from U.S. workforce data in shared, patterned ways: White and Black workers are underrepresented while Hispanic and Asian work...

  11. Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

    cs.CV 2025-10 conditional novelty 6.0

    When countries are not named, image models default to US-like modern styles, and iterative image editing erodes cultural fidelity that CLIPScore misses but human raters and a culture-aware VQA metric catch.

  12. VideoPhy: Evaluating Physical Commonsense for Video Generation

    cs.CV 2024-06 conditional novelty 6.0

    VideoPhy benchmark shows state-of-the-art text-to-video models follow physical commonsense and text prompts in only 39.6% of cases for the best model.

  13. Controllable Image Generation with Composed Parallel Token Prediction

    cs.CV 2024-05 unverdicted novelty 6.0

    A derived formulation for composing discrete probabilistic generative processes enables novel condition combinations in image generation, yielding 63.4% relative error reduction and FID gains on CLEVR and FFHQ datasets.

  14. Who Defines Fairness? Target-Based Prompting for Demographic Representation in Generative Models

    cs.AI 2026-04 unverdicted novelty 5.0

    Target-based prompting lets users define fairness distributions for skin tones in generative AI, shifting outputs closer to chosen targets across 36 tested prompts for occupations and contexts.

  15. Beyond Categories of Caste: Examining Caste Bias and Morality in Text-to-Image AI Models

    cs.CY 2026-04 unverdicted novelty 4.0

    The paper reframes caste as relational rather than categorical and combines algorithmic audit with critical discourse analysis to examine nuanced caste biases in T2I models while proposing an anti-caste framework for ...