Pith. sign in

REVIEW 13 cited by

Survey of Bias In Text-to-Image Generation: Definition, Evaluation, and Mitigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01030 v3 pith:GLP42B6R submitted 2024-04-01 cs.CV cs.AIcs.CY

classification cs.CVcs.AIcs.CY
keywords biasbiasesgenderworkscurrentmitigationmodelsskintone
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent advancement of large and powerful models with Text-to-Image (T2I) generation abilities -- such as OpenAI's DALLE-3 and Google's Gemini -- enables users to generate high-quality images from textual prompts. However, it has become increasingly evident that even simple prompts could cause T2I models to exhibit conspicuous social bias in generated images. Such bias might lead to both allocational and representational harms in society, further marginalizing minority groups. Noting this problem, a large body of recent works has been dedicated to investigating different dimensions of bias in T2I systems. However, an extensive review of these studies is lacking, hindering a systematic understanding of current progress and research gaps. We present the first extensive survey on bias in T2I generative models. In this survey, we review prior studies on dimensions of bias: Gender, Skintone, and Geo-Culture. Specifically, we discuss how these works define, evaluate, and mitigate different aspects of bias. We found that: (1) while gender and skintone biases are widely studied, geo-cultural bias remains under-explored; (2) most works on gender and skintone bias investigated occupational association, while other aspects are less frequently studied; (3) almost all gender bias works overlook non-binary identities in their studies; (4) evaluation datasets and metrics are scattered, with no unified framework for measuring biases; and (5) current mitigation methods fail to resolve biases comprehensively. Based on current limitations, we point out future research directions that contribute to human-centric definitions, evaluations, and mitigation of biases. We hope to highlight the importance of studying biases in T2I systems, as well as encourage future efforts to holistically understand and tackle biases, building fair and trustworthy T2I technologies for everyone.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Bias in MM-DiTs is mediated by sparse stage-wise semantic binding hubs, and sparse inference-time steering at those hubs mitigates gender, race, and intersectional stereotypes with low overhead.

  2. Investigating Social Bias in Narrative Image Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Across six text-to-image models, stereotyped outputs rise from 25.9% of single photos to about 36% of storyboards and 44% of four-panel comics, with bias expressed through plot, character placement, and dialogue.

  3. EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

    cs.MM 2026-08 conditional novelty 6.0 of 10

    In emergency scenes, text-to-image models skew who appears and who helps, favoring men, middle-aged, and lighter-skinned people, and a soft-token tweak reduces the gap.

  4. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  5. GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    GASS enhances fixed-prompt diversity in T2I models by expanding CLIP embedding spread along the text direction and a computed orthogonal background direction.

  6. Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Prototypicality bias: common text-to-image metrics systematically prefer plausible-but-wrong images over correct non-prototypical ones; PROTOSCORE mitigates but does not eliminate the failure.

  7. Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

    cs.HC 2025-10 conditional novelty 6.0 of 10

    Four major LLMs generate occupational personas whose race and gender distributions deviate from U.S. workforce data in shared, patterned ways: White and Black workers are underrepresented while Hispanic and Asian work...

  8. Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models

    cs.CV 2025-10 conditional novelty 6.0 of 10

    When countries are not named, image models default to US-like modern styles, and iterative image editing erodes cultural fidelity that CLIPScore misses but human raters and a culture-aware VQA metric catch.

  9. From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.

  10. Adultification Bias in LLMs and Text-to-Image Models

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.

  11. Multi-Group Proportional Representation for Text-to-Image Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors apply the MPR metric (an integral probability metric) to text-to-image generation, derive tractable forms for linear and decision-tree function classes, and use it as a fine-tuning objective that reduces i...

  12. Can we Debias Social Stereotypes in AI-Generated Images? Examining Text-to-Image Outputs and User Perceptions

    cs.HC 2025-05 conditional novelty 5.0 of 10

    A rubric-based Social Stereotype Index shows prompt refinement lowers measured stereotypes in text-to-image outputs, but users often still prefer the stereotypical versions.

  13. Hidden Bias in the Machine: Stereotypes in Text-to-Image Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Text-to-image models reproduce and amplify stereotypes about gender, race, age, and body type across a broad set of everyday prompt categories.

Pith tools