Pith. sign in

REVIEW 3 major objections 5 minor 112 references

Text-to-image models reduce disability to wheelchairs and other narrow markers, generating less diverse images that match real-world high-warmth, low-competence stereotypes.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 06:07 UTC pith:WTLS2N3A

load-bearing objection First real large-scale T2I disability benchmark; wheelchair default and diversity collapse are solid, SCM claim is more circular than presented. the 3 major comments →

arxiv 2607.08515 v1 pith:WTLS2N3A submitted 2026-07-09 cs.CV

Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH

classification cs.CV
keywords text-to-imagedisability biasINCLUDE-BENCHStereotype Content Modelrepresentational harmVendi scoreintersectionalityCLIP alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper builds INCLUDE-BENCH, the first large-scale test set for disability bias in text-to-image models: 352 carefully varied prompts covering generic and specific impairments, intersectional attributes, and static-to-dynamic everyday contexts, yielding 119,680 images from 17 models. It shows that mobility and default disability prompts almost always produce wheelchair images, that disability-conditioned outputs are systematically less diverse, and that the most stereotypical pictures align most strongly with the disability text. A new SCM Score places these depictions as warmer yet less competent, exactly as sociological stereotype research predicts. Readers should care because these models already shape media and design; their compressed, demographically skewed portrayals can erase the real variety of disabled lives and reinforce low-agency assumptions.

Core claim

Across 17 text-to-image models and 119,680 images, mobility-impaired and default disability prompts predominantly yield wheelchair depictions; disability-conditioned generations exhibit reduced visual diversity; stereotypical portrayals show stronger disability-text alignment; and the introduced SCM Score demonstrates that the models systematically reproduce real-world high-warmth/low-competence associations for people with disabilities.

What carries the argument

INCLUDE-BENCH (352 prompts in four subsets varying impairment specificity, intersectional identity, and context) plus the SCM Score, which projects CLIP image embeddings onto normalized warmth and competence directions derived from positive/negative attribute texts, quantifying how closely generations match sociological stereotype content.

Load-bearing premise

The automated pipeline of person crops, vision-language captions and demographics, CLIP embeddings, and diversity scores is a valid enough stand-in for real representational harm and stereotype content even without large-scale validation by people with disabilities.

What would settle it

A large, diverse panel of people with disabilities rates the same generated images and finds high intra-group variety, low stereotyping, and competence/warmth judgments that reverse the SCM pattern while the automated metrics continue to report the opposite.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Disability must become a core, not optional, axis in any serious T2I bias benchmark.
  • Models that overweight diagnostic visual markers will keep sacrificing diversity for label alignment.
  • Adding context or intersectional cues alone does not break the alignment-diversity trade-off for physical disabilities.
  • Creative tools built on current models will continue defaulting to demographically narrow, low-agency portrayals of disability unless the training and evaluation regimes change.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same diagnostic-feature over-weighting likely explains why other underrepresented groups collapse to single visual tropes in generative models.
  • Community-validated SCM ratings could turn the metric into a practical fairness audit for future model releases.
  • Targeted diversification or filtering of wheelchair-dominant training images may be sufficient to raise diversity without collapsing prompt alignment.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces INCLUDE-BENCH, a large-scale benchmark for disability-related bias in text-to-image models. It constructs 352 prompts spanning generic and specific impairments (mobility, sensory), intersectional attributes (age, race, gender), and static/dynamic contexts grounded in WHO ICF domains, then generates 20 images per prompt from 17 T2I models (15 open, 2 closed) for a total of ~119K images. Evaluation uses person-centric crops (SAM3), clustering + captioning (Qwen3-VL), CLIPScore for disability-text alignment, Vendi Score for spectral diversity, intersectional frequency ratios, and a new multimodal SCM Score that projects image CLIP embeddings onto warmth/competence directions derived from fixed adjective sets. Main empirical claims are that mobility/default disability prompts overwhelmingly produce wheelchair imagery, disability-conditioned outputs show reduced diversity and higher alignment for stereotypical depictions, and SCM scores track real-world high-warmth/low-competence patterns for PWD.

Significance. Disability is a clear gap in T2I bias evaluation relative to gender, race, and occupation. The scale (17 models, 119K images, four controlled prompt subsets) and the multi-metric toolbox (clustering, alignment-diversity trade-off, intersectional ratios, SCM) constitute a useful community resource. Explicit credit is due for releasing a structured prompt design that isolates impairment, identity, and context, for evaluating both open and closed models under identical conditions, and for attempting to ground evaluation in the sociological Stereotype Content Model rather than ad-hoc visual checklists. If the descriptive findings hold under external validation, the benchmark can become a standard diagnostic for representational harm toward PWD.

major comments (3)
  1. §3.2.4 and finding (4): The SCM Score is defined as the projection of image CLIP embeddings onto normalized warmth/competence directions obtained from the same CLIP text encoder (adjectives adapted from Fraser et al.). CLIPScore (§3.2.2) and Vendi (§3.2.3) also live in CLIP space, and demographic labels for Table 2 come from Qwen3-VL. Because these evaluators themselves encode the social stereotypes under study, elevated SCM scores for wheelchair-heavy clusters (C1–C2) and older-white over-representation may re-project evaluator priors rather than independently confirm that T2I outputs match real-world SCM patterns. The Limitations section (§6) acknowledges the absence of PWD human annotation, yet the abstract and conclusion still treat the automated SCM numbers as evidence that “T2I models reflect real-world stereotypical associations.” An external anchor (human SCM ratings on a stratif
  2. §3.1 and axiom of visual observability: The benchmark deliberately excludes cognitive and intellectual disabilities because they are “less visually apparent.” While this is a practical design choice, it is not neutral: the paper’s own sociological framing (representativeness, diagnostic features) predicts that models will over-weight the most visible markers. Restricting the functional groups therefore risks confirming the very visual-shorthand bias the authors criticize, and weakens the claim that INCLUDE-BENCH comprehensively evaluates “disability-related bias.” At minimum the paper should quantify how much of the observed wheelchair/blindfold dominance is an artifact of this exclusion, or provide a parallel analysis for prompts that name non-visible disabilities.
  3. §4.0.1 / Table 3 and §3.2.1: Clustering (MiniBatchKMeans, k=10) and subsequent token-frequency / VQA analyses are performed on the pooled set of all models and prompts. Without model-stratified or prompt-subset-stratified cluster statistics, it is impossible to tell whether the wheelchair dominance and age/gender compression are uniform or driven by a subset of architectures (e.g., older SD variants vs. FLUX/Janus). The central claim that the patterns hold “across all models” therefore rests on an untested aggregation. Reporting per-model cluster membership or silhouette diagnostics would make the claim falsifiable.
minor comments (5)
  1. Abstract and §1: “15 open-source and 2 closed-source models” vs. later “17 state-of-the-art”; keep the count consistent and list the closed models (NanoBanana, GPT-Image-1-mini) in the abstract or early introduction.
  2. Table 2 caption and body: intersection codes (o f w, y m af, etc.) are dense; a short legend or expanded first column would improve readability.
  3. Figure 3 / 4 / 5 / 6: axis labels and Δ definitions are only partially explained in the main text; move the precise baseline subtraction formula into the caption or §4.0.2.
  4. §3.2.3: Vendi is reported primarily with CLIP embeddings; the DINO results are deferred to the supplement without a one-sentence summary of whether the diversity ranking is stable.
  5. References: several arXiv preprints and model cards lack stable version identifiers or access dates; standardize for archival purposes.

Circularity Check

0 steps flagged

Empirical benchmark measurements with external metrics; no derivation reduces to inputs by construction.

full rationale

INCLUDE-BENCH is an observational evaluation paper: prompts are constructed along impairment/context/identity axes, 119K images are generated from 17 T2I models, and properties are measured with off-the-shelf tools (SAM3 crops, Qwen3-VL captions/VQA, CLIP embeddings for CLIPScore and Vendi, MiniBatchKMeans clusters). The SCM Score (§3.2.4) is a fixed projection of image CLIP embeddings onto normalized directions δ_W and δ_C obtained from external positive/negative warmth and competence adjective lists (adapted from Fraser et al. 2021, itself grounded in the classic Fiske/Cuddy SCM literature). No parameters are fitted to the generated images and then re-used as “predictions”; no equation equates a claimed result to its own definition; no uniqueness theorem or ansatz is imported via self-citation. The authors’ own prior work is not load-bearing. Mild evaluator-bias concerns (CLIP/Qwen may themselves encode stereotypes) are validity/limitations issues (§6), not circularity of the derivation chain. Findings remain self-contained empirical measurements against the constructed prompt set.

Axiom & Free-Parameter Ledger

3 free parameters · 6 axioms · 2 invented entities

The central claims rest on standard ML evaluation tools plus domain assumptions that visual diagnostic features and SCM warmth/competence directions are valid stereotype measures for T2I outputs. Free choices include cluster count, sample size per prompt, and attribute prompt sets. Invented entities are the benchmark and the multimodal SCM Score packaging; neither is a new physical entity, but both are paper-specific constructs without external independent validation beyond the paper’s own runs.

free parameters (3)
  • n_clusters (MiniBatchKMeans)
    Fixed at 10 clusters for visual pattern discovery; choice affects which tropes dominate reported clusters.
  • images_per_prompt
    20 random-seed images per prompt; diversity and alignment statistics depend on this sample size.
  • SCM attribute prompt set size R
    Mean embeddings of positive/negative warmth and competence phrases adapted from prior work; the specific phrase inventory is a design choice that defines the SCM directions.
axioms (6)
  • domain assumption CLIP image–text cosine similarity is a valid measure of disability-label semantic alignment and of warmth/competence stereotype directions.
    Used for CLIPScore and SCM Score in §3.2.2–3.2.4; validity is assumed, not proven for disability stereotypes.
  • domain assumption Vendi Score on CLIP (and DINO) embeddings quantifies representational diversity of disability depictions.
    §3.2.3; diversity claims rest on this entropy-of-eigenvalues construction.
  • domain assumption SAM3 person detections and largest-box crops yield person-centric images free enough of background confounds for stereotype analysis.
    §3.2.1; 108 images without detections excluded.
  • domain assumption Qwen3-VL captions and VQA demographic labels are accurate enough for cluster characterization and intersectional ratios R_i,g.
    §3.2.1 and Table 2; no human gold labels reported.
  • domain assumption Stereotype Content Model warmth and competence dimensions transfer from social psychology / NLP to T2I image embeddings.
    §3.2.4; core interpretive claim that low SCM scores mean real-world stereotype reproduction.
  • ad hoc to paper Focusing on visually observable impairments (excluding cognitive/intellectual disabilities) is sufficient for a disability bias benchmark.
    §3.1 design choice; limits scope of ‘disability’ claims.
invented entities (2)
  • INCLUDE-BENCH no independent evidence
    purpose: Large-scale prompt and image suite for differential evaluation of disability bias in T2I models across contexts and intersections.
    Paper-specific benchmark construct; independent evidence would require external reuse and PWD-centered validation, not yet shown.
  • Multimodal SCM Score no independent evidence
    purpose: Project CLIP image embeddings onto normalized warmth and competence text directions to quantify stereotypical semantic alignment.
    New packaging of SCM for T2I; depends on chosen attribute texts and CLIP geometry rather than an external validated instrument for images.

pith-pipeline@v1.1.0-grok45 · 21736 in / 3554 out tokens · 42939 ms · 2026-07-10T06:07:27.580620+00:00 · methodology

0 comments
read the original abstract

Text-to-image (T2I) models have been shown to exhibit social biases. Prior work has mainly focused on gender, skin tone, and cultural representation within restricted occupational associations, and emerging benchmarks increasingly incorporate these dimensions. However, disability remains systematically underexplored. Current evaluation practices often fail to align with sociologically grounded definitions of stereotyping, limiting principled assessment of representational harms toward people with disabilities (PWD). To address this, we introduce INCLUDE-BENCH, the first large-scale benchmark for evaluating disability-related bias in T2I models. INCLUDE-BENCH comprises 119K generated images based on prompt design across multiple bias dimensions and both static and dynamic contexts. We evaluate 15 open-source and 2 closed-source models. Our key findings reveal that: (1) mobility-impaired and default disability prompts predominantly yield wheelchair depictions across all models; (2) disability-conditioned generations consistently exhibit less diversity; (3) stereotypical portrayals demonstrate stronger disability-text alignment; and (4) we introduce the Stereotype Content Model (SCM) Score, demonstrating that T2I models reflect real-world stereotypical associations.

Figures

Figures reproduced from arXiv: 2607.08515 by Albert Gatt, Judith Masthoff, Sophia Lichtenberg.

Figure 1
Figure 1. Figure 1: Sample images from INCLUDE-BENCH across seven function groups categories: Blind, Deaf, Mute, Deafblind, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: INCLUDE-BENCH pipeline: Dataset creation based on different disability groups across diverse context sets and [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: Vendi Score and CLIPScore for each Impairment [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 6
Figure 6. Figure 6: SCM Score for each Dataset and Impairment [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

112 extracted references · 112 canonical work pages · 5 internal anchors

  1. [1]

    Qwen3-vl technical report, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qi- dong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shu- tong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shix...

  2. [2]

    Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023

    Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elhoseiny. Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023. 3

  3. [3]

    How well can text-to-image generative models un- derstand ethical natural language interventions?, 2022

    Hritik Bansal, Da Yin, Masoud Monajatipoor, and Kai-Wei Chang. How well can text-to-image generative models un- derstand ethical natural language interventions?, 2022. 3

  4. [4]

    A systematic review of open datasets used in text-to-image (t2i) gen ai model safety.IEEE Access, 13:32661–32680, 2025

    Trupti Bavalatti, Osama Ahmed, Dhaval Potdar, Rakeen Rouf, Faraz Jawed, Manish Kumar Govind, and Siddharth Krishnan. A systematic review of open datasets used in text-to-image (t2i) gen ai model safety.IEEE Access, 13:32661–32680, 2025. 1

  5. [5]

    it’s complicated

    Cynthia L Bennett, Cole Gleason, Morgan Klaus Scheuer- man, Jeffrey P Bigham, Anhong Guo, and Alexandra To. “it’s complicated”: Negotiating accessibility and (mis) rep- resentation in image descriptions of race, gender, and dis- ability. InProceedings of the 2021 chi conference on human factors in computing systems, pages 1–19, 2021. 2

  6. [6]

    Toward community-led evaluations of text-to- image ai representations of disability, health, and accessi- bility

    Cynthia L Bennett, Shaun K Kane, and Christina N Har- rington. Toward community-led evaluations of text-to- image ai representations of disability, health, and accessi- bility. InProceedings of the 5th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 256–270, 2025. 1, 2, 3

  7. [7]

    Easily accessible text-to-image generation amplifies demographic stereotypes at large scale

    Federico Bianchi, Pratyusha Kalluri, Esin Durmus, Faisal Ladhak, Myra Cheng, Debora Nozza, Tatsunori Hashimoto, Dan Jurafsky, James Zou, and Aylin Caliskan. Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. InProceedings of the 2023 ACM conference on fairness, accountability, and transparency, pages 1493–1504, 2023. 2

  8. [8]

    Flux dev.https:// huggingface.co/black-forest-labs/FLUX

    Black-Forest-Labs. Flux dev.https:// huggingface.co/black-forest-labs/FLUX. 1-dev, 2024. AI text-to-image generation model. 5

  9. [9]

    Flux schnell.https: //huggingface.co/black-forest-labs/ FLUX.1-schnell, 2024

    Black-Forest-Labs. Flux schnell.https: //huggingface.co/black-forest-labs/ FLUX.1-schnell, 2024. AI text-to-image generation model. 5

  10. [10]

    Flux2 dev.https: //huggingface.co/black-forest-labs/ FLUX.2-dev, 2026

    Black-Forest-Labs. Flux2 dev.https: //huggingface.co/black-forest-labs/ FLUX.2-dev, 2026. AI text-to-image generation model. 5

  11. [11]

    Language (technology) is power: A critical survey of “bias” in nlp

    Su Lin Blodgett, Solon Barocas, Hal Daum ´e Iii, and Hanna Wallach. Language (technology) is power: A critical survey of “bias” in nlp. InProceedings of the 58th annual meet- ing of the association for computational linguistics, pages 5454–5476, 2020. 1

  12. [12]

    Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets

    Su Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim, and Hanna Wallach. Stereotyping norwegian salmon: An inventory of pitfalls in fairness benchmark datasets. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Pro- cessing (Volume 1: Long Pap...

  13. [13]

    On the Opportunities and Risks of Foundation Models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Alt- man, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021. 2 9

  14. [14]

    Stereotypes.The Quarterly journal of eco- nomics, 131(4):1753–1794, 2016

    Pedro Bordalo, Katherine Coffman, Nicola Gennaioli, and Andrei Shleifer. Stereotypes.The Quarterly journal of eco- nomics, 131(4):1753–1794, 2016. 1

  15. [15]

    Envisioning equitable speech technologies for black older adults

    Robin N Brewer, Christina Harrington, and Courtney Hel- dreth. Envisioning equitable speech technologies for black older adults. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 379– 388, 2023. 2

  16. [16]

    Hidream-i1: A high-efficient image generative foundation model with sparse diffusion transformer, 2025

    Qi Cai, Jingwen Chen, Yang Chen, Yehao Li, Fuchen Long, Yingwei Pan, Zhaofan Qiu, Yiheng Zhang, Fengbin Gao, Peihan Xu, Yimeng Wang, Kai Yu, Wenxuan Chen, Zi- wei Feng, Zijian Gong, Jianzhuang Pan, Yi Peng, Rui Tian, Siyu Wang, Bo Zhao, Ting Yao, and Tao Mei. Hidream-i1: A high-efficient image generative foundation model with sparse diffusion transformer, 2025. 5

  17. [17]

    The stereotype content model and disabilities.The Journal of Social Psychology, 163(4):480–500, 2023

    Emily Canton, Darren Hedley, and Jennifer R Spoor. The stereotype content model and disabilities.The Journal of Social Psychology, 163(4):480–500, 2023. 8

  18. [18]

    Sam 3: Segment anything with concepts, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoub- hik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, Jie Lei, Tengyu Ma, Baishan Guo, Arpit Kalla, Markus Marks, Joseph Greer, Meng Wang, Peize Sun, Roman R¨adle, Triantafyllos Afouras, Effrosyni Mavroudi, Kather- ine Xu, Tsung-Han Wu, Yu Zhou, Lil...

  19. [19]

    Disability im- pacts all of us infographic, 2025

    Centers for Disease Control and Prevention. Disability im- pacts all of us infographic, 2025. Accessed: 2025-09-12. 1

  20. [20]

    Pixart-α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast train- ing of diffusion transformer for photorealistic text-to-image synthesis, 2023. 5

  21. [21]

    Janus- pro: Unified multimodal understanding and generation with data and model scaling, 2025

    Xiaokang Chen, Zhiyu Wu, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, and Chong Ruan. Janus- pro: Unified multimodal understanding and generation with data and model scaling, 2025. 5

  22. [22]

    Tibet: Identifying and evaluating biases in text-to-image generative models, 2024

    Aditya Chinchure, Pushkar Shukla, Gaurav Bhatt, Kiri Salij, Kartik Hosanagar, Leonid Sigal, and Matthew Turk. Tibet: Identifying and evaluating biases in text-to-image generative models, 2024. 3

  23. [23]

    Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

    Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models. InProceedings of the IEEE/CVF international conference on computer vision, pages 3043– 3054, 2023. 2

  24. [24]

    Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models, 2023

    Jaemin Cho, Abhay Zala, and Mohit Bansal. Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models, 2023. 3

  25. [25]

    Mapping the margins: Inter- sectionality, identity politics, and violence against women of color

    Kimberl ´e Williams Crenshaw. Mapping the margins: Inter- sectionality, identity politics, and violence against women of color. InThe public nature of private violence, pages 93–118. Routledge, 2013. 2

  26. [26]

    The bias map: behaviors from intergroup affect and stereotypes

    Amy JC Cuddy, Susan T Fiske, and Peter Glick. The bias map: behaviors from intergroup affect and stereotypes. Journal of personality and social psychology, 92(4):631,

  27. [27]

    Garrido-Merch´an

    Adriana Fern ´andez de Caleya V ´azquez and Eduardo C. Garrido-Merch´an. A taxonomy of the biases of the images created by generative artificial intelligence, 2024. 1

  28. [28]

    OASIS Uncovers: High-Quality T2I Models, Same Old Stereotypes

    Sepehr Dehdashtian, Gautam Sreekumar, and Vishnu Naresh Boddeti. Oasis uncovers: High- quality t2i models, same old stereotypes.arXiv preprint arXiv:2501.00962, 2025. 3

  29. [29]

    Harms of gender exclusivity and challenges in non-binary repre- sentation in language technologies

    Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff Phillips, and Kai-Wei Chang. Harms of gender exclusivity and challenges in non-binary repre- sentation in language technologies. InProceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 1968–1994, 2021. 2

  30. [30]

    Addressing age-related bias in sentiment analysis

    Mark D ´ıaz, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. Addressing age-related bias in sentiment analysis. InProceedings of the 2018 chi confer- ence on human factors in computing systems, pages 1–14,

  31. [31]

    Openbias: Open-set bias detection in text-to-image generative models, 2024

    Moreno D’Inc `a, Elia Peruzzo, Massimiliano Mancini, Dejia Xu, Vidit Goel, Xingqian Xu, Zhangyang Wang, Humphrey Shi, and Nicu Sebe. Openbias: Open-set bias detection in text-to-image generative models, 2024. 3

  32. [32]

    NYU Press, 2017

    Elizabeth Ellcessor and Bill Kirkpatrick.Disability media studies. NYU Press, 2017. 2

  33. [33]

    A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition

    Susan T Fiske, Amy JC Cuddy, Peter Glick, and Jun Xu. A model of (often mixed) stereotype content: Competence and warmth respectively follow from perceived status and competition. InSocial cognition, pages 162–214. Rout- ledge, 2018. 5

  34. [34]

    Fraser, Isar Nejadgholi, and Svetlana Kir- itchenko

    Kathleen C. Fraser, Isar Nejadgholi, and Svetlana Kir- itchenko. Understanding and countering stereotypes: A computational approach to the stereotype content model. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedings of the 59th Annual Meeting of the As- sociation for Computational Linguistics and the 11th Inter- national Joint...

  35. [35]

    Association for Computational Linguistics. 5

  36. [36]

    The Vendi Score: A Diversity Evaluation Metric for Machine Learning

    Dan Friedman and Adji Bousso Dieng. The vendi score: A diversity evaluation metric for machine learning.arXiv preprint arXiv:2210.02410, 2022. 4

  37. [37]

    ” i wouldn’t say offensive but...”: Disability-centered perspectives on large language models

    Vinitha Gadiraju, Shaun Kane, Sunipa Dev, Alex Taylor, Ding Wang, Remi Denton, and Robin Brewer. ” i wouldn’t say offensive but...”: Disability-centered perspectives on large language models. InProceedings of the 2023 ACM conference on fairness, accountability, and transparency, pages 205–216, 2023. 2

  38. [38]

    Word embeddings quantify 100 years of gender and ethnic stereotypes.Proceedings of the National Academy of Sciences, 115(16):E3635–E3644, 2018

    Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. Word embeddings quantify 100 years of gender and ethnic stereotypes.Proceedings of the National Academy of Sciences, 115(16):E3635–E3644, 2018. 2

  39. [39]

    Disability and representa- tion.Pmla, 120(2):522–527, 2005

    Rosemarie Garland-Thomson. Disability and representa- tion.Pmla, 120(2):522–527, 2005. 2 10

  40. [40]

    Feminist disability stud- ies.Signs: Journal of women in Culture and Society, 30(2):1557–1587, 2005

    Rosemarie Garland-Thomson. Feminist disability stud- ies.Signs: Journal of women in Culture and Society, 30(2):1557–1587, 2005. 2

  41. [41]

    Misfits: A feminist materi- alist disability concept.Hypatia, 26(3):591–609, 2011

    Rosemarie Garland-Thomson. Misfits: A feminist materi- alist disability concept.Hypatia, 26(3):591–609, 2011. 2, 3

  42. [42]

    Integrating disability, trans- forming feminist theory

    Rosemarie Garland-Thomson. Integrating disability, trans- forming feminist theory. InFeminist theory reader, pages 181–191. Routledge, 2020. 2

  43. [43]

    The politics of staring: Visual rhetorics of disability in popular photography.Dis- ability studies: Enabling the humanities, 1, 2002

    Rosemarie Garland-Thomson et al. The politics of staring: Visual rhetorics of disability in popular photography.Dis- ability studies: Enabling the humanities, 1, 2002. 2

  44. [44]

    A large scale analysis of gender bi- ases in text-to-image generative models.arXiv preprint arXiv:2503.23398, 2025

    Leander Girrbach, Stephan Alaniz, Genevieve Smith, and Zeynep Akata. A large scale analysis of gender bi- ases in text-to-image generative models.arXiv preprint arXiv:2503.23398, 2025. 3, 4

  45. [45]

    Ai, the storyteller: Content analy- sis of disability representation in stories created for chil- dren.Interactions: Studies in Communication & Culture, 13(3):289–306, 2022

    Luda Gogolushko. Ai, the storyteller: Content analy- sis of disability representation in stories created for chil- dren.Interactions: Studies in Communication & Culture, 13(3):289–306, 2022. 2

  46. [46]

    Nanobanana: Gemini-2.5-flash- image.https://deepmind.google/models/ gemini-image/flash/, 2025

    Google/DeepMind. Nanobanana: Gemini-2.5-flash- image.https://deepmind.google/models/ gemini-image/flash/, 2025. 5

  47. [47]

    Disability stereotyping is shaped by stigma characteristics.Group Processes & In- tergroup Relations, 27(6):1403–1422, 2024

    Marine Granjon, Odile Rohmer, Maria Popa-Roch, Benoite Aub´e, and Camille Sanrey. Disability stereotyping is shaped by stigma characteristics.Group Processes & In- tergroup Relations, 27(6):1403–1422, 2024. 2, 5, 8

  48. [48]

    Unpacking the interdependent systems of discrimina- tion: Ableist bias in nlp systems through an intersectional lens

    Saad Hassan, Matt Huenerfauth, and Cecilia Ovesdotter Alm. Unpacking the interdependent systems of discrimina- tion: Ableist bias in nlp systems through an intersectional lens. InFindings of the Association for Computational Lin- guistics: EMNLP 2021, pages 3116–3123, 2021. 2

  49. [49]

    Ap- plying the stereotype content model to assess disability bias in popular pre-trained nlp models underlying ai-based as- sistive technologies

    Brienna Herold, James Waller, and Raja Kushalnagar. Ap- plying the stereotype content model to assess disability bias in popular pre-trained nlp models underlying ai-based as- sistive technologies. InNinth workshop on speech and lan- guage processing for assistive technologies (SLPAT-2022), pages 58–65, 2022. 2, 8

  50. [50]

    Clipscore: A reference-free evalu- ation metric for image captioning, 2022

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evalu- ation metric for image captioning, 2022. 4

  51. [51]

    vulnerable, victimized, and objectified

    Sharon Heung, Lucy Jiang, Shiri Azenkot, and Aditya Vashistha. “vulnerable, victimized, and objectified”: Un- derstanding ableist hate and harassment experienced by dis- abled content creators on social media. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–19, 2024. 2

  52. [52]

    Stereotypes.An- nual review of psychology, 47(1):237–271, 1996

    James L Hilton and William V on Hippel. Stereotypes.An- nual review of psychology, 47(1):237–271, 1996. 1

  53. [53]

    Hsin-Ping Huang, Xinyi Wang, Yonatan Bitton, Hagai Taitelbaum, Gaurav Singh Tomar, Ming-Wei Chang, Xuhui Jia, Kelvin C. K. Chan, Hexiang Hu, Yu-Chuan Su, and Ming-Hsuan Yang. Kitten: A knowledge-intensive evalua- tion of image generation on visual entities, 2024. 3

  54. [54]

    Unin- tended machine learning biases as social barriers for per- sons with disabilitiess.ACM SIGACCESS Accessibility and Computing, (125):1–1, 2020

    Ben Hutchinson, Vinodkumar Prabhakaran, Remi Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl. Unin- tended machine learning biases as social barriers for per- sons with disabilitiess.ACM SIGACCESS Accessibility and Computing, (125):1–1, 2020. 2

  55. [55]

    Reddy, and Sunipa Dev

    Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan K. Reddy, and Sunipa Dev. Visage: A global-scale analysis of visual stereotypes in text-to-image generation, 2024. 3

  56. [56]

    Beyond aesthetics: Cultural competence in text-to-image models, 2025

    Nithish Kannen, Arif Ahmad, Marco Andreetto, Vinodku- mar Prabhakaran, Utsav Prabhu, Adji Bousso Dieng, Push- pak Bhattacharyya, and Shachi Dave. Beyond aesthetics: Cultural competence in text-to-image models, 2025. 3

  57. [57]

    Reference-based metrics are biased against blind and low-vision users’ image de- scription preferences

    Rhea Kapur and Elisa Kreiss. Reference-based metrics are biased against blind and low-vision users’ image de- scription preferences. In Daryna Dementieva, Oana Ignat, Zhijing Jin, Rada Mihalcea, Giorgio Piatti, Joel Tetreault, Steven Wilson, and Jieyu Zhao, editors,Proceedings of the Third Workshop on NLP for Positive Impact, pages 308– 314, Miami, Florid...

  58. [58]

    Taxonomizing and measuring representational harms: A look at image tagging

    Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. Taxonomizing and measuring representational harms: A look at image tagging. InProceedings of the AAAI Conference on artificial intelligence, volume 37, pages 14277–14285, 2023. 1

  59. [59]

    Examining gender and race bias in two hundred sentiment analysis sys- tems

    Svetlana Kiritchenko and Saif Mohammad. Examining gender and race bias in two hundred sentiment analysis sys- tems. InProceedings of the seventh joint conference on lex- ical and computational semantics, pages 43–53, 2018. 2

  60. [60]

    Fairface: Face at- tribute dataset for balanced race, gender, and age, 2019

    Kimmo K ¨arkk¨ainen and Jungseock Joo. Fairface: Face at- tribute dataset for balanced race, gender, and age, 2019. 3

  61. [61]

    Holis- tic evaluation of text-to-image models, 2023

    Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Min- guk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, and Percy Liang. Holis- tic evaluation of text-to-image models, 2023. 3

  62. [62]

    Playground v2.5: Three in- sights towards enhancing aesthetic quality in text-to-image generation, 2024

    Daiqing Li, Aleks Kamko, Ehsan Akhgari, Ali Sabet, Lin- miao Xu, and Suhail Doshi. Playground v2.5: Three in- sights towards enhancing aesthetic quality in text-to-image generation, 2024. 5

  63. [63]

    Towards equitable representation in text-to-image synthesis models with the cross-cultural un- derstanding benchmark (ccub) dataset, 2023

    Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu, Youngsik Yun, Lia Coleman, Peter Schaldenbrand, Jihie Kim, and Jean Oh. Towards equitable representation in text-to-image synthesis models with the cross-cultural un- derstanding benchmark (ccub) dataset, 2023. 3

  64. [64]

    Stable bias: Analyzing so- cietal representations in diffusion models, 2023

    Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. Stable bias: Analyzing so- cietal representations in diffusion models, 2023. 3

  65. [65]

    Faintbench: A holistic and precise benchmark for bias eval- uation in text-to-image models, 2025

    Hanjun Luo, Ziye Deng, Ruizhe Chen, and Zuozhu Liu. Faintbench: A holistic and precise benchmark for bias eval- uation in text-to-image models, 2025. 3

  66. [66]

    Bigbench: A unified benchmark for evaluating multi- dimensional social biases in text-to-image models, 2025

    Hanjun Luo, Haoyu Huang, Ziye Deng, Xinfeng Li, Hewei Wang, Yingbin Jin, Yang Liu, Wenyuan Xu, and Zuozhu Liu. Bigbench: A unified benchmark for evaluating multi- dimensional social biases in text-to-image models, 2025. 3

  67. [67]

    they only care to show us 11 the wheelchair

    Kelly Avery Mack, Rida Qadri, Remi Denton, Shaun K. Kane, and Cynthia L. Bennett. “they only care to show us 11 the wheelchair”: disability representation in text-to-image ai models. InProceedings of the CHI Conference on Hu- man Factors in Computing Systems, CHI ’24, page 1–23. ACM, May 2024. 1, 2

  68. [68]

    Bias against 93 stigmatized groups in masked language models and downstream sentiment classification tasks

    Katelyn Mei, Sonia Fereidooni, and Aylin Caliskan. Bias against 93 stigmatized groups in masked language models and downstream sentiment classification tasks. InProceed- ings of the 2023 ACM Conference on Fairness, Account- ability, and Transparency, pages 1699–1710, 2023. 2

  69. [69]

    Social biases through the text-to-image generation lens

    Ranjita Naik and Besmira Nushi. Social biases through the text-to-image generation lens. InProceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 786–808, 2023. 2

  70. [70]

    Gpt-image-1-mini.https:// developers.openai.com/api/docs/models/ gpt-image-1-mini, 2025

    OpenAI. Gpt-image-1-mini.https:// developers.openai.com/api/docs/models/ gpt-image-1-mini, 2025. 5

  71. [71]

    Accesseval: Benchmarking disability bias in large lan- guage models, 2025

    Srikant Panda, Amit Agarwal, and Hitesh Laxmichand Pa- tel. Accesseval: Benchmarking disability bias in large lan- guage models, 2025. 2

  72. [72]

    Ableist: Intersectional dis- ability bias in llm-generated hiring scenarios, 2025

    Mahika Phutane, Hayoung Jung, Matthew Kim, Tanushree Mitra, and Aditya Vashistha. Ableist: Intersectional dis- ability bias in llm-generated hiring scenarios, 2025. 2

  73. [73]

    cold, calculated, and condescending

    Mahika Phutane, Ananya Seelam, and Aditya Vashistha. “cold, calculated, and condescending”: How ai identifies and explains ableism compared to disabled people. InPro- ceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, pages 1927–1941, 2025. 2

  74. [74]

    Disability across cultures: A human-centered audit of ableism in western and indic llms

    Mahika Phutane and Aditya Vashistha. Disability across cultures: A human-centered audit of ableism in western and indic llms. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 8, pages 2000–2014, 2025. 2

  75. [75]

    Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023. 5

  76. [76]

    Hidden bias in the machine: Stereotypes in text-to-image models, 2025

    Sedat Porikli and Vedat Porikli. Hidden bias in the machine: Stereotypes in text-to-image models, 2025. 3

  77. [77]

    Participa- tory machine learning using community-based system dy- namics.Health and Human Rights, 22(2):71, 2020

    Vinodkumar Prabhakaran and Donald Martin Jr. Participa- tory machine learning using community-based system dy- namics.Health and Human Rights, 22(2):71, 2020. 2

  78. [78]

    Ai’s regimes of representation: A community- centered study of text-to-image models in south asia

    Rida Qadri, Renee Shelby, Cynthia L Bennett, and Remi Denton. Ai’s regimes of representation: A community- centered study of text-to-image models in south asia. In Proceedings of the 2023 ACM Conference on Fairness, Ac- countability, and Transparency, pages 506–517, 2023. 2

  79. [79]

    Lumina-image 2.0: A unified and efficient image generative framework, 2025

    Qi Qin, Le Zhuo, Yi Xin, Ruoyi Du, Zhen Li, Bin Fu, Yiting Lu, Jiakang Yuan, Xinyue Li, Dongyang Liu, Xiangyang Zhu, Manyuan Zhang, Will Beddow, Erwann Millon, Victor Perez, Wenhai Wang, Conghui He, Bo Zhang, Xiaohong Liu, Hongsheng Li, Yu Qiao, Chang Xu, and Peng Gao. Lumina-image 2.0: A unified and efficient image generative framework, 2025. 5

  80. [80]

    Columbia University Press, 2007

    Ato Quayson.Aesthetic nervousness: Disability and the crisis of representation. Columbia University Press, 2007. 2

Showing first 80 references.