Pith. sign in

REVIEW 6 cited by

DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.04053 v3 pith:3THH3WOE submitted 2022-02-08 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords modelsgenerationskillstext-to-imagebiasesreasoninggenderobject
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, DALL-E, a multimodal transformer language model, and its variants, including diffusion models, have shown high-quality text-to-image generation capabilities. However, despite the realistic image generation results, there has not been a detailed analysis of how to evaluate such models. In this work, we investigate the visual reasoning capabilities and social biases of different text-to-image models, covering both multimodal transformer language models and diffusion models. First, we measure three visual reasoning skills: object recognition, object counting, and spatial relation understanding. For this, we propose PaintSkills, a compositional diagnostic evaluation dataset that measures these skills. Despite the high-fidelity image generation capability, a large gap exists between the performance of recent models and the upper bound accuracy in object counting and spatial relation understanding skills. Second, we assess the gender and skin tone biases by measuring the gender/skin tone distribution of generated images across various professions and attributes. We demonstrate that recent text-to-image generation models learn specific biases about gender and skin tone from web image-text pairs. We hope our work will help guide future progress in improving text-to-image generation models on visual reasoning skills and learning socially unbiased representations. Code and data: https://github.com/j-min/DallEval

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering

    stat.ML 2026-07 conditional novelty 6.0 of 10

    A covariance quantity is proven exactly equal to a subgroup-fairness gap for clustering, yielding COVA-FC, a scalable algorithm that can also enforce marginal fairness.

  2. Discovering Divergent Representations between Text-to-Image Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An evolutionary algorithm discovers visual attributes that appear in one text-to-image model's outputs but not another's, and identifies the prompt concepts that trigger them.

  3. Evaluating and comparing gender bias across four text-to-image models

    cs.CY 2025-09 conditional novelty 6.0 of 10

    Across 30 professions and 6,000 images, DALL-E 3 over-represented women, Stable Diffusion XL and Cascade over-represented men in high-status roles, and Emu was more balanced.

  4. Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.

  5. Adultification Bias in LLMs and Text-to-Image Models

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.

  6. Federated Learning Inspired Fuzzy Systems: Decentralized Rule Updating for Privacy and Scalable Decision Making

    cs.LG 2025-07 reject novelty 2.0 of 10

    The paper suggests federated-learning-style updates for fuzzy rule sets and a machine-learning-augmented fuzzy system, without providing implementation, derivation, or evidence.

Pith tools