REVIEW 6 cited by
DALL-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recently, DALL-E, a multimodal transformer language model, and its variants, including diffusion models, have shown high-quality text-to-image generation capabilities. However, despite the realistic image generation results, there has not been a detailed analysis of how to evaluate such models. In this work, we investigate the visual reasoning capabilities and social biases of different text-to-image models, covering both multimodal transformer language models and diffusion models. First, we measure three visual reasoning skills: object recognition, object counting, and spatial relation understanding. For this, we propose PaintSkills, a compositional diagnostic evaluation dataset that measures these skills. Despite the high-fidelity image generation capability, a large gap exists between the performance of recent models and the upper bound accuracy in object counting and spatial relation understanding skills. Second, we assess the gender and skin tone biases by measuring the gender/skin tone distribution of generated images across various professions and attributes. We demonstrate that recent text-to-image generation models learn specific biases about gender and skin tone from web image-text pairs. We hope our work will help guide future progress in improving text-to-image generation models on visual reasoning skills and learning socially unbiased representations. Code and data: https://github.com/j-min/DallEval
Forward citations
Cited by 6 Pith papers
-
COVAriance-Induced Fairness Gap Penalty for Subgroup-Fair Clustering
A covariance quantity is proven exactly equal to a subgroup-fairness gap for clustering, yielding COVA-FC, a scalable algorithm that can also enforce marginal fairness.
-
Discovering Divergent Representations between Text-to-Image Models
An evolutionary algorithm discovers visual attributes that appear in one text-to-image model's outputs but not another's, and identifies the prompt concepts that trigger them.
-
Evaluating and comparing gender bias across four text-to-image models
Across 30 professions and 6,000 images, DALL-E 3 over-represented women, Stable Diffusion XL and Cascade over-represented men in high-status roles, and Emu was more balanced.
-
Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
A new contrastive real-image benchmark shows most vision-language models fail spatial relation tasks, while chain-of-thought reasoning models approach human-level accuracy.
-
Adultification Bias in LLMs and Text-to-Image Models
Large language and text-to-image models show measurable adultification bias, portraying Black girls as more mature, culpable, and sexualized than White girls in several tested models.
-
Federated Learning Inspired Fuzzy Systems: Decentralized Rule Updating for Privacy and Scalable Decision Making
The paper suggests federated-learning-style updates for fuzzy rule sets and a machine-learning-augmented fuzzy system, without providing implementation, derivation, or evidence.
Discussion (0). Sign in to comment.