Pith. sign in

REVIEW 2 cited by

Balancing the Picture: Debiasing Vision-Language Datasets with Synthetic Contrast Sets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15407 v1 pith:CF7QDC5P submitted 2023-05-24 cs.CV

classification cs.CV
keywords biasimagesdatasetcococontrastdebiasingmodelssets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language models are growing in popularity and public visibility to generate, edit, and caption images at scale; but their outputs can perpetuate and amplify societal biases learned during pre-training on uncurated image-text pairs from the internet. Although debiasing methods have been proposed, we argue that these measurements of model bias lack validity due to dataset bias. We demonstrate there are spurious correlations in COCO Captions, the most commonly used dataset for evaluating bias, between background context and the gender of people in-situ. This is problematic because commonly-used bias metrics (such as Bias@K) rely on per-gender base rates. To address this issue, we propose a novel dataset debiasing pipeline to augment the COCO dataset with synthetic, gender-balanced contrast sets, where only the gender of the subject is edited and the background is fixed. However, existing image editing methods have limitations and sometimes produce low-quality images; so, we introduce a method to automatically filter the generated images based on their similarity to real images. Using our balanced synthetic contrast sets, we benchmark bias in multiple CLIP-based models, demonstrating how metrics are skewed by imbalance in the original COCO images. Our results indicate that the proposed approach improves the validity of the evaluation, ultimately contributing to more realistic understanding of bias in vision-language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PrefPalette: Personalized Preference Modeling with Latent Attributes

    cs.AI 2025-07 conditional novelty 6.0 of 10

    Decomposing text into latent attributes and learning community-specific attribute weights improves preference prediction on Reddit and yields interpretable community profiles.

  2. Vision-Language Models display a strong gender bias

    cs.CV 2025-08 reject novelty 3.0 of 10

    Using cosine similarity in CLIP embedding space, the paper finds that male and female face sets are differentially associated with occupation and activity statements across all four tested models.

Pith tools