REVIEW 3 cited by
Survey of Social Bias in Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Survey of Social Bias in Vision-Language Models
read the original abstract
In recent years, the rapid advancement of machine learning (ML) models, particularly transformer-based pre-trained models, has revolutionized Natural Language Processing (NLP) and Computer Vision (CV) fields. However, researchers have discovered that these models can inadvertently capture and reinforce social biases present in their training datasets, leading to potential social harms, such as uneven resource allocation and unfair representation of specific social groups. Addressing these biases and ensuring fairness in artificial intelligence (AI) systems has become a critical concern in the ML community. The recent introduction of pre-trained vision-and-language (VL) models in the emerging multimodal field demands attention to the potential social biases present in these models as well. Although VL models are susceptible to social bias, there is a limited understanding compared to the extensive discussions on bias in NLP and CV. This survey aims to provide researchers with a high-level insight into the similarities and differences of social bias studies in pre-trained models across NLP, CV, and VL. By examining these perspectives, the survey aims to offer valuable guidelines on how to approach and mitigate social bias in both unimodal and multimodal settings. The findings and recommendations presented here can benefit the ML community, fostering the development of fairer and non-biased AI models in various applications and research endeavors.
Forward citations
Cited by 3 Pith papers
-
When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models
Vision-language models confidently misread face-like objects as people, while pure-vision classifiers stay diffuse and detectors suppress responses, showing uncertainty and bias are independent.
-
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
FairCLIP's claimed fairness and performance gains over CLIP do not reproduce on two datasets, and its official implementation diverges from the paper's own formulation.
-
Vision-Language Models display a strong gender bias
Using cosine similarity in CLIP embedding space, the paper finds that male and female face sets are differentially associated with occupation and activity statements across all four tested models.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.