Pith. sign in

Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged. Such models can incorporate images as input in addition to text, and perform tasks such as visual question answering, image captioning, story generation, etc. Here, we examine potential gender and racial biases in such systems, based on the perceived characteristics of the people in the input images. To accomplish this, we present a new dataset PAIRS (PArallel Images for eveRyday Scenarios). The PAIRS dataset contains sets of AI-generated images of people, such that the images are highly similar in terms of background and visual content, but differ along the dimensions of gender (man, woman) and race (Black, white). By querying the LVLMs with such images, we observe significant differences in the responses according to the perceived gender or race of the person depicted.

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • AHELM: A Holistic Evaluation of Audio-Language Models cs.AI · 2025-08-29 · conditional · none · ref 15 · internal anchor

    AHELM standardizes evaluation of audio-language models across 10 aspects and shows simple ASR+LLM systems are competitive with multimodal models.