Pith. sign in

REVIEW 4 cited by

Examining Gender and Racial Bias in Large Vision-Language Models Using a Novel Dataset of Parallel Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05779 v1 pith:Z2NOCZDR submitted 2024-02-08 cs.CY cs.CLcs.CV

classification cs.CYcs.CLcs.CV
keywords imagesmodelsgenderdatasetlargeinputlvlmspairs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged. Such models can incorporate images as input in addition to text, and perform tasks such as visual question answering, image captioning, story generation, etc. Here, we examine potential gender and racial biases in such systems, based on the perceived characteristics of the people in the input images. To accomplish this, we present a new dataset PAIRS (PArallel Images for eveRyday Scenarios). The PAIRS dataset contains sets of AI-generated images of people, such that the images are highly similar in terms of background and visual content, but differ along the dimensions of gender (man, woman) and race (Black, white). By querying the LVLMs with such images, we observe significant differences in the responses according to the perceived gender or race of the person depicted.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AHELM: A Holistic Evaluation of Audio-Language Models

    cs.AI 2025-08 conditional novelty 6.0 of 10

    AHELM standardizes evaluation of audio-language models across 10 aspects and shows simple ASR+LLM systems are competitive with multimodal models.

  2. Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances

    cs.CY 2025-07 accept novelty 6.0 of 10

    A Fair Equality of Chances-based framework decomposes GenAI unfairness into harms/benefits, morally arbitrary factors, and morally decisive factors to improve measurement validity.

  3. From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.

  4. Vision-Language Models display a strong gender bias

    cs.CV 2025-08 reject novelty 3.0 of 10

    Using cosine similarity in CLIP embedding space, the paper finds that male and female face sets are differentially associated with occupation and activity statements across all four tested models.

Pith tools