Pith. sign in

REVIEW 3 cited by

Semantic and Expressive Variation in Image Captions Across Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.14356 v5 pith:TCDCWSXF submitted 2023-10-22 cs.CV cs.CLcs.CYcs.HC

classification cs.CVcs.CLcs.CYcs.HC
keywords differentmodelsacrosscaptionsdatadatasetsdescriptionslanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Computer vision often treats human perception as homogeneous: an implicit assumption that visual stimuli are perceived similarly by everyone. This assumption is reflected in the way researchers collect datasets and train vision models. By contrast, literature in cross-cultural psychology and linguistics has provided evidence that people from different cultural backgrounds observe vastly different concepts even when viewing the same visual stimuli. In this paper, we study how these differences manifest themselves in vision-language datasets and models, using language as a proxy for culture. By comparing textual descriptions generated across 7 languages for the same images, we find significant differences in the semantic content and linguistic expression. When datasets are multilingual as opposed to monolingual, descriptions have higher semantic coverage on average, where coverage is measured using scene graphs, model embeddings, and linguistic taxonomies. For example, multilingual descriptions have on average 29.9% more objects, 24.5% more relations, and 46.0% more attributes than a set of monolingual captions. When prompted to describe images in different languages, popular models (e.g. LLaVA) inherit this bias and describe different parts of the image. Moreover, finetuning models on captions from one language performs best on corresponding test data from that language, while finetuning on multilingual data performs consistently well across all test data compositions. Our work points towards the need to account for and embrace the diversity of human perception in the computer vision community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation of Cultural Competence of Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The paper proposes five theory-informed frameworks from visual cultural studies for evaluating cultural competence in vision-language models.

  2. Contrasting Cognitive Styles in Vision-Language Models: Holistic Attention in Japanese Versus Analytical Focus in English

    cs.CL 2025-07 reject novelty 5.0 of 10

    Japanese-prompted vision-language models produce more background-first captions than English-prompted ones, but the effect is confounded by the evaluator and by language grammar.

  3. Hidden Bias in the Machine: Stereotypes in Text-to-Image Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Text-to-image models reproduce and amplify stereotypes about gender, race, age, and body type across a broad set of everyday prompt categories.

Pith tools