Pith. sign in

REVIEW 2 cited by

From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00263 v1 pith:K7URBVKN submitted 2024-06-28 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords modelsacrossconceptsculturalculturesimagesvision-languagecountries
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural inclusivity, but they have limited coverage of cultures and do not adequately assess cultural diversity across universal as well as culture-specific local concepts. To address these limitations, we introduce the GlobalRG benchmark, comprising two challenging tasks: retrieval across universals and cultural visual grounding. The former task entails retrieving culturally diverse images for universal concepts from 50 countries, while the latter aims at grounding culture-specific concepts within images from 15 countries. Our evaluation across a wide range of models reveals that the performance varies significantly across cultures -- underscoring the necessity for enhancing multicultural understanding in vision-language models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RusCode: Russian Cultural Code Benchmark for Text-to-Image Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    RusCode is a new 1,250-prompt Russian/English benchmark for cultural awareness in text-to-image models, with human evaluation showing Russian-trained models outperform general models.

  2. CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

    cs.AI 2025-01 conditional novelty 6.0 of 10

    CultureVerse is a 188-country, 19k-concept visual QA benchmark, and fine-tuning open VLMs on it improves cultural accuracy, but the main evaluation shares concepts between training and test sets.

Pith tools