REVIEW 10 cited by
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
read the original abstract
Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.
Forward citations
Cited by 10 Pith papers
-
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
A unified survey that consolidates Indian NLP resources by task, language, domain, and modality while identifying gaps in coverage and generalization.
-
Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images
A new cross-cultural benchmark shows vision-language models infer structured cultural metadata from images inconsistently, with fragmented signals and large performance gaps across regions and metadata types.
-
Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
Pre-generation visual tokens and prompt hidden states can diagnose whether a VLM will fail from recognition, visual evidence, or factual knowledge, enabling targeted interventions.
-
Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
Pre-generation visual tokens and prompt hidden states can predict whether a VLM will fail from recognition bottlenecks or from post-recognition factual gaps, enabling targeted interventions.
-
Failing to See or Failing to Know? Attributing Errors in Vision-Language Models
VLM wrong answers in knowledge-intensive visual QA can be attributed to four decision points—recognition, visual evidence, answer success, factual access—with different pre-generation representations best predicting e...
-
Computer-Aided Tagging on Wikimedia Commons: Designing for Human-AI Collaboration in Open Knowledge Work
Qualitative study of the CAT tool on Wikimedia Commons identifies seven issues from community comments and interviews that led to mixed reception and deactivation, with suggestions for human-AI collaboration in open k...
-
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.
-
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
MAVEN is a multi-agent prompt refinement framework that improves cultural fidelity in text-to-video generation, demonstrated on a new benchmark of 243 prompts and 972 videos across Chinese, American, and Romanian cultures.
-
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
MAVEN introduces a multi-agent system for refining prompts in multicultural text-to-video generation and releases a benchmark of 243 prompts and 972 videos showing improved cultural relevance via parallel agent specia...
-
Large Language Model Agent: A Survey on Methodology, Applications and Challenges
A survey that deconstructs LLM agent systems via a methodology-centered taxonomy linking design principles to emergent behaviors, applications, and challenges.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.