Pith. sign in

REVIEW 10 cited by

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.01282 v1 pith:I6AHPNSR submitted 2025-01-02 cs.AI cs.CLcs.CV

CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

classification cs.AI cs.CLcs.CV
keywords culturalmodelsunderstandingconceptsperformancevlmscharacterizingcountries
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Vision-language models (VLMs) have advanced human-AI interaction but struggle with cultural understanding, often misinterpreting symbols, gestures, and artifacts due to biases in predominantly Western-centric training data. In this paper, we construct CultureVerse, a large-scale multimodal benchmark covering 19, 682 cultural concepts, 188 countries/regions, 15 cultural concepts, and 3 question types, with the aim of characterizing and improving VLMs' multicultural understanding capabilities. Then, we propose CultureVLM, a series of VLMs fine-tuned on our dataset to achieve significant performance improvement in cultural understanding. Our evaluation of 16 models reveals significant disparities, with a stronger performance in Western concepts and weaker results in African and Asian contexts. Fine-tuning on our CultureVerse enhances cultural perception, demonstrating cross-cultural, cross-continent, and cross-dataset generalization without sacrificing performance on models' general VLM benchmarks. We further present insights on cultural generalization and forgetting. We hope that this work could lay the foundation for more equitable and culturally aware multimodal AI systems.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

    cs.CL 2026-04 unverdicted novelty 7.0

    A unified survey that consolidates Indian NLP resources by task, language, domain, and modality while identifying gaps in coverage and generalization.

  2. Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images

    cs.CV 2026-04 unverdicted novelty 7.0

    A new cross-cultural benchmark shows vision-language models infer structured cultural metadata from images inconsistently, with fragmented signals and large performance gaps across regions and metadata types.

  3. Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.5

    Pre-generation visual tokens and prompt hidden states can diagnose whether a VLM will fail from recognition, visual evidence, or factual knowledge, enabling targeted interventions.

  4. Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0

    Pre-generation visual tokens and prompt hidden states can predict whether a VLM will fail from recognition bottlenecks or from post-recognition factual gaps, enabling targeted interventions.

  5. Failing to See or Failing to Know? Attributing Errors in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0

    VLM wrong answers in knowledge-intensive visual QA can be attributed to four decision points—recognition, visual evidence, answer success, factual access—with different pre-generation representations best predicting e...

  6. Computer-Aided Tagging on Wikimedia Commons: Designing for Human-AI Collaboration in Open Knowledge Work

    cs.HC 2026-05 unverdicted novelty 6.0

    Qualitative study of the CAT tool on Wikimedia Commons identifies seven issues from community comments and interviews that led to mixed reception and deactivation, with suggestions for human-AI collaboration in open k...

  7. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 conditional novelty 6.0

    Parallel, role-specialized prompt agents improve cultural relevance in text-to-video generation, with a new cross-cultural benchmark showing the largest gains for location cues.

  8. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN is a multi-agent prompt refinement framework that improves cultural fidelity in text-to-video generation, demonstrated on a new benchmark of 243 prompts and 972 videos across Chinese, American, and Romanian cultures.

  9. When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    MAVEN introduces a multi-agent system for refining prompts in multicultural text-to-video generation and releases a benchmark of 243 prompts and 972 videos showing improved cultural relevance via parallel agent specia...

  10. Large Language Model Agent: A Survey on Methodology, Applications and Challenges

    cs.CL 2025-03 accept novelty 3.0

    A survey that deconstructs LLM agent systems via a methodology-centered taxonomy linking design principles to emergent behaviors, applications, and challenges.