Pith. sign in

REVIEW 7 cited by

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16508 v4 pith:OUSH23QF submitted 2024-11-25 cs.CV cs.CL

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

classification cs.CV cs.CL
keywords languagesalm-benchlmmsdiversemodelsbenchmarkculturalculturally
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model's ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond 'One Language, One Script': Quantifying Orthographic Bias in Multilingual VLMs with PuMVR

    cs.CL 2026-06 unverdicted novelty 7.0

    PuMVR benchmark shows VLMs exhibit script-dependent bias on Punjabi tasks with accuracy gaps up to 16% and script consistency rates as low as 24.8%, even when visual input is provided.

  2. Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

    cs.CV 2026-06 conditional novelty 7.0

    A new benchmark for Punjabi reveals VLMs have large script-dependent performance gaps on identical tasks, with consistency as low as 24.8 percent.

  3. ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

    cs.DB 2026-06 unverdicted novelty 7.0

    ArtiFact is a new multi-modal dataset of 651k museum records used to benchmark cross-modal error detection with seven error categories and semantic query processing challenges.

  4. NRITYAM: Language Models Meet Art and Heritage of Dance

    cs.CL 2026-06 unverdicted novelty 6.0

    NRITYAM creates the largest multilingual benchmark for evaluating language models' understanding of dance traditions through expert-curated QA pairs.

  5. STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery

    eess.IV 2025-09 conditional novelty 6.0

    A new 1,000-video benchmark of stroke patients performing box-and-block sub-actions, with raw frames and 2D skeletons, establishes baseline action classification accuracy for seven models.

  6. An Effective Router for Vision-Language Model Selection

    cs.AI 2026-06 conditional novelty 5.0

    ARMS is a learned router for VLM selection trained on a new 32k-query multimodal dataset that outperforms GPT-4o on both in- and out-of-distribution tests after incremental adaptation.

  7. Toward LLMs Beyond English-Centric Development

    cs.CL 2026-05 unverdicted novelty 4.0

    Analysis of open-weight LLMs reveals strong English bias in generated sequences, with continual pre-training providing no cost benefit over from-scratch training for non-English adaptation.