Pith. sign in

REVIEW 8 cited by

The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.15521 v1 pith:S4UB4TFU submitted 2025-04-22 cs.CL

classification cs.CL
keywords benchmarksmultilingualhumanbenchmarkingcorrelationscountriesenglishevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As large language models (LLMs) continue to advance in linguistic capabilities, robust multilingual evaluation has become essential for promoting equitable technological progress. This position paper examines over 2,000 multilingual (non-English) benchmarks from 148 countries, published between 2021 and 2024, to evaluate past, present, and future practices in multilingual benchmarking. Our findings reveal that, despite significant investments amounting to tens of millions of dollars, English remains significantly overrepresented in these benchmarks. Additionally, most benchmarks rely on original language content rather than translations, with the majority sourced from high-resource countries such as China, India, Germany, the UK, and the USA. Furthermore, a comparison of benchmark performance with human judgments highlights notable disparities. STEM-related tasks exhibit strong correlations with human evaluations (0.70 to 0.85), while traditional NLP tasks like question answering (e.g., XQuAD) show much weaker correlations (0.11 to 0.30). Moreover, translating English benchmarks into other languages proves insufficient, as localized benchmarks demonstrate significantly higher alignment with local human judgments (0.68) than their translated counterparts (0.47). This underscores the importance of creating culturally and linguistically tailored benchmarks rather than relying solely on translations. Through this comprehensive analysis, we highlight six key limitations in current multilingual evaluation practices, propose the guiding principles accordingly for effective multilingual benchmarking, and outline five critical research directions to drive progress in the field. Finally, we call for a global collaborative effort to develop human-aligned benchmarks that prioritize real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

    cs.CL 2026-08 conditional novelty 7.0 of 10

    MameLoshnLM, trained by continuing pretraining Llama 3.1 8B on a curated Yiddish corpus, outperforms similar-scale open models on a new multi-task Yiddish benchmark and better retains Yiddish-specific loshn-koydesh vo...

  2. Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Across 30 languages, commercial LLMs outscore all open-weight models in every EU language, and non-English service costs more and scores lower, suggesting equality requires resources beyond public web crawls.

  3. BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

    cs.CL 2025-02 conditional novelty 7.0 of 10

    BOUQuET is a handcrafted, multicentric, paragraph-level machine translation evaluation dataset in 8 non-English pivot languages, designed to be community-extendable.

  4. KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

    cs.CL 2026-07 conditional novelty 6.0 of 10

    KyrgyzLLM-Bench evaluates 26 LLMs on native and translated Kyrgyz tasks, showing rankings transfer only partially and translated HellaSwag is unreliable.

  5. Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fine-grained Korean street-view benchmark shows that adding captions with place names to photos lets vision-language models pinpoint locations at high rates, highlighting privacy exposure.

  6. The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It

    cs.CL 2025-05 accept novelty 6.0 of 10

    LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.

  7. Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark

    cs.CL 2026-08 conditional novelty 5.0 of 10

    ProverbIT shows that large language models can complete Italian proverbs but often fail to select 'none of the above' when the exact ending is absent, revealing a gap between memorized knowledge and discriminative reasoning.

  8. From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    On a new 490-question Arabic depth dataset, Claude 3.5 Sonnet answered about 30 percent correctly, while GPT-4 answered about 9 percent, showing current models are weak on culturally specialized Arabic knowledge.

Pith tools