Pith. sign in

REVIEW 4 cited by

MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10497 v2 pith:7I3KCDAM submitted 2025-03-13 cs.CL

classification cs.CL
keywords llmslanguagemmlu-proxmultilingualbenchmarkevaluationlanguagesquestions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-linguistic reasoning abilities. This dual limitation makes it challenging to comprehensively assess LLMs' performance in the multilingual setting. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-linguistic comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves multiple powerful LLMs for translation, followed by expert review to ensure accurate expression, consistent terminology, and cultural relevance. Building on this, we systematically evaluate 36 state-of-the-art LLMs, including reasoning-enhanced and multilingual-optimized LLMs. The results reveal significant disparities in the multilingual capabilities of LLMs: While they perform well in high-resource languages, their performance declines markedly in low-resource languages, with gaps of up to 24.3%. Through MMLU-ProX, we aim to advance the development of more inclusive AI systems and promote equitable access to technology across global contexts.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiNRC: A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs

    cs.CL 2025-07 conditional novelty 7.0 of 10

    A native-authored French, Spanish, and Chinese reasoning benchmark shows current LLMs score below 50% and improve by about 10% on math when questions are in English.

  2. CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting

    cs.LG 2025-05 accept novelty 7.0 of 10

    Capping achievable accuracy with randomized correct answers turns any model that exceeds the cap into a detectable contamination alarm.

  3. Disentangling Language Modeling and Boundaries

    cs.CL 2026-08 reject novelty 6.0 of 10

    The paper hypothesizes that next-byte and boundary distributions in byte-level LMs can be disentangled, proposes two experiments to test it, but provides no experimental results.

  4. Do LLMs exhibit the same commonsense capabilities across languages?

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs produce more commonsensical sentences in English than in Spanish, Dutch, or Valencian, across automatic, LLM-judge, and human evaluations on the new MULTICOM benchmark.

Pith tools