Pith. sign in

REVIEW 5 cited by

TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12402 v2 pith:NTUEHVNV submitted 2024-07-17 cs.CL

classification cs.CL
keywords turkishllmsevaluationlanguagequestionsturkishmmluevaluateanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automatic translation for multilingual evaluation, this approach is error-prone and potentially introduces culturally biased questions, especially in social sciences. We introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU, to evaluate LLMs' understanding of the Turkish language. TurkishMMLU includes over 10,000 questions, covering 9 different subjects from Turkish high-school education curricula. These questions are written by curriculum experts, suitable for the high-school curricula in Turkey, covering subjects ranging from natural sciences and math questions to more culturally representative topics such as Turkish Literature and the history of the Turkish Republic. We evaluate over 20 LLMs, including multilingual open-source (e.g., Gemma, Llama, MT5), closed-source (GPT 4o, Claude, Gemini), and Turkish-adapted (e.g., Trendyol) models. We provide an extensive evaluation, including zero-shot and few-shot evaluation of LLMs, chain-of-thought reasoning, and question difficulty analysis along with model performance. We provide an in-depth analysis of the Turkish capabilities and limitations of current LLMs to provide insights for future LLMs for the Turkish language. We publicly release our code for the dataset and evaluation: https://github.com/ArdaYueksel/TurkishMMLU.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMzSz{\L}: a comprehensive LLM benchmark for Polish

    cs.CL 2025-01 conditional novelty 6.0 of 10

    LLMzSzŁ is a new benchmark of almost 19,000 Polish national exam questions with evaluations of 38 language models and comparisons to human results.

  2. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  3. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

    cs.CL 2024-11 conditional novelty 6.0 of 10

    INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.

  4. Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

    cs.CL 2024-12 reject novelty 4.0 of 10

    TR-MMLU is a 6,200-question Turkish multiple-choice benchmark for LLMs, but the paper lacks data access, contamination checks, and baselines.

  5. A Comparative Study of Text Retrieval Models on DaReCzech

    cs.IR 2024-11 conditional novelty 4.0 of 10

    A benchmark on the Czech DaReCzech dataset finds Gemma2 most accurate, Contriever least accurate, and SPLADE/PLAID the best efficiency-quality trade-off.

Pith tools