REVIEW 5 cited by
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs). While existing benchmarks employ automatic translation for multilingual evaluation, this approach is error-prone and potentially introduces culturally biased questions, especially in social sciences. We introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU, to evaluate LLMs' understanding of the Turkish language. TurkishMMLU includes over 10,000 questions, covering 9 different subjects from Turkish high-school education curricula. These questions are written by curriculum experts, suitable for the high-school curricula in Turkey, covering subjects ranging from natural sciences and math questions to more culturally representative topics such as Turkish Literature and the history of the Turkish Republic. We evaluate over 20 LLMs, including multilingual open-source (e.g., Gemma, Llama, MT5), closed-source (GPT 4o, Claude, Gemini), and Turkish-adapted (e.g., Trendyol) models. We provide an extensive evaluation, including zero-shot and few-shot evaluation of LLMs, chain-of-thought reasoning, and question difficulty analysis along with model performance. We provide an in-depth analysis of the Turkish capabilities and limitations of current LLMs to provide insights for future LLMs for the Turkish language. We publicly release our code for the dataset and evaluation: https://github.com/ArdaYueksel/TurkishMMLU.
Forward citations
Cited by 5 Pith papers
-
LLMzSz{\L}: a comprehensive LLM benchmark for Polish
LLMzSzŁ is a new benchmark of almost 19,000 Polish national exam questions with evaluations of 38 language models and comparisons to human results.
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.
-
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.
-
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
TR-MMLU is a 6,200-question Turkish multiple-choice benchmark for LLMs, but the paper lacks data access, contamination checks, and baselines.
-
A Comparative Study of Text Retrieval Models on DaReCzech
A benchmark on the Czech DaReCzech dataset finds Gemma2 most accurate, Contriever least accurate, and SPLADE/PLAID the best efficiency-quality trade-off.
Discussion (0). Continue with ORCID to comment.