REVIEW 14 cited by
ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts, evaluating their performance in Arabic remains challenging due to the limited availability of relevant datasets. To bridge this gap, we present \datasetname{}, the first multi-task language understanding benchmark for the Arabic language, sourced from school exams across diverse educational levels in different countries spanning North Africa, the Levant, and the Gulf regions. Our data comprises 40 tasks and 14,575 multiple-choice questions in Modern Standard Arabic (MSA) and is carefully constructed by collaborating with native speakers in the region. Our comprehensive evaluations of 35 models reveal substantial room for improvement, particularly among the best open-source models. Notably, BLOOMZ, mT0, LLaMA2, and Falcon struggle to achieve a score of 50%, while even the top-performing Arabic-centric model only achieves a score of 62.3%.
Forward citations
Cited by 14 Pith papers
-
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset
MALAMUTE is a 116k-prompt cloze-style dataset derived from 71 university textbooks in three languages, used to probe language models' fine-grained subject knowledge.
-
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
An adaptive, per-language data filtering and deduplication pipeline produces multilingual LLM pre-training corpora that beat prior public datasets on 11 of 14 evaluated languages, and a 20TB, 1,868 language-script dat...
-
How well can LLMs Grade Essays in Arabic?
Generative LLMs, including Arabic-specific ones, underperform a fine-tuned BERT model on Arabic essay scoring, with bilingual prompting giving the best LLM results.
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.
-
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge
INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.
-
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
ALM-bench is a 100-language, 19-domain cultural visual QA benchmark on which GPT-4o reaches 78.8% and the best open model, GLM-4V-9B, reaches 51.9%.
-
RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment
A fully open 518M Arabic-specialized LLM, built by vocabulary injection and standard post-training on Qwen2.5-0.5B, beats same-class multilingual baselines and ships at 398 MB quantized.
-
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering
KoLasSimpleQA is a new multilingual simple factuality benchmark with general and language-specific domains, showing LLMs are consistently weaker on language-specific knowledge across nine languages.
-
AIN: The Arabic INclusive Large Multimodal Model
AIN, a 7B-parameter Arabic English multimodal model fine-tuned from Qwen2-VL on 3.6M samples, reports state-of-the-art Arabic scores including a 3.4-point average gain over GPT-4o on CAMEL-Bench.
-
Fanar: An Arabic-Centric Multimodal Generative AI Platform
Fanar presents two Arabic LLMs, a morphology-aware tokenizer, and multimodal services, with strong but uneven benchmark evidence for top Arabic performance.
-
Arabic Stable LM: Adapting Stable LM 2 1.6B to Arabic
A 1.6B Arabic language model trained with synthetic multiple-choice instruction data outperforms 7B-13B models on several Arabic multiple-choice benchmarks.
-
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
Arabic instructions on the R2R navigation task preserve much of GPT-4o mini's success rate but push Phi-3 and Jais to zero, indicating model capability rather than language is the decisive factor.
-
A Comparative Study of Text Retrieval Models on DaReCzech
A benchmark on the Czech DaReCzech dataset finds Gemma2 most accurate, Contriever least accurate, and SPLADE/PLAID the best efficiency-quality trade-off.
-
Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks
Across Arabic, English, and Kannada benchmarks, 4-bit and 8-bit quantization preserves most accuracy while aggressive pruning degrades larger multilingual models more than smaller ones.
Discussion (0). Continue with ORCID to comment.