REVIEW 5 cited by
IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench -- a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference~(AfriXNLI), mathematical reasoning~(AfriMGSM), and multi-choice knowledge-based question answering~(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings~(where test sets are translated into English) across 10 open and six proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages~(such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63\% of the best-performing proprietary model GPT-4o performance. In addition, machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, such as Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.
Forward citations
Cited by 5 Pith papers
-
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
Training-data mix—not base-model language coverage—was the main driver of improved African-language performance, with Qwen-3 bases gaining the most after continued pre-training.
-
Against 'softmaxing' culture
A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.
-
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents
Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.
-
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
Data mixing with math, code, and synthetic translations during continued pre-training improves LLM performance on African languages and reasoning tasks, with data composition as the primary driver.
-
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.
Discussion (0). Sign in to comment.