Pith. sign in

REVIEW 5 cited by

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03368 v2 pith:JGO6QKOC submitted 2024-06-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagesafricanmodelsenglishhigh-resourceirokobenchlanguagellms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (\eg African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench -- a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference~(AfriXNLI), mathematical reasoning~(AfriMGSM), and multi-choice knowledge-based question answering~(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings~(where test sets are translated into English) across 10 open and six proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages~(such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63\% of the best-performing proprietary model GPT-4o performance. In addition, machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, such as Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Training-data mix—not base-model language coverage—was the main driver of improved African-language performance, with Qwen-3 bases gaining the most after continued pre-training.

  2. Against 'softmaxing' culture

    cs.HC 2025-06 unverdicted novelty 5.0 of 10

    A position paper arguing that AI evaluations should shift from defining culture to understanding when culture becomes relationally valid.

  3. Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents

    cs.CL 2026-05 unverdicted novelty 4.0 of 10

    Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.

  4. AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages

    cs.CL 2026-01 conditional novelty 4.0 of 10

    Data mixing with math, code, and synthetic translations during continued pre-training improves LLM performance on African languages and reasoning tasks, with data composition as the primary driver.

  5. The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.

Pith tools