Pith. sign in

REVIEW 4 cited by

AfroBench: How Good are Large Language Models on African Languages?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07978 v5 pith:U6T5G3W4 submitted 2023-11-14 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesafricanperformanceafrobenchdatasetstasksacrossllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African languages due to the scarcity of high-quality evaluation data and the limited discoverability of existing African datasets. This lack of representation hinders comprehensive LLM evaluation across a diverse range of languages and tasks. To address these challenges, we introduce AfroBench -- a multi-task benchmark for evaluating the performance of LLMs across 64 African languages, 15 tasks and 22 datasets. AfroBench consists of nine natural language understanding datasets, six text generation datasets, six knowledge and question answering tasks, and one mathematical reasoning task. We present results comparing the performance of prompting LLMs to fine-tuned baselines based on BERT and T5-style models. Our results suggest large gaps in performance between high-resource languages, such as English, and African languages across most tasks; but performance also varies based on the availability of monolingual data resources. Our findings confirm that performance on African languages continues to remain a hurdle for current LLMs, underscoring the need for additional efforts to close this gap. https://mcgill-nlp.github.io/AfroBench/

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. mRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages

    cs.CL 2025-07 conditional novelty 5.0 of 10

    mRAKL reformulates multilingual knowledge graph completion as question answering and shows that retrieving context from Wikipedia improves tail-entity prediction for Tigrinya and Amharic, with gains up to 8.79 points ...

  2. mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

    cs.CL 2025-06 conditional novelty 5.0 of 10

    mSTEB is a new 200+ language speech and text benchmark showing that LLMs perform substantially worse on low-resource African and Americas/Oceania languages, especially in speech tasks.

  3. The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A continuously updated multilingual benchmark dashboard aggregates existing tasks to rank LLMs across up to 200 languages.

  4. LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa

    cs.CE 2025-08 unverdicted novelty 3.0 of 10

    LLMs and agentic AI are presented as a transformative opportunity for African insurance, with a call for African-led, equitable AI strategies.

Pith tools