Pith. sign in

REVIEW 1 cited by

AfroLID: A Neural Language Identification Tool for African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.11744 v3 pith:KQZ7G46O submitted 2022-10-21 cs.CL cs.LG

AfroLID: A Neural Language Identification Tool for African Languages

classification cs.CL cs.LG
keywords afrolidlanguagesafricanlanguagefiveidentificationneuralnumber
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Language identification (LID) is a crucial precursor for NLP, especially for mining web data. Problematically, most of the world's 7000+ languages today are not covered by LID technologies. We address this pressing issue for Africa by introducing AfroLID, a neural LID toolkit for $517$ African languages and varieties. AfroLID exploits a multi-domain web dataset manually curated from across 14 language families utilizing five orthographic systems. When evaluated on our blind Test set, AfroLID achieves 95.89 F_1-score. We also compare AfroLID to five existing LID tools that each cover a small number of African languages, finding it to outperform them on most languages. We further show the utility of AfroLID in the wild by testing it on the acutely under-served Twitter domain. Finally, we offer a number of controlled case studies and perform a linguistically-motivated error analysis that allow us to both showcase AfroLID's powerful capabilities and limitations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Language Diversity: Evaluating Language Usage and AI Performance on African Languages in Digital Spaces

    cs.CL 2025-12 conditional novelty 4.0

    Language detection models are near-perfect on clean news text in Yoruba, Kinyarwanda, and Amharic but perform poorly on code-switched Reddit posts, suggesting curated news data is more reliable for training.