Pith. sign in

REVIEW

Mining Wikidata for Name Resources for African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.00558 v1 pith:Y2GGIVFE submitted 2021-04-01 cs.CL cs.AI

classification cs.CLcs.AI
keywords languagesdatalistsnameafricanproduceresourcewikidata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work supports further development of language technology for the languages of Africa by providing a Wikidata-derived resource of name lists corresponding to common entity types (person, location, and organization). While we are not the first to mine Wikidata for name lists, our approach emphasizes scalability and replicability and addresses data quality issues for languages that do not use Latin scripts. We produce lists containing approximately 1.9 million names across 28 African languages. We describe the data, the process used to produce it, and its limitations, and provide the software and data for public use. Finally, we discuss the ethical considerations of producing this resource and others of its kind.

Discussion (0). Sign in to comment.

Pith tools