Pith. sign in

REVIEW 1 cited by

Taxi1500: A Multilingual Dataset for Text Classification in 1500 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.08487 v2 pith:EHRN45CA submitted 2023-05-15 cs.CL

classification cs.CL
keywords languagesdatasetclassificationdatatextannotateddatasetsextensively
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While natural language processing tools have been developed extensively for some of the world's languages, a significant portion of the world's over 7000 languages are still neglected. One reason for this is that evaluation datasets do not yet cover a wide range of languages, including low-resource and endangered ones. We aim to address this issue by creating a text classification dataset encompassing a large number of languages, many of which currently have little to no annotated data available. We leverage parallel translations of the Bible to construct such a dataset by first developing applicable topics and employing a crowdsourcing tool to collect annotated data. By annotating the English side of the data and projecting the labels onto other languages through aligned verses, we generate text classification datasets for more than 1500 languages. We extensively benchmark several existing multilingual language models using our dataset. To facilitate the advancement of research in this area, we will release our dataset and code.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An empirical study showing that smaller monolingual language models are more vulnerable to adversarial attacks than larger multilingual models across 70 languages, though multilinguality alone does not guarantee security.

Pith tools