Pith. sign in

REVIEW 1 cited by

A Survey of Corpora for Germanic Low-Resource Languages and Dialects

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.09805 v1 pith:2EQIB4NZ submitted 2023-04-19 cs.CL

classification cs.CL
keywords corporalanguagelanguageslow-resourceavailableresearchresourcessurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particular non-standardized low-resource languages. Even within branches of major language families, often considered well-researched, little is known about the extent and type of available resources and what the major NLP challenges are for these language varieties. The first step to address this situation is a systematic survey of available corpora (most importantly, annotated corpora, which are particularly valuable for NLP research). Focusing on Germanic low-resource language varieties, we provide such a survey in this paper. Except for geolocation (origin of speaker or document), we find that manually annotated linguistic resources are sparse and, if they exist, mostly cover morphosyntax. Despite this lack of resources, we observe that interest in this area is increasing: there is active development and a growing research community. To facilitate research, we make our overview of over 80 corpora publicly available. We share a companion website of this overview at https://github.com/mainlp/germanic-lrl-corpora .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Indigenous Languages Spoken in Argentina: A Survey of NLP and Speech Resources

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A survey of Argentina's Indigenous languages and their available computational resources, with speaker estimates drawn from the 2022 national census.

Pith tools