Pith. sign in

REVIEW 3 cited by

NusaCrowd: Open Source Initiative for Indonesian NLP Resources

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.09648 v4 pith:S3ZD6JZ5 submitted 2022-12-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords indonesianlanguagesnusacrowdinitiativeresourcescreationdatadatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the datasets has been assessed manually and automatically, and their value is demonstrated through multiple experiments. NusaCrowd's data collection enables the creation of the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. Furthermore, NusaCrowd brings the creation of the first multilingual automatic speech recognition benchmark in Indonesian and the local languages of Indonesia. Our work strives to advance natural language processing (NLP) research for languages that are under-represented despite being widely spoken.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A translated temporal reasoning dataset plus a cross-lingual example retriever that outperforms semantic-alignment baselines on low-resource language temporal questions.

  2. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

    cs.CL 2024-11 conditional novelty 6.0 of 10

    INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.

  3. The Multilingual Divide and Its Impact on Global AI Safety

    cs.AI 2025-05 conditional novelty 3.0 of 10

    The language gap in AI models creates safety disparities across languages, and closing it requires funding multilingual datasets, transparency, and research.

Pith tools