Pith. sign in

REVIEW 6 cited by

Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02353 v2 pith:53SSHP6W submitted 2020-10-05 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesresearchlow-resourcedparticipatorytranslationafricanbenchmarkscase
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Research in NLP lacks geographic diversity, and the question of how NLP can be scaled to low-resourced languages has not yet been adequately solved. "Low-resourced"-ness is a complex problem going beyond data availability and reflects systemic problems in society. In this paper, we focus on the task of Machine Translation (MT), that plays a crucial role for information accessibility and communication worldwide. Despite immense improvements in MT over the past decade, MT is centered around a few high-resourced languages. As MT researchers cannot solve the problem of low-resourcedness alone, we propose participatory research as a means to involve all necessary agents required in the MT development process. We demonstrate the feasibility and scalability of participatory research with a case study on MT for African languages. Its implementation leads to a collection of novel translation datasets, MT benchmarks for over 30 languages, with human evaluations for a third of them, and enables participants without formal training to make a unique scientific contribution. Benchmarks, models, data, code, and evaluation results are released under https://github.com/masakhane-io/masakhane-mt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ModelCitizens: Representing Community Voices in Online Safety

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.

  2. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  3. RelAItionship Building: Analyzing Recruitment Strategies for Participatory AI

    cs.CY 2025-08 conditional novelty 5.0 of 10

    Across 37 participatory AI projects and 5 interviews, the paper finds recruitment practice is under-documented and relationship-driven, and recommends reflexive documentation and institutional support.

  4. Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Researchers created and released 30,000 Kiswahili-translated sentences and over 260 hours of speech for three Kenyan languages.

  5. The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes

    cs.CY 2025-02 conditional novelty 4.0 of 10

    A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.

  6. Task-Oriented Dialog Systems for the Senegalese Wolof Language

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A Rasa-based Wolof task-oriented dialog system, trained on French MASSIVE data projected through an in-house French-Wolof machine translation system, achieves near-French intent classification but weaker slot filling.

Pith tools