REVIEW 6 cited by
Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Research in NLP lacks geographic diversity, and the question of how NLP can be scaled to low-resourced languages has not yet been adequately solved. "Low-resourced"-ness is a complex problem going beyond data availability and reflects systemic problems in society. In this paper, we focus on the task of Machine Translation (MT), that plays a crucial role for information accessibility and communication worldwide. Despite immense improvements in MT over the past decade, MT is centered around a few high-resourced languages. As MT researchers cannot solve the problem of low-resourcedness alone, we propose participatory research as a means to involve all necessary agents required in the MT development process. We demonstrate the feasibility and scalability of participatory research with a case study on MT for African languages. Its implementation leads to a collection of novel translation datasets, MT benchmarks for over 30 languages, with human evaluations for a third of them, and enables participants without formal training to make a unique scientific contribution. Benchmarks, models, data, code, and evaluation results are released under https://github.com/masakhane-io/masakhane-mt.
Forward citations
Cited by 6 Pith papers
-
ModelCitizens: Representing Community Voices in Online Safety
A community-annotated toxicity dataset with conversational context shows that models trained on ingroup labels outperform state-of-the-art moderation APIs.
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.
-
RelAItionship Building: Analyzing Recruitment Strategies for Participatory AI
Across 37 participatory AI projects and 5 interviews, the paper finds recruitment practice is under-documented and relationship-driven, and recommends reflexive documentation and institutional support.
-
Building low-resource African language corpora: A case study of Kidawida, Kalenjin and Dholuo
Researchers created and released 30,000 Kiswahili-translated sentences and over 260 hours of speech for three Kenyan languages.
-
The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes
A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.
-
Task-Oriented Dialog Systems for the Senegalese Wolof Language
A Rasa-based Wolof task-oriented dialog system, trained on French MASSIVE data projected through an in-house French-Wolof machine translation system, achieves near-French intent classification but weaker slot filling.
Discussion (0). Continue with ORCID to comment.