Nsanku benchmark shows current LLMs achieve only modest zero-shot translation scores on 43 Ghanaian languages, with no model reaching both high average performance and high cross-language consistency.
Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages
6 Pith papers cite this work, alongside 64 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 2representative citing papers
A seed dataset of 89 Nigerian machinery indicators plus 94 domain-grounded CoT rows, raising domain-grounded prompts from 1/78 to 94/94 and retrieval fidelity to 84/84.
MUDIDI introduces a two-stage LLM pipeline for multilingual dictionary digitization, releases a human-annotated dataset from 30 dictionaries, and shows LLMs outperforming OCR and VLMs on character recognition, markup, and entry segmentation.
Ethnographic study of feminist civic-tech data work argues reparative AI dataset production requires resetting accountability ties to center those harmed by current practices.
AI integration in newsrooms drives internal deferral of judgment to LLMs and external shifts of power to platforms, making fairness, accountability, and transparency harder to sustain unless participatory mechanisms redistribute authority.
citing papers explorer
-
Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages
Nsanku benchmark shows current LLMs achieve only modest zero-shot translation scores on 43 Ghanaian languages, with no model reaching both high average performance and high cross-language consistency.
-
Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer
A seed dataset of 89 Nigerian machinery indicators plus 94 domain-grounded CoT rows, raising domain-grounded prompts from 1/78 to 94/94 and retrieval fidelity to 84/84.
-
MUDIDI: A Two-Stage Framework for Multilingual Dictionary Digitization with Language Models
MUDIDI introduces a two-stage LLM pipeline for multilingual dictionary digitization, releases a human-annotated dataset from 30 dictionaries, and shows LLMs outperforming OCR and VLMs on character recognition, markup, and entry segmentation.
-
Can Data Work be Reparative?
Ethnographic study of feminist civic-tech data work argues reparative AI dataset production requires resetting accountability ties to center those harmed by current practices.
-
FAccT-Checked: A Narrative Review of Authority Reconfigurations and Retention in AI-Mediated Journalism
AI integration in newsrooms drives internal deferral of judgment to LLMs and external shifts of power to platforms, making fairness, accountability, and transparency harder to sustain unless participatory mechanisms redistribute authority.
- Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages