REVIEW 6 major objections 6 minor 40 references
Building a Functional Machine Translation Corpus for Kpelle
T0 review · 6 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read First English-Kpelle corpus reaches BLEU 30
desk verdict A useful new Kpelle-English corpus, but the paper's own numbers are inconsistent enough that the headline BLEU scores should be treated as provisional until the dataset is cleaned and re-described. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the dataset itself: 3,234 English-Kpelle translation pairs, 30,021 words total, 4,702 unique Kpelle words, and two corpus versions (V1 with 1,518 pairs and V2 with 2,005 pairs) that isolate the effect of adding data. The same NLLB model is fine-tuned on each version at 10k, 30k, and 60k steps, with evaluation by BLEU and chrF2++ (a character-based metric); the V1-to-V2 comparison in Kpelle-to-English translation is the experiment that carries the claim that the corpus is functional and that more data helps.
What would settle it
Randomly sample 200 held-out sentence pairs, have two additional native Kpelle speakers independently produce reference translations, and recompute BLEU for the released fine-tuned model against the new references; if the scores fall well below the reported range of 24–30, the performance is an artifact of the original reference set. A second check would be to renormalize the entire corpus under a documented orthographic standard and see whether fine-tuning results change materially.
Extended reading notes
Core claim
The central claim is that a carefully assembled bilingual dataset can make Kpelle machine translation work despite the language's low-resource status. Fine-tuning NLLB on Version 2 reaches a best BLEU of 30.28 for Kpelle-to-English at 60,000 training steps and 24.46 for English-to-Kpelle at 30,000 steps on Version 1; these scores are presented as consistent with NLLB-200's reported range for other African languages. The paper also claims that expanding the corpus from 1,518 to 2,005 translation pairs produces measurable gains in the Kpelle-to-English direction, demonstrating the value of data augmentation in a low-resource setting.
Load-bearing premise
The corpus's usefulness rests on the assumption that expert sentence alignments and the tone-marked Latin orthography are consistently accurate across all pairs; if translations are misaligned or tone marks are inconsistent, the reported BLEU scores are not a reliable signal of translation quality.
Editorial extensions
If this is right
- Kpelle becomes a reproducible testbed for low-resource translation: released sentence pairs allow other teams to fine-tune and evaluate models without recollecting data.
- Expanding a low-resource corpus by a few hundred in-domain pairs can move BLEU measurably in one translation direction, suggesting data curation is a high-leverage investment.
- The same corpus can seed speech recognition, language modeling, and sentiment analysis because it provides normalized orthography with tone marks.
- The reported Kpelle scores give future systems a concrete baseline to beat and a point of comparison against NLLB-200 languages such as Wolof, Luo, Yoruba, and Luganda.
Reading between the lines
- If the corpus added Guinean Kpelle variants, translation quality would probably shift; an experiment separating Liberian and Guinean sources could quantify how much the macro-language split costs current models.
- The high number of Kpelle hapax legomena (2,714 words appearing once) suggests the reported BLEU scores may not reflect fluency on rare words; a human evaluation or a rare-word breakdown would test this.
- Because the domain table shows Religion, Health, and Education are thinly covered, targeted expansion of those domains could improve scores more than a generic increase in pair count.
- A useful extension would be to benchmark the same corpus on a second model architecture; divergence between models would reveal which errors come from the data and which from NLLB's inductive biases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a new English–Kpelle parallel corpus, assembled from travel phrasebooks, religious texts, and educational materials, with expert translations and verification. The authors fine-tune Meta's NLLB model on two versions of the corpus and report BLEU scores of up to 30.28 for Kpelle-to-English and 24.46 for English-to-Kpelle, interpreting these as evidence that Kpelle can reach performance levels comparable to other low-resource African languages in NLLB-200. The paper also describes domain statistics, vocabulary analyses, and a roadmap for future expansion.
Significance. If the corpus and evaluation are validated, this is a meaningful contribution: Kpelle is a major Liberian language with very few digital resources, and a publicly released parallel corpus with expert involvement would support machine translation and downstream NLP work. The authors are appropriately explicit about some limitations, and the dataset is placed on Hugging Face. However, the significance is conditional on resolving internal count inconsistencies, verifying orthographic and tonal consistency, and strengthening the evaluation design; the current experiments do not yet establish the headline BLEU claims.
major comments (6)
- [Section 5.1, Section 6.1, Table 2] The dataset size is reported inconsistently: Section 5.1 says 3,234 translation pairs, footnote 2 says Version 2 extends to 2,005 pairs, Section 6.1 gives 2,202 Kpelle and 2,167 English sentences for V2, and Table 2's domain counts sum to 2,005. These numbers cannot all be correct under a one-sentence-per-pair definition. Please define the exact unit (sentence pair versus multi-sentence entry) and reconcile all counts, since the abstract's 'over 2,000 pairs' and Contribution (a)'s '3,234 translation pairs' are in direct tension.
- [Section 4.2.2, Section 5.5, Section 3.1.4] The paper first states that diacritical marks were standardized to represent tonal variations accurately, but Section 5.5 admits 'spelling and tone-marking variations throughout the dataset.' Because Section 3.1.4 shows that tone is lexically contrastive (e.g., lá 'mouth' vs. là 'if'), any unresolved tone variation in the reference data can invalidate BLEU estimates and downstream use. Report a quantitative consistency check, such as inter-annotator agreement on tone marking or a normalized-token audit, and, if residual variation remains, assess its effect on the held-out test set.
- [Section 6.1, Table 3] No zero-shot NLLB scores are reported on the Kpelle held-out sets. Without these baselines, the reported BLEU gains cannot be attributed to fine-tuning or to the corpus; the model may already produce comparable scores before any training on Kpelle. Please add zero-shot NLLB (and ideally another untrained baseline) evaluated on the same test splits.
- [Section 6.1, Section 4] The comparison between V1 and V2 is presented as evidence for the benefit of 'data augmentation,' but the augmentation method is never described, and the two versions differ in size and likely in source composition. This confound makes the V1/V2 comparison uninterpretable as an augmentation study. Define the augmentation procedure explicitly and, if possible, ablate it while holding the underlying source data fixed.
- [Section 6.3, Table 4] The cross-lingual BLEU comparisons are not valid because the scores come from different test sets. The Kpelle numbers are computed on a small in-domain held-out split of this corpus, whereas the cited NLLB-200, M2M-100, and MMTAfrica scores are based on different benchmark test sets. Direct statements such as 'surpasses NLLB-200's lower-bound performances' should be removed or explicitly reframed as non-comparable contextual references; otherwise the claim that Kpelle reaches NLLB-200-level performance is unsupported.
- [Section 6.2, Table 3] The experiments report a single run with no confidence intervals or multiple seeds, and the headline BLEU of 30.28 is the maximum over three step counts. Please report variance across at least three random seeds and state the model-selection rule (e.g., choose the step count on a validation split) before selecting the best result.
minor comments (6)
- [References] The TangaleNLP reference contains what appears to be a typo ('po tangle'); please correct the title and provide a persistent DOI or URL for the dataset repository.
- [Section 5.3] The sentence 'Even though we remove common stop words' is confusing because the listed frequent English words are content words, not stop words; please describe the stop-word list used and align the Kpelle function-word examples with the actual top-frequency list.
- [Table 3] The row label 'NLLB' is ambiguous; use 'NLLB V1' and 'NLLB V2' to distinguish the fine-tuned versions from the un-fine-tuned baseline model.
- [Figure 3] Figure 3(b) shows a single translation example; please state whether it was selected as representative or as a best-case output, and consider adding a few examples from different domains.
- [Section 5.5] The paragraph repeats the same point about underrepresented domains twice; please condense to one statement.
- [Section 3.1.3] The emphatic particle appears as 'b´ e' with a stray accent; please use a consistent diacritic encoding throughout the paper.
Circularity Check
No circularity: the BLEU results come from held-out sacreBLEU evaluation after fine-tuning NLLB, and the only self-citation is a non-load-bearing tone-marking reference.
full rationale
This paper does not claim a first-principles derivation; it constructs a dataset and benchmarks it. The central quantitative claims (BLEU ~30 kpe_Latn→eng_Latn and ~24 eng_Latn→kpe_Latn) are produced by fine-tuning NLLB on training splits and scoring with sacreBLEU on held-out test splits (Section 6.1, Table 3). The evaluation is therefore not circular with respect to the test set by construction. The V1 versus V2 comparison is a simple more-data experiment; although it is confounded by differing corpus sizes and sources, it does not reduce to an identity between input and output. The only citation to the authors' own prior work is [Weako 2024] in Table 1, used to illustrate tone-marking conventions; this citation is not load-bearing for the dataset creation, alignment, normalization, or the reported BLEU scores. Admitted internal inconsistencies, such as the discrepancy between '3,234 translation pairs' in Section 5.1 and the V2 counts of 2,005 pairs or 2,202/2,167 sentences, and the acknowledged 'spelling and tone-marking variations throughout the dataset' in Section 5.5, are data-quality and reproducibility concerns rather than circularity: they do not make the reported metrics equivalent to the input by definition. The claim of being 'the first publicly available English-Kpelle parallel dataset' is an existence claim, not a self-referential derivation. No load-bearing step in the paper's chain reduces a prediction to a fitted parameter, a self-citation, or a redefined known result. The appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Fine-tuning step count (best-of-three selection) =
60k steps (BLEU 30.28 for kpe->eng)
assumptions (4)
- domain assumption The expert translations and sentence alignments are accurate and idiomatic.
- domain assumption The Latin-based orthography was normalized consistently, preserving tonal distinctions.
- domain assumption The NLLB-200 model's kpe_Latn language code corresponds to the same Liberian Kpelle variant used in the corpus.
- standard math SacreBLEU evaluated on the held-out 10% test set is a reliable measure of translation quality for Kpelle.
Cite this review
Pith. "Pith review of Building a Functional Machine Translation Corpus for Kpelle." pith.science (2026). https://pith.science/paper/T25WVEFI
@misc{pith2026250518905,
author = {Pith},
title = {Pith review of: Building a Functional Machine Translation Corpus for Kpelle},
year = {2026},
howpublished = {\url{https://pith.science/paper/T25WVEFI}},
note = {Machine review of arXiv:2505.18905}
}
read the original abstract
In this paper, we introduce the first publicly available English-Kpelle dataset for machine translation, comprising over 2000 sentence pairs drawn from everyday communication, religious texts, and educational materials. By fine-tuning Meta's No Language Left Behind(NLLB) model on two versions of the dataset, we achieved BLEU scores of up to 30 in the Kpelle-to-English direction, demonstrating the benefits of data augmentation. Our findings align with NLLB-200 benchmarks on other African languages, underscoring Kpelle's potential for competitive performance despite its low-resource status. Beyond machine translation, this dataset enables broader NLP tasks, including speech recognition and language modelling. We conclude with a roadmap for future dataset expansion, emphasizing orthographic consistency, community-driven validation, and interdisciplinary collaboration to advance inclusive language technology development for Kpelle and other low-resourced Mande languages.
Figures
Reference graph
Works this paper leans on
-
[1]
Mark Abadi. 2018. https://www.businessinsider.com/travel-language-phrases-to-learn-2018-6 I've been to 25 countries and i can tell you there are only 11 phrases you need to get by anywhere
work page 2018
-
[2]
David Ifeoluwa Adelani, Jesujoba Oluwadara Alabi, Angela Fan, Julia Kreutzer, Xiaoyu Shen, Machel Reid, Dana Ruiter, Dietrich Klakow, Peter Nabende, Ernie Chang, Tajuddeen Gwadabe, Freshia Sackey, Bonaventure F. P. Dossou, Chris Chinenye Emezue, Colin Leong, Michael Beukman, Shamsuddeen Hassan Muhammad, Guyo Dub Jarso, Oreen Yousuf, Andre Niyongabo Rubung...
work page Pith review arXiv doi:10.48550/arxiv.2205.02022 2022
-
[3]
University of Wisconsin-Madison Students in African 671. 2019. https://wisc.pb.unizin.org/lctlresources/chapter/kpelle-history-and-brief-intro-lesson-on-syllabary-and-alphabet/ Kpelle- history and brief intro. Lesson on syllabary and alphabet . Publisher: Pressbooks
work page 2019
-
[4]
Emmanuel Agyei, Xiaoling Zhang, Stephen Bannerman, Ama Bonuah Quaye, Sophyani Banaamwini Yussi, and Victor Kwaku Agbesi. 2024. https://doi.org/10.1007/s10791-024-09451-8 Low resource twi-english parallel corpus for machine translation in multiple domains (twi-2-eng) . Deleted Journal, 27
-
[5]
Adewale Akinfaderin. 2020. https://doi.org/10.18653/v1/2020.winlp-1.38 H ausa MT v1.0: Towards E nglish -- H ausa neural machine translation . In Proceedings of the Fourth Widening Natural Language Processing Workshop, pages 144--147, Seattle, USA. Association for Computational Linguistics
-
[6]
D. Asamoah Owusu, A. Korsah, B. Quartey, S. Nwolley Jnr., D. Sampah, D. Adjepon-Yamoah, and L. Omane Boateng. 2022. Github - ashesi-org/financial-inclusion-speech-dataset: A speech dataset to support financial inclusion created by ashesi university and nokwary technologies with funding from lacuna fund. https://github.com/Ashesi-Org/Financial-Inclusion-Sp...
work page 2022
-
[7]
Gebremeskel , and Abel Aregawi
Asmelash Teka Hadgu , Gebrekirstos G. Gebremeskel , and Abel Aregawi . 2022. https://github.com/asmelashteka/HornMT Machine Translation Benchmark Dataset for Languages in the Horn of Africa . Original-date: 2021-12-05T14:04:38Z
work page 2022
-
[8]
Sara B. 2018. https://www.ef.edu/blog/language/13-important-phrases-know-second-language/ 13 important phrases to know in your second language
work page 2018
Show all 40 references
-
[9]
Gideon George, Olubayo Adekanmbi, and Anthony Soronnadi. 2024. https://openreview.net/forum?id=xLNUQTWsS7 Tangale NLP : Building po tangle to english parallel corpora and machine translation of the tangle (tangale) language . In 5th Workshop on African Natural Language Processing
2024
-
[10]
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc'Aurelio Ranzato, Francisco Guzman, and Angela Fan. 2021. https://arxiv.org/abs/2106.03193 The flores-101 evaluation benchmark for low-resource and multilingual machine t...
2021 arXiv
-
[11]
Heine and M
B. Heine and M. Reh. 1984. https://books.google.com.gh/books?id=36YOAAAAYAAJ Grammaticalization and Reanalysis in African Languages . H. Buske
1984
-
[12]
Maria Konoshenko. 2024. https://doi.org/10.30842/alp23065737193558583 Quotatives in guinean and liberian kpelle: A study of parallel bible corpora and non-biblical texts . Acta Linguistica Petropolitana, 19(3):558--583
2024 doi
-
[13]
Maria Yu Konoshenko. 2008. https://api.semanticscholar.org/CorpusID:148565767 Tonal systems in three dialects of the kpelle language . Mandenkan
2008
-
[14]
Taku Kudo and John Richardson. 2018. https://arxiv.org/abs/1808.06226 Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing . CoRR, abs/1808.06226
2018 arXiv
-
[15]
Siva Subrahamanyam Varma Kusampudi, Anudeep Chaluvadi, and Radhika Mamidi. 2021. https://aclanthology.org/2021.ranlp-1.85 Corpus creation and language identification in low-resource code-mixed T elugu- E nglish text . In Proceedings of the International Conference on Recent Ad...
2021
-
[16]
Leidenfrost and John S
Theodore E. Leidenfrost and John S. McKay. 2005. Kpelle-English Dictionary, with a Grammar Sketch and English-Kpelle Finder List. Language-Literacy-Literature and Bible Translation Center- Lutheran Church in Liberia, Totota
2005
-
[17]
Lonely Planet Global Limited. 2018. https://premiki.si/wp-content/uploads/2018/05/accessible-travel-phrasebook-1.pdf 35 languages covered accessible travel phrasebook
2018
-
[18]
Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, and Francisco Guzman. 2023. https://doi.org/10.18653/v1/2023.acl-long.154 Small data, big impact: Leveraging minimal data for effective machine translation . In Proce...
2023 doi
-
[19]
Joyce Nakatumba-Nabende, Claire Babirye, Peter Nabende, Jeremy Francis Tusubira, Jonathan Mukiibi, Eric Peter Wairagala, Chodrine Mutebi, Tobius Saul Bateesa, Alvin Nahabwe, Hewitt Tusiime, and Andrew Katumba. 2024. https://doi.org/10.1002/ail2.92 Building text and speech benc...
2024 doi
-
[20]
Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Mere...
2020
-
[21]
Vinh Van Nguyen, Ha Nguyen, Huong Thanh Le, Thai Phuong Nguyen, Tan Van Bui, Luan Nghia Pham, Anh Tuan Phan, Cong Hoang-Minh Nguyen, Viet Hong Tran, and Anh Huu Tran. 2022. https://aclanthology.org/2022.lrec-1.588 KC 4 MT : A high-quality corpus for multilingual machine transl...
2022
-
[22]
Iroro Orife, Julia Kreutzer, Blessing Sibanda, Daniel Whitenack, Kathleen Siminyu, Laura Martinus, Jamiil Toure Ali, Jade Abbott, Vukosi Marivate, Salomon Kabongo, Musie Meressa, Espoir Murhabazi, Orevaoghene Ahia, Elan van Biljon, Arshath Ramkilowan, Adewale Akinfaderin, Alp ...
2020 arXiv
-
[23]
Denis Paperno. 2014. https://doi.org/10.4000/mandenkan.563 Sample texts in beng . Mandenkan, pages 106--111
2014 doi
-
[24]
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec, Bettina Messmer, Negar Foroutan, Martin Jaggi, Leandro von Werra, and Thomas Wolf. 2024. https://doi.org/10.57967/hf/3744 Fineweb2: A sparkling update with 1000s of languages
2024 doi
-
[25]
Olivia Christine Perez. 2022. https://www.gooverseas.com/blog/language-phrases-before-travel Helpful language phrases to learn before you travel | go overseas
2022
-
[26]
Matt Post. 2018. https://www.aclweb.org/anthology/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, pages 186--191, Belgium, Brussels. Association for Computational Linguistics
2018
-
[27]
Paul Kanmu Ricks. 2009. Kwaa Pa Kpelee-Woo Maa Kori(We Have Come to Learn Kpelle), first edition. Cuttington University
2009
-
[28]
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...
2022 arXiv
-
[29]
Sharon V Thach. 1981. https://nla.gov.au/nla.cat-vn5443671 A Learner Directed Approach to Kpelle. A Handbook on Communication and Culture with Dialogs, Texts, Cultural Notes, Exercises, Drills and Instructions [microform] / Sharon V. Thach and Others . Distributed by ERIC Clea...
1981
-
[30]
Thach, D.J
S.V. Thach, D.J. Dwyer, and Michigan State University. African Studies Center. 1981. https://books.google.com.gh/books?id=B3oOAAAAYAAJ Kpelle, a Reference Handbook of Phonetics, Grammar, Lexicon and Learning Procedures . [Prepared] for the United States Peace Corps at the Afri...
1981
-
[31]
Valentin Vydrin. 2018. https://doi.org/10.1093/acrefore/9780199384655.013.397 Mande languages
2018
-
[32]
Valentin Vydrin, Jean-Jacques Meric, Kirill Maslinsky, Andrij Rovenchak, Allashera Auguste Tapo, Sebastien Diarra, Christopher Homan, Marco Zampieri, and Michael Leventhal. 2022. Machine learning dataset development for manding languages. url https://github.com/robotsmali-ai/datasets
2022
-
[33]
Alexandra Vydrina. 2017. https://shs.hal.science/tel-03203594 A corpus-based description of Kakabe, a Western Mande language: prosody in grammar . Theses, Institut National des Langues et Civilisations Orientales
2017
-
[34]
Barack Wanjawa, Lilian D. A. Wanzare, Florence Indede, Owen McOnyango, Edward Ombui, and Lawrence Muchemi. 2024. https://doi.org/10.7910/DVN/6N5V1K Kencorpus: Kenyan Languages Corpus
2024 doi
-
[35]
Jackson Weako. 2024. https://doi.org/1005580.ingest.sentry.io/5983515 Libtralo kpelle keyboard help . Keyman.com
2024
-
[36]
Contributors Wikivoyage. 2005. https://en.wikivoyage.org/wiki/Afrikaans_phrasebook West germanic language, spoken in south africa and namibia
2005
-
[37]
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A broad-coverage challenge corpus for sentence understanding through inference . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computation...
2018 doi
-
[38]
Liam G.-Staff Writer. 2017. https://onlineteachersuk.com/english-for-tourism-travel/ English for tourism: Essential uk travel phrases with examples
2017
-
[39]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[40]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.