{"id":"cf7e9be2-bba9-438f-9c3c-27a8513f335b","arxiv_id":"2501.11003","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Researchers created and released 30,000 Kiswahili-translated sentences and over 260 hours of speech for three Kenyan languages.","lead":"The paper reports a year-long effort to collect and release parallel text and speech corpora for three under-resourced Kenyan languages: Kidaw'ida, Kalenjin, and Dholuo. It is a practical case study in using paid community contributors and open platforms to build linguistic resources for languages that are largely absent from modern NLP.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim of 30,000 aligned sentence pairs per language and the reported speech hours is externally verifiable, but the paper supplies no in-text or linked evidence; the corpora's availability and exact contents are unconfirmed.","rationale":"The reader's weakest assumption correctly identifies the central load-bearing concern: the paper's primary claims are empirical resource claims, and they rest entirely on the accuracy of the authors' accounting and on the contents of external repositories. My stress test agrees and sharpens the concern into a concrete verifiability check. The paper is a case study and may describe a genuine community effort, but the manuscript itself contains no direct evidence for the 30,000-sentence-per-language figure or for the Table 1 speech statistics. Because the Zenodo record ID is given, the text-corpus claim can be checked directly; because Common Voice is a public platform, the speech claim can also be checked if the language codes or pages are identified. The proposed test would settle whether the concern lands. Since this is the same concern the reader already raised, the verdict should remain CONDITIONAL; no new objection beyond the existing verifiability gap was identified.","tokens_in":8711,"tokens_out":2869,"duration_ms":34526,"concrete_test":"Download the Zenodo record at https://zenodo.org/records/13355021 and programmatically count unique non-empty sentence pairs per language, verifying that at least 30,000 aligned Kidaw'ida-Kiswahili, Kalenjin-Kiswahili, and Dholuo-Kiswahili pairs are present and that both sides are non-empty. Then query the Mozilla Common Voice public dataset for the language codes dav, kln, and luo, and compare validated hours and speaker counts against Table 1. If the counts or repository contents do not match the paper's numbers, the central resource claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is the existence of three parallel corpora (Kidaw'ida-Kiswahili, Kalenjin-Kiswahili, Dholuo-Kiswahili) with 30,000 sentences each, plus the speech data in Table 1. Section 5 states these numbers, and the only evidence offered is a single Zenodo URL and a general reference to Mozilla Common Voice. No sample sentences, file inventory, token counts, sentence-pair counts per language, duplication checks, or license details appear in the manuscript. The Common Voice entries are even harder to audit because the paper does not give the language page URLs or dataset version, and Common Voice distinguishes total recorded hours from validated hours; Table 1 does not say which is reported. If the Zenodo record does not contain 90,000 aligned sentence pairs (30,000 per language), or if the Common Voice datasets do not match Table 1's hours and speaker counts, the central claim fails. This is not an internal logical inconsistency, but it is a load-bearing evidential gap: the scientific claim is about resources that must exist and be usable, and the manuscript does not demonstrate either property.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a one-year, grant-funded project to build three parallel text corpora (Kidaw'ida–Kiswahili, Kalenjin–Kiswahili, Dholuo–Kiswahili) of 30,000 sentences each, together with speech datasets contributed to Mozilla Common Voice. The authors describe a 'selective crowdsourcing' methodology, a quality-assurance process based on Data Collection Leads, and community-engagement challenges. The principal contribution is the claimed public release of these resources for low-resource African language NLP.","tokens_in":8909,"tokens_out":3490,"duration_ms":38693,"significance":"If the resources exist as described, they constitute the first NLP corpora for Kidaw'ida and Kalenjin and an additional Dholuo–Kiswahili parallel corpus, with a non-trivial speech component. The paper's participatory approach, attention to gender balance, and emphasis on open licensing are strengths, and the work addresses a genuine resource gap. However, the paper does not currently substantiate the central empirical claims: the text corpus sizes and speech statistics are asserted without in-paper evidence, making the significance conditional on external verification.","major_comments":[{"comment":"The claim 'We collected 30,000 text sentences for each of the three languages' is not supported within the paper. No file inventory, per-language sentence-pair counts, token or character statistics, duplication checks, or sample entries are provided, and the Zenodo URL alone does not allow a reader to verify the count or the contents of the deposit. Please include a dataset summary table with per-language counts, at least one sample aligned sentence pair per language, and confirmation of the license and file format of the deposited corpus.","section":"Section 5"},{"comment":"The speech data table reports hours and speaker counts for each language on Mozilla Common Voice, but it does not state the dataset version, the language page URLs, or whether the hours are total recorded hours or validated hours. Common Voice distinguishes these two metrics, and without specifying which is reported the numbers cannot be audited. Please state the dataset snapshot and report both total and validated hours and speaker counts.","section":"Table 1"},{"comment":"The quality-assurance process is described only qualitatively: Data Collection Leads 'checked the contributors' data for correct spelling, grammar, fluency, and proper translation.' No numbers are given, such as the number of DCLs per language, the fraction of sentences reviewed, the correction rate, or how disagreements were resolved. Because the paper presents 'selective crowdsourcing' as a methodological contribution, these quantitative details are needed to support the quality claims.","section":"Section 4.1"}],"minor_comments":[{"comment":"The sentence describing Nakatumba-Nabende et al. reads 'A total of sentences of monolingual data for five languages was collected.' The number is missing; please complete the sentence.","section":"Section 3"},{"comment":"In the description of Ogayo et al., 'Dhouo' appears to be a typo for 'Dholuo'; please correct it.","section":"Section 3"},{"comment":"The methodology behind the 'Distribution of NLP activity in Africa' figure is not described. Please specify the data source, the search terms used, and the time period covered.","section":"Figure 1"},{"comment":"The text states 'The same repository is on Github' but no GitHub URL is provided; please add it so readers can access the version-controlled source.","section":"Section 5"},{"comment":"The caption contains a stray space: 'T able 1' should be 'Table 1'.","section":"Table 1 caption"},{"comment":"The sentence 'The issue of cultural appropriateness of data is often cited as militating against using texts of foreign origin, but we propose that such text can catalyse ideas' is presented without a supporting reference; please either cite relevant literature or reframe it as an observation from the project.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper reads more like a project report than a research paper: the central claim is a data release, but the evidence is deferred almost entirely to external repositories. If the editor can arrange an independent check of the Zenodo and Common Voice records, that would substantially strengthen the revision. In addition, the authors might consider adding a small demonstration of the corpus's utility, such as a baseline machine-translation or ASR experiment, which would help establish that the resource is usable as claimed. The writing also needs a careful proofread for incomplete sentences and typos."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Useful news: Kidaw'ida and Kalenjin appear to get their first NLP corpora, and Dholuo gets another parallel resource, built through a year of funded community work. If the resources exist as described, that is a concrete addition. The paper also does well at reporting practical obstacles honestly: the Common Voice localization burden, why Living Dictionaries is not fit for ASR data collection, and the need to pay contributors.\n\nWhat is new is the data, not the method. Selective crowdsourcing with community leads is established practice; the authors cite the relevant work. Their contribution is the application to three Kenyan languages, especially the two without prior resources. They also foreground gender balance and licensing concerns.\n\nThe soft spot is real: the central claim of 30,000 sentences per language and the stated speech hours is unsupported in-paper. No sample sentences, token counts, file inventory, duplication checks, or quality metrics appear. Table 1 does not say whether hours are recorded or validated. The only external pointer is a Zenodo URL; Common Voice pages are not linked. A reader cannot verify the repositories match the numbers. This is not an internal contradiction, but it is a load-bearing evidential gap. The paper's scientific claim is that the resources exist and are usable, and that claim is not demonstrated.\n\nThe paper is otherwise clear and honest, and it acknowledges its own limitations. The authors need to open the data: an appendix with statistics, samples, license details, and direct Common Voice links would fix most of the problem.\n\nI would send this to peer review, but I would ask for that evidence before acceptance. The resource is plausibly valuable; the manuscript as written does not yet demonstrate it. For a reader working on African NLP, the methodology discussion is useful, and the corpora, once verified, worth citing. I would not cite the current version yet.","headline":"The corpora are a real contribution; the manuscript just doesn't show them, and that gap is fixable before publication.","tokens_in":712,"tokens_out":1631,"would_cite":false,"duration_ms":35815,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A year of grassroots collection delivered 30,000 sentences per language and 268 hours of speech for three Kenyan languages, the paper reports.","keywords":["Natural language processing","Low-resource languages","African languages","Corpus building","Crowdsourcing","Kidaw'ida","Kalenjin","Dholuo"],"falsifier":"Open the Zenodo record 13355021 and the Mozilla Common Voice language pages for Kidaw'ida, Kalenjin, and Dholuo, and count the sentence pairs, unique sentences, recording hours, and validated speakers; if any language has substantially fewer than 30,000 sentence pairs or fewer recording hours than the table reports, the paper's central claim fails.","tokens_in":8542,"feed_emoji":"🗣️","tokens_out":5874,"duration_ms":58518,"temperature":0.7,"pith_summary":"The paper reports a one-year, community-based effort to build the first NLP corpora for Kidaw'ida and Kalenjin, and an additional parallel corpus for Dholuo, three under-resourced Kenyan languages. The authors claim they collected 30,000 sentences per language, translated each into Kiswahili, and made the resulting three parallel corpora freely downloadable on Zenodo. They also report 56, 92, and 120 hours of speech data for the three languages on Mozilla Common Voice, with voice collection still ongoing. If the repositories contain what the paper describes, these resources lower the barrier to machine translation, speech recognition, and other NLP applications for languages that currently have almost no digital data.","feed_headline":"Kidaw'ida, Kalenjin, Dholuo get 90,000 sentences, 268 speech hours","feed_subtitle":"Freely downloadable text and voice data aim to power translation and speech tools for three low-resource Kenyan languages.","key_machinery":"The load-bearing mechanism is a two-track collection pipeline: contributors write or transcribe sentences in the target language, bilingual contributors translate them into Kiswahili, and Data Collection Leads—native speakers with high language proficiency—check spelling, grammar, fluency, and translation quality. The text is then uploaded to Mozilla Common Voice, which provides the infrastructure for recording and validating the same sentences across many voices, while the parallel sentences are stored in spreadsheets mirrored on GitHub and Zenodo. This pipeline converts native-speaker availability and local language knowledge into structured parallel text and speech data without relying on web crawling or existing digital sources.","core_discovery":"The paper's central claim is that a small team can create substantial parallel text and speech resources for languages that lack them by recruiting native speakers known to the team, paying small stipends, recording conversations, transcribing them, and translating the results into Kiswahili. The reported result is 30,000 Kidaw'ida–Kiswahili, 30,000 Kalenjin–Kiswahili, and 30,000 Dholuo–Kiswahili sentence pairs, together with 56, 92, and 120 hours of speech recordings on Mozilla Common Voice, with speaker counts of 24, 41, and 44 respectively. The paper states that these are, to its knowledge, the first NLP corpora for Kidaw'ida and Kalenjin, and it makes all of the resources freely available under open licenses so that baseline models can be trained and community expansion can continue.","pith_inferences":["If the reported counts survive independent checks, the paper's main practical lesson is that the bottleneck for low-resource African language NLP is not technical but organizational: trusted bilingual community members, paid modestly, can produce useful corpora in about a year.","The decision to strip code-switched words during transcription, while sensible for clean parallel corpora, may make the data less representative of everyday mixed-language speech, so downstream systems may need separate treatment of code-switching.","Because the paper provides no sample sentences or quality metrics, an immediate extension would be to publish a held-out test set with human-validated references, allowing future work to measure translation and speech-recognition quality consistently.","The selective-crowdsourcing method is unlikely to scale to the millions of sentences required by large language models without layering in automated validation, so the corpora may be more immediately useful for baselines and community tools than for foundation-model pretraining."],"forward_implications":["Baseline machine translation and speech recognition models can now be trained for Kidaw'ida, Kalenjin, and Dholuo using freely downloadable data, giving developers a starting point to improve on.","The corpora give Kidaw'ida and Kalenjin a digital presence they previously lacked, making it possible to build NLP tools for health, agriculture, education, and commerce for their speakers.","The reported gender balance among contributors and Data Collection Leads makes the speech data more likely to represent female voices, which is often missing in low-resource speech datasets.","As the open repositories grow, community members can add more sentences and recordings, which should improve model accuracy over time."],"supporting_citations":[{"why":"Describes the Kencorpus project for Swahili, Dholuo, and Luhya, the main existing Kenyan corpus that this paper builds on and distinguishes its Dholuo data from.","marker":"[13]"},{"why":"MasakhaNER example of community-based annotation for ten African languages, cited as a model for recruiting motivated volunteer annotators.","marker":"[14]"},{"why":"Parallel text and speech corpus building for five East African languages, the closest methodological precedent for translating text into a lingua franca and recording it.","marker":"[15]"},{"why":"Speech synthesis dataset covering eleven African languages, used to compare scale and to justify the recording and voice-talent methods.","marker":"[16]"},{"why":"Llama 2's 2-trillion-token scale, cited to quantify the gap between low-resource corpus sizes and current large language model training demands.","marker":"[17]"}],"fun_headline_variants":["90,000 sentences and 268 hours of speech for 3 Kenyan languages","First corpora for Kidaw'ida and Kalenjin, open to all","How one year built 90,000 sentence pairs for 3 Kenyan languages","Open corpora for Kidaw'ida, Kalenjin, and Dholuo now on Zenodo","From recordings to datasets: 90K pairs for three Kenyan languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the accuracy of the reported counts: that the Zenodo record and Mozilla Common Voice pages actually contain the stated 30,000 sentences per language and the stated speech hours, even though the paper shows no sample sentences, file counts, or quality checks.","fun_headline_variants_meta":{"raw":{"variants":["90,000 sentences and 268 hours of speech for 3 Kenyan languages","First corpora for Kidaw'ida and Kalenjin, open to all","How one year built 90,000 sentence pairs for 3 Kenyan languages","Open corpora for Kidaw'ida, Kalenjin, and Dholuo now on Zenodo","From recordings to datasets: 90K pairs for three Kenyan languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000629,"raw_usage":{"total_tokens":2948,"prompt_tokens":1026,"completion_tokens":1922,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1816}},"tokens_in":642,"tokens_out":1922,"duration_ms":12966,"temperature":1.0,"reasoning_tokens":1816,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:43:17.886172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the Zenodo record 13355021 and the Mozilla Common Voice language pages for Kidaw'ida, Kalenjin, and Dholuo, and count the sentence pairs, unique sentences, recording hours, and validated speakers; if any language has substantially fewer than 30,000 sentence pairs or fewer recording hours than the table reports, the paper's central claim fails.","supporting_citations":[{"cited_title":"In: Wartena, C","cited_arxiv_id":null,"evidence_quote":"Describes the Kencorpus project for Swahili, Dholuo, and Luhya, the main existing Kenyan corpus that this paper builds on and distinguishes its Dholuo data from."},{"cited_title":"Transactions of the Association for Computational Linguistics 9, 1116–1131 (2021)","cited_arxiv_id":null,"evidence_quote":"MasakhaNER example of community-based annotation for ten African languages, cited as a model for recruiting motivated volunteer annotators."},{"cited_title":"Applied AI Letters 5(2), 92 (2024)","cited_arxiv_id":null,"evidence_quote":"Parallel text and speech corpus building for five East African languages, the closest methodological precedent for translating text into a lingua franca and recording it."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"Speech synthesis dataset covering eleven African languages, used to compare scale and to justify the recording and voice-talent methods."}],"review_version":1}