{"id":"37c31546-bbe5-4cf0-8e8a-5e8ea701d333","arxiv_id":"2412.03334","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Yankari is a curated 30-million-token monolingual Yoruba dataset built from 13 web sources without religious or machine-translated content.","lead":"This paper introduces Yankari, a new collection of more than 51,000 Yoruba web documents (over 30 million tokens) from 13 sources, covering news, culture, and general knowledge while excluding religious and machine-translated texts. It matters because Yoruba is spoken by over 30 million people but is severely under-resourced in natural language processing.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quality claims rest on single-author manual filtering with no stated criteria, no inter-annotator agreement, and no external audit; the abstract's promised automated evaluations are absent from the body.","rationale":"The reader's weakest assumption correctly identifies the manual native-speaker filter in §4.6 as the least secure load-bearing element. I checked the numbers in Table 1: the document counts sum exactly to 51,407, and the reported average tokens per document is consistent with the total, so the primary statistics are internally coherent. The question is not arithmetic but evidence: the paper's distinctiveness claims depend on having removed machine-translated and religious content, and the only stated mechanism is an undocumented single-author manual review. The abstract explicitly promises automated evaluations that do not appear in the manuscript; in my reading, that mismatch makes the quality claim unsupported rather than false. Since the resource itself may be valuable and the gap is correctable by adding the promised evaluations and a versioned release, the reader's CONDITIONAL verdict remains appropriate. My stress-test does not move the verdict to ACCEPT (evidence still missing), REJECT (no demonstrated fatal flaw), or UNVERDICTED (the paper is evaluable). The cleanest settlement is to keep the conditional verdict and require the concrete quality audit before unconditional acceptance.","tokens_in":5619,"tokens_out":2809,"duration_ms":29998,"concrete_test":"Publish a versioned release (commit hash) and an audit script that (1) runs fastText language identification on a random sample, (2) computes exact and near-duplicate rates, (3) counts specific religious-content signals (e.g., \"Jésù\", \"Bíbélì\", \"Kristi\") in Yankari and in Wura/mC4 Yorùbá subsets, and (4) asks two independent native Yorùbá speakers to blind-annotate a random 500-document sample for machine-translation artifacts using a written rubric, reporting inter-annotator agreement. If the machine-translation flag rate or religious fraction is comparable to existing corpora, the non-religious, manually cleaned claim fails; if the audit rates are low, the central quality claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central value proposition is that Yankari is a large, clean, diverse, ethically sourced Yoruba corpus, specifically avoiding machine-translated and religious content. The only direct support for this is §4.6: filtering was \"primarily a manual process conducted by the author, a native Yoruba speaker, during data spot-checking and review,\" with no explicit exclusion criteria, no inter-annotator validation, no audit trail, and no examples of excluded documents. Meanwhile, the abstract says \"we provide thorough automated evaluations of the dataset, demonstrating its quality compared to existing resources,\" but the body contains no such automated evaluations: §4.8 describes only native-speaker involvement and spot checks, and §6 defers downstream evaluation to future work. The dataset is released without a version or commit hash, making independent re-audit difficult. This is a missing-evidence problem rather than a demonstrated flaw: the corpus could be excellent, but the manuscript as written does not provide the evidence needed to support the load-bearing quality and authenticity claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Yankari, a monolingual Yoruba dataset of 51,407 documents and 30,438,702 tokens collected from 13 web sources. The methodology includes HTML parsing, exact and near-duplicate deduplication, manual filtering to exclude machine-translated and inappropriate content, and an ethical stance against using restricted sources such as JW300. The authors claim Yankari is the first large-scale, non-religious-domain monolingual resource for Yoruba and argue that it addresses quality and ethical problems in existing corpora such as Wura and Yorùbá Text C3. The Abstract promises 'thorough automated evaluations' and comparisons with existing resources, but the body contains no such evaluations: §4.8 reports only manual spot checks, and §6 defers downstream evaluation to future work. The paper is transparent about several limitations, but its central claims of cleanliness, authenticity, and superiority over prior resources are not supported by the evidence presented.","tokens_in":5743,"tokens_out":4231,"duration_ms":41464,"significance":"If the quality and authenticity claims were verified, Yankari would be a valuable addition to Yoruba NLP resources, particularly because it deliberately avoids religious and machine-translated content and provides a source-attributed, JSONL-formatted corpus. The paper's ethical motivation, including the critique of JW300 usage, is a strength, and the acknowledgment of limitations such as internet bias and diacritization issues is commendable. However, the manuscript currently provides no quantitative evaluation, no reproducible audit trail for the manual filtering, no code for the Wura analysis, and no versioned dataset identifier. The resource may be useful, but the paper does not yet meet the evidentiary standard for its own claims.","major_comments":[{"comment":"The Abstract states that the paper provides 'thorough automated evaluations of the dataset, demonstrating its quality compared to existing resources,' but the body does not contain these evaluations. §4.8 describes only native-speaker involvement and spot checks, §4.9 lists qualitative limitations, and §6 explicitly defers downstream evaluation to future work. No quality metrics, no comparisons with Wura or other Yoruba corpora, and no language-model or probing experiments are reported. Because the central claim of the paper is that Yankari is clean, diverse, and superior to existing resources, this missing evidence is load-bearing. The authors should either add the promised automated evaluations or revise the Abstract and Section 1 to reflect what the paper actually provides.","section":"Abstract; §4.8; §6"},{"comment":"The exclusion of machine-translated and inappropriate content is described as 'primarily a manual process conducted by the author, a native Yoruba speaker,' with no explicit exclusion criteria, no inter-annotator agreement, no audit trail, and no examples of excluded documents. Since the paper's value proposition rests on authenticity and quality—'avoiding ... machine-translated content' and 'rigorous quality control'—this procedure must be specified to a reproducible degree. The authors should state the concrete criteria used to identify 'unnatural phrasing' and 'translation artifacts,' report the number of documents removed at each filtering step, and provide a sample of excluded items. Without this, the quality claim cannot be independently checked.","section":"§4.6"},{"comment":"The analysis of the Wura dataset reports specific figures (18.01% of documents contain 'asteroidi', 17,103 unique entries after cleaning, and 45% of the original dataset remaining) but provides no methodology, code, or scripts for how these numbers were obtained. This analysis is used to motivate Yankari's curation approach, so the figures need to be reproducible. Please specify the deduplication method, the cleaning steps, and the exact computation behind each percentage, or release the analysis scripts.","section":"§4.2"},{"comment":"The pipeline description omits the values of the free parameters that determine the dataset's composition: the MinHash/LSH similarity threshold for near-duplicate paragraph removal and the threshold for removing 'very short texts' are never specified. Because these thresholds directly affect the reported 51,407 documents and 30,438,702 tokens, their absence makes the dataset statistics non-reproducible even with access to the same sources. Please state the threshold values and the heuristics used for short-text removal.","section":"§4.3.2 and §4.6"},{"comment":"The dataset is released via a Hugging Face URL, but the paper provides no version identifier, commit hash, or data card describing fields, license, and provenance. Since the paper does not include the full source URL list or processing code, an independent audit cannot confirm that the published artifact matches the described statistics. A persistent versioned identifier and a datasheet should be provided.","section":"Dataset availability; §4.7"}],"minor_comments":[{"comment":"There are inconsistent orthographic renderings of 'Yoruba' (e.g., 'Y oruba' in the title and 'Yorùbá' in §2.1.1); please normalize the orthography throughout the manuscript.","section":"Title; §1"},{"comment":"The numeric columns in Table 1 are not aligned and the 'Total' row uses a space separator; the formatting should be cleaned for readability.","section":"Table 1"},{"comment":"The sample JSON entry contains unusual spacing in the text field ('O ma s e o ! Ijamba oko o f u r u f u gba emi eeyan marun - un') and the dataset URL in the main text contains a space ('Y ANKARI'); these appear to be rendering artifacts that should be fixed.","section":"§4.7; Dataset link"},{"comment":"Several references lack full publication details or consistent formatting (e.g., Adelani et al. 2021, Alabi et al. 2020); please harmonize the reference list and add stable identifiers where available.","section":"References"},{"comment":"The claim of 'first large-scale, non-religious domain monolingual resource' should be qualified, since Wura contains a larger number of Yoruba documents (approximately 68,000), even if it is multilingual and religious-heavy; clarify the comparison criteria. In addition, the 'diverse sources' claim is weakened by the fact that the top three domains account for about 69% of documents; consider reporting an effective diversity metric or discussing this skew.","section":"§3; §4.5"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is positioned as a dataset paper, but the central quality claims are stated without verification. The missing automated evaluations, the unspecified manual filtering criteria, and the lack of a versioned dataset identifier are all fixable within the scope of a revision, so I recommend major revision rather than rejection. The paper would fit a resource-oriented venue well if the authors add the promised evaluations or adjust the claims accordingly. I noticed that no code or analysis scripts are provided despite the quantitative claims about Wura; this should be addressed for reproducibility. There is no indication of improper citation or novelty concealment, though the related-work section could more explicitly compare with other Yoruba corpora beyond those cited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the Yankari corpus is a genuinely new resource that fills a real gap — a large monolingual Yoruba corpus that avoids religious and machine-translated content. The paper also does something useful that isn't advertised: it documents concrete quality problems in the Wura dataset (18% of entries contain 'asteroidi', only 45% unique after cleaning). That analysis is checkable and gives me some confidence the author knows what bad data looks like.\n\nWhat's good: the source selection is transparent (13 domains with counts), the dedup pipeline (exact + MinHash/LSH) is reasonable, and the limitations section is unusually honest for a dataset paper. The dataset is on Hugging Face, so anyone can inspect it.\n\nThe soft spots are about evidence, not the artifact. The abstract promises 'thorough automated evaluations' and quality comparisons against existing resources. The body doesn't deliver those. Section 4.8 describes spot checks and native-speaker review, but no automated metrics, no external benchmark, no inter-annotator agreement. Section 4.6 says machine-translated and inappropriate content were filtered out by the author, a native speaker, but gives no exclusion criteria and no examples. The HF release has no version or commit hash, so re-auditing is harder than it should be. Together these mean the central claim — 'large, clean, diverse, ethically sourced' — is plausible but unverified.\n\nI don't think any of this is fatal. The corpus could easily be as good as claimed; the author just hasn't shown the work. The fix is straightforward: add the missing evaluations (e.g., language-id score, duplicated n-gram rates, comparison with Yorùbá Text C3 and Wura), specify the manual filtering criteria in detail, and pin the dataset version. If the author does that, this is a solid contribution to Yoruba NLP.\n\nWho's this for: people building resources for low-resource languages, and anyone who wants a non-religious Yoruba corpus for language modeling or intrinsic evaluation. It's not a methods paper, and 30 million tokens is modest by LLM standards, but for Yoruba it's a real step up.\n\nBottom line: worth a serious referee. It deserves conditional acceptance, not rejection — the value is in the artifact, and the missing evidence is clearly obtainable. If I were reviewing, I'd ask for a revised version that delivers the promised evaluations and a versioned dataset.","headline":"Useful new Yoruba corpus, but the paper promises more evidence than it delivers; the quality claim needs external checks or clearer criteria.","tokens_in":6312,"tokens_out":1983,"would_cite":true,"duration_ms":17878,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces Yankari, a 51,407-document, 30-million-token monolingual Yoruba corpus from 13 non-religious web sources, and claims it is the first large-scale, ethically curated resource of its kind.","keywords":["Yoruba","monolingual corpus","low-resource NLP","dataset curation","African languages","language modeling","machine translation","text classification"],"falsifier":"An independent audit would settle it: take a random sample of Yankari documents, have at least two native Yoruba speakers who did not build the corpus label each for machine-translation artifacts, non-Yoruba content, and duplicates, and compare their labels with an automatic duplicate detector; if the flagged share approaches the 18% repetition or 24.48% duplication figures reported for Wura or CulturaX's Yoruba subset, or if annotators disagree substantially, the paper's quality claim is refuted.","tokens_in":5348,"feed_emoji":"📚","tokens_out":11414,"duration_ms":92935,"temperature":0.7,"pith_summary":"Yankari is a proposed answer to a concrete gap: Yoruba, spoken by more than 30 million people, has no large, contemporary, non-religious text corpus for NLP. The paper claims to fill that gap with a corpus of 51,407 documents and over 30 million tokens drawn from 13 news, blog, encyclopedia, and government sources. Its central argument is that avoiding religious texts and machine-translated content, and applying dedicated duplication filtering, makes Yankari more balanced, more ethical, and better suited for language modeling and generation than existing Yoruba resources. A careful reader would care because a language's digital future depends on the quality and diversity of the text available to train its models.","feed_headline":"New Yoruba corpus offers 30 million non-religious tokens","feed_subtitle":"Yankari draws 51,407 documents from 13 web sources, filtering out religious and machine-translated text.","key_machinery":"The mechanism that carries the argument is the curation pipeline, not any one algorithm. Source selection chooses 13 publicly accessible domains with an eye to news, blogs, culture, sports, and general knowledge; HTML parsing converts pages to structured text; exact document matching plus MinHash locality-sensitive hashing, a hashing technique that efficiently finds nearly identical text blocks, removes duplicate paragraphs; and a manual native-speaker review, performed by the author, excludes machine-translated, offensive, or non-Yoruba content. Each stage targets a failure documented in prior corpora: religious over-representation, duplication, machine translation, and the legal and ethical restrictions of JW300. The pipeline is what makes Yankari's claimed quality a property of the dataset's construction rather than a matter of luck.","core_discovery":"The paper's core claim is that Yankari is the first large-scale monolingual Yoruba dataset built without the religious skew that dominates prior resources: 51,407 documents, 30,438,702 tokens, an average of 592 tokens per document, across 13 sources led by yo.wikipedia.org, alaroye.org, and BBC Yoruba. The author asserts that by excluding JW300-derived and other restricted content, filtering out machine-translated pages through manual native-speaker review, and removing exact and near-duplicate paragraphs with MinHash and LSH, Yankari supplies a cleaner and more representative sample of written Yoruba. As evidence, the paper reports corpus statistics, domain distribution, and a quality audit of the existing Wura dataset showing 18.01% repetition of a single word, only 45% unique entries after cleaning, and formatting and language errors. The intended consequence is that Yankari can support language modeling, machine translation, text classification, and comparative linguistic studies without the ethical and legal problems tied to religious corpora.","pith_inferences":["Because every source is a website, Yankari will still under-represent spoken Yoruba, informal text-speak, and regional orthographic variants; models trained only on it may miss everyday conversational usage, an extension of the representation biases the paper itself lists.","The per-document URL and source metadata make Yankari a natural seed for benchmark tasks: holding out one domain and testing transfer across the other twelve would measure how domain diversity affects generalization.","The paper's comparison with Wura suggests an independent quantitative check: running the same repetition and duplication analysis on Yankari and publishing the numbers would let the non-religious, low-duplication claim be verified without trusting the curation process.","To make the pipeline reproducible for other languages, the manual exclusion step would need written guidelines and inter-annotator agreement measures; without them, the method cannot be cleanly transplanted."],"forward_implications":["If Yankari is as clean and diverse as claimed, Yoruba NLP gains a training corpus that is not dominated by religious text, which should improve natural language generation and text classification.","Researchers can build on Yankari without the legal and ethical exposure that comes from JW300-derived content, because restricted sources were explicitly excluded.","The size and domain spread make Yankari a credible base for language modeling and for comparative studies of written Yoruba against other corpora.","The documented pipeline, including deduplication and manual filtering steps, provides a template for creating similar ethically sourced datasets for other low-resource languages.","The paper's planned downstream evaluations in language modeling, machine translation, and classification would directly test whether curation translates into better model performance."],"supporting_citations":[{"why":"Introduced Yorùbá Text C3, the prior monolingual Yoruba corpus whose religious skew motivates the need for a non-religious resource.","marker":"Alabi et al. (2020)"},{"why":"Provided JW300, the corpus cited as the source of religious bias and of legal and ethical restrictions in earlier Yoruba datasets.","marker":"Agić and Vulić (2019)"},{"why":"Introduced MENYO-20k, the small multi-domain English-Yoruba translation corpus whose limited size marks the gap Yankari fills.","marker":"Adelani et al. (2021)"},{"why":"Introduced Wura, whose 18% repetition, high duplication, and formatting errors are the empirical evidence for Yankari's stricter curation.","marker":"Oladipo et al. (2023)"},{"why":"Argues that using religious texts in NLP is ethically and legally problematic, grounding Yankari's decision to exclude JW300-derived content.","marker":"Hutchinson (2024)"},{"why":"Shows that data repetition hurts language model performance, justifying Yankari's exact and near-duplicate deduplication steps.","marker":"Hernandez et al. (2022)"},{"why":"Introduced CulturaX, whose Yoruba subset's 24.48% duplication and machine-translated pages serve as the quality baseline Yankari aims to beat.","marker":"Nguyen et al. (2023)"},{"why":"Introduced the AfriBERTa corpus, whose limited domain diversity supports Yankari's push for broader source coverage.","marker":"Ogueji et al. (2021)"}],"fun_headline_variants":["Yankari: 30M Yoruba tokens without religious bias","Clean Yoruba corpus: 51K docs, 30M tokens, 13 sources","First large-scale non-religious Yoruba dataset","Yoruba NLP boost: 30M clean tokens from 13 sources","Yankari dataset: 30M tokens of diverse Yoruba text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The dataset's quality rests on the author's own manual review of documents for machine-translated or inappropriate content, with no published criteria, no second reviewer, and no independent check; if that review is inconsistent, Yankari's advantage over earlier Yoruba corpora is not established.","fun_headline_variants_meta":{"raw":{"variants":["Yankari: 30M Yoruba tokens without religious bias","Clean Yoruba corpus: 51K docs, 30M tokens, 13 sources","First large-scale non-religious Yoruba dataset","Yoruba NLP boost: 30M clean tokens from 13 sources","Yankari dataset: 30M tokens of diverse Yoruba text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0003,"raw_usage":{"total_tokens":1715,"prompt_tokens":912,"completion_tokens":803,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":708}},"tokens_in":528,"tokens_out":803,"duration_ms":6912,"temperature":1.0,"reasoning_tokens":708,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:30:27.537424+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit would settle it: take a random sample of Yankari documents, have at least two native Yoruba speakers who did not build the corpus label each for machine-translation artifacts, non-Yoruba content, and duplicates, and compare their labels with an automatic duplicate detector; if the flagged share approaches the 18% repetition or 24.48% duplication figures reported for Wura or CulturaX's Yoruba subset, or if annotators disagree substantially, the paper's quality claim is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced Yorùbá Text C3, the prior monolingual Yoruba corpus whose religious skew motivates the need for a non-religious resource."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced MENYO-20k, the small multi-domain English-Yoruba translation corpus whose limited size marks the gap Yankari fills."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced Wura, whose 18% repetition, high duplication, and formatting errors are the empirical evidence for Yankari's stricter curation."},{"cited_title":"On topological representation theory from quivers","cited_arxiv_id":"2011.03823","evidence_quote":"Introduced the AfriBERTa corpus, whose limited domain diversity supports Yankari's push for broader source coverage."}],"review_version":1}