{"id":"9e854f57-fb35-4615-a61f-69ff9a10f843","arxiv_id":"2412.08279","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Y-NQ is a 358-question open-book reading comprehension benchmark for English and Yorùbá, and the paper reports that GPT-4o, o1-mini, and Llama-3.1-8b all perform worse on Yorùbá than on English.","lead":"The authors introduce Y-NQ, a new English-Yorùbá question-answering dataset with 358 human-annotated question-answer pairs built from Natural Questions and Wikipedia. They report that three large language models score lower on Yorùbá than on English, even though the Yorùbá documents are shorter, suggesting that current models do not transfer reading comprehension to this low-resource language.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The x2.5 disparity figure is computed on 4–6 documents with inconsistent denominators; the claim should be re-reported with the full 358-question result and explicit per-bucket counts.","rationale":"The reader identified the validity of ROUGE for Yorùbá morphology as the weakest assumption. That is a real general limitation and the paper itself acknowledges the lack of human evaluation. However, in this stress-test my view is that the more load-bearing problem is the statistical and textual basis of the headline 'x2.5' disparity. The ROUGE validity concern would only change the magnitude of the measured gap, not the existence of some gap, and it would apply equally to both the full-set and the comparable-subset analyses. The much narrower, more fragile basis is the 'x2.5 times' figure: it comes from a subset of 4–6 documents, is reported at document level rather than per-question, and the '2.5' is actually the ROUGE-2 ratio while ROUGE-1 and ROUGE-L ratios are about 1.4–1.6. This is a concrete, checkable inconsistency that directly undermines the headline number as stated. The paper's own full-table results (Table 4) show modest absolute ROUGE gaps of 0.04–0.15, and those differences are the more defensible evidence for 'does not extend to Yorùbá.' But the abstract's specific 'x2.5' claim requires the document-level subset to carry the load, and that subset is too small and too inconsistently described. I do not think this requires rejection; the dataset is a useful contribution and the main qualitative conclusion likely still holds. It does require the authors to re-report the headlined result honestly: per-question paired scores, explicit counts, and confidence intervals, plus a statement that the '2.5x' refers to ROUGE-2 on a handful of documents. Hence CONDITIONAL, in line with the reader's verdict but with a different stated condition.","tokens_in":6104,"tokens_out":1903,"duration_ms":16464,"concrete_test":"Recompute the comparable-length comparison at the per-question level using the released dataset: identify all Yorùbá–English question pairs whose documents are within a defined length band (e.g., both > 900 words, or matched by topic), compute ROUGE-1/ROUGE-2/ROUGE-L for every model per pair, and report the mean difference, the ratio, and a bootstrap confidence interval. If the per-question mean ROUGE-2 ratio is not significantly above 1.5x or the confidence interval crosses 1.0, the 'x2.5' claim does not survive.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—'Yorùbá performance drops by x2.5 times' and 'current English LLMs do not extend to Yorùbá'—rests on a tiny, internally inconsistent subset. The abstract says 'a small set of documents with comparable length' and 'x2.5 times,' but Section 2 says 'a subset of six documents' and Section 3 says 'only 4 documents that are over 900 words long,' with Table 5 labeled 'six comparable English and Yorùbá documents.' The table reports AVG W. 3299 vs 3070, so the six-document average is consistent with the table, but the text says only four exceed 900 words. More importantly, the comparison is at the document level with one ROUGE score per language, not per question; one anomalous Yorùbá document or one very short English document can dominate. The abstract also states 'performance of Yorùbá drops by x2.5 times,' but Table 5 shows ROUGE-1 ratios of 0.45/0.32 = 1.41x, ROUGE-L of 0.30/0.19 = 1.58x, and only ROUGE-2 gives 2.56x (0.23/0.09). The 'x2.5' is the maximum ratio across metrics, not a robust central estimate. The claim would be more defensible if the paper reported per-question paired ROUGE scores with confidence intervals and counts, and if the full-dataset results (Table 4) were emphasized as the primary evidence.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Y-NQ, an English–Yorùbá dataset for open-book reading comprehension and generative question answering, built by extending Natural Questions with 358 human-annotated question–answer pairs over 338 English and 208 Yorùbá Wikipedia documents. The authors evaluate GPT-4o, o1-mini, and Llama-3.1-8b using ROUGE scores, report consistently lower automatic scores for Yorùbá than English, and further claim that on a small set of comparable-length documents Yorùbá performance drops by a factor of 2.5. The paper also reports an annotation by-product: 26 incorrect English Wikipedia answers identified during annotation. The main contribution is the dataset itself, and the experiments are presented as evidence that current English-oriented LLMs do not extend their reading comprehension capabilities to Yorùbá.","tokens_in":6539,"tokens_out":3143,"duration_ms":32454,"significance":"If the central claim holds, Y-NQ is a useful resource for multilingual open-book reading comprehension: it is human-annotated, provides document-level context, includes parallel English–Yorùbá documents on shared topics, and is one of few datasets targeting Yorùbá generative QA rather than multiple-choice span retrieval. The documentation of annotation guidelines, the detailed pre-annotation pipeline, and the finding about English Wikipedia inaccuracies are also valuable. However, the experimental evidence for the headline claim is currently fragile: the strongest quantitative statement relies on at most six documents, and the automatic metric (ROUGE) is not validated for Yorùbá. The dataset and the qualitative direction of the results are solid enough to merit revision, but the paper should not be accepted in its current form without strengthening or carefully re-scoping the performance claims.","major_comments":[{"comment":"The headline claim that 'performance of Yorùbá drops by x2.5 times' is not supported by the reported evidence. Table 5 shows ROUGE-1 and ROUGE-L ratios of 1.41x and 1.58x; only ROUGE-2 gives 2.56x (0.23 vs 0.09). Selecting the maximum ratio across metrics is not a robust central estimate, and the comparison is computed at the document level on six documents, with the text elsewhere referring to 'only 4 documents that are over 900 words long' and Table 5 labeled as 'six comparable English and Yorùbá documents.' The authors should either re-report the comparison per question with counts and confidence intervals, or de-emphasize the x2.5 statement and make the full-dataset results in Table 4 the primary evidence for the disparity.","section":"Abstract; §3, Table 5"},{"comment":"The paper uses ROUGE-1, ROUGE-2, and ROUGE-L as the sole automatic metrics for Yorùbá without any validation that these metrics track answer correctness in a morphologically rich language. Since the conclusion that 'reading comprehension capabilities of current English LLMs do not extend to Yorùbá' depends on this assumption, the authors should provide either a human evaluation of model outputs, a cross-lingual metric validation (e.g., correlation with manual judgments on a sample), or a clear statement that the results are about n-gram overlap rather than comprehension. This concern is reinforced by the Table 4 caption, which says 'Human Score is computed on 358 questions' but no human scores appear anywhere in the paper; the human evaluation should be reported, or the caption corrected.","section":"§3, Evaluation and Table 4"},{"comment":"The sentence 'e.g., losing 0.4 in Rouge-1' is numerically inconsistent with Table 4: the ROUGE-1 gaps are 0.05 (GPT-4o), 0.15 (o1-mini), and 0.11 (Llama-3.1-8b), not 0.4. This appears to be a typo, but it should be corrected because it changes the magnitude of the reported disparity.","section":"§3, Automatic metrics"},{"comment":"There are several internal inconsistencies in the dataset counts that need to be reconciled. The abstract and Table 2 state 358 questions, while §2 states '356 unique questions'; Table 2 lists 338 English and 208 Yorùbá documents, whereas the abstract repeats these numbers, and §2 says '664 Yorùbá documents and 1,566 questions were sent for human annotation.' The authors should clarify the exact final counts, the relationship between the 1,566 annotated questions and the 358 released questions, and why the English and Yorùbá question counts differ if the dataset is parallel.","section":"§2, Dataset statistics; Table 2"}],"minor_comments":[{"comment":"The sentence ending '(d) answers existing in multiple paragraphs in the document for which they annotated the row with all paragraphs where' is incomplete and should be finished to describe what was recorded for such cases.","section":"§2, Annotator findings"},{"comment":"There are several typos and naming inconsistencies: 'Rouge' should be 'ROUGE', 'LlaMA-3.1-8b' should be 'Llama-3.1-8B', 'Bebebele' should be 'Belebele', 'close-book' should be 'closed-book', and 'therevy' should be 'thereby'.","section":"Throughout"},{"comment":"Several citations use the placeholder form 'et al., 2024a' and 'et al., 2024b' without the full author lists; the reference entries should be completed to standard citation style.","section":"References"},{"comment":"The caption for Figure 1 does not state how the equal-size length buckets were constructed, how many documents fall into each bucket, or whether the curves are for GPT-4o only; adding this information would make the length analysis reproducible.","section":"Figure 1"},{"comment":"The dataset availability statement says Y-NQ is 'freely available on HuggingFace' but no URL or dataset identifier is given; this should be included for provenance and reproducibility.","section":"§4, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The dataset is a genuine contribution and the annotation process is described in unusual detail, but the paper currently overstates the strength of its experimental evidence. The inconsistencies around the six-document subset and the missing human evaluation need to be addressed before this can be accepted. The fit with the journal's interests in multilingual and low-resource NLP is good."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Y-NQ is a genuinely useful resource: 358 human-annotated open-book QA pairs for Yorùbá built on NQ, with parallel English documents. If you work on multilingual evaluation, this is worth having. The annotation pipeline is described in useful detail, including the failed SONAR pre-annotation attempt, cleaning of English-contaminated Wikipedia pages, and the discovery of 26 incorrect English Wikipedia answers. The paper is honest about its limits: small size, Wikipedia domain, contamination risk, and the non-comparability of document lengths across languages.\n\nThe main experimental claim is the problem. The abstract's '2.5x drop' comes from a subset described as six documents in one place and four in another. The ROUGE ratios in Table 5 are 1.41x for R-1, 1.58x for R-L, and 2.56x for R-2, so the 2.5x is the maximum metric, not a central estimate. With four to six documents, one anomalous Yorùbá document can dominate. That claim should be re-reported with per-question paired scores and confidence intervals, or dropped in favor of the full-dataset result, which is more robust: consistent Yorùbá underperformance of roughly 0.05–0.15 ROUGE points across all three models.\n\nThe other soft spot is ROUGE itself. No human evaluation is actually reported, despite the Table 4 caption mentioning 'Human Score.' For Yorùbá, ROUGE-1/2/L against a single reference can under-score valid answers that use different word forms or word order. The limitations section acknowledges human evaluation as a possible complement, but it is not included. This leaves the central disparity partly a measurement concern. It is fixable: a small human judgment sample or an analysis of ROUGE failure modes would help.\n\nNone of this undermines the dataset as a contribution. The direction of the finding likely holds—English LLMs do worse in Yorùbá even with shorter documents—but the paper oversells the strength of the evidence. A careful revision that anchors the conclusion in the full-dataset numbers, reports the comparable-length subset with explicit counts and error bars, and adds a small human validation of ROUGE would make this a solid resource paper.\n\nRecommendation: send it to review. With revisions, this is citable and useful.","headline":"A useful Yorùbá QA dataset whose headline 2.5x disparity claim rests on a tiny, inconsistently described subset and on unvalidated ROUGE scores.","tokens_in":6954,"tokens_out":2858,"would_cite":true,"duration_ms":26639,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An English–Yorùbá evaluation dataset shows that current large language models answer open-book reading-comprehension questions worse in Yorùbá than in English, despite Yorùbá documents being far shorter and the task therefore easier.","keywords":["Yoruba","reading comprehension","open-book question answering","text generation","low-resource languages","multilingual evaluation","large language models","ROUGE"],"falsifier":"Randomly sample about 50 Y-NQ questions, have bilingual annotators write two or three acceptable Yorùbá answers per question, and recompute ROUGE-L against the best-matching reference; if Yorùbá scores rise to English parity, the reported gap is a single-reference scoring artifact, while if they stay lower, the comprehension deficit is real.","tokens_in":5939,"feed_emoji":"📖","tokens_out":9400,"duration_ms":86631,"temperature":0.7,"pith_summary":"This paper introduces Y-NQ, a parallel English–Yorùbá dataset for open-book reading comprehension and text generation, built from Wikipedia pages paired with Natural Questions. It asks whether the reading-comprehension abilities of current English-centric large language models transfer to Yorùbá, a low-resource language. Across GPT-4o, o1-mini, and Llama-3.1-8B, the answer is no: Yorùbá scores are consistently lower on ROUGE-1/2/L even though Yorùbá documents average about 430 words versus roughly 10,000 for English, which should make the task easier. On a small matched-length subset, Yorùbá performance is about 2.5 times lower on ROUGE-2, and Yorùbá scores collapse for documents over 1,500 words while English does not. The paper positions the dataset as a reusable benchmark for testing whether multilingual capabilities extend to Yorùbá.","feed_headline":"English LLMs do not extend reading comprehension to Yorùbá","feed_subtitle":"Even with far shorter Yorùbá articles, every tested model scores worse on the same questions.","key_machinery":"The central object is the Y-NQ dataset itself: a curated parallel set of English and Yorùbá Wikipedia documents, each paired with the same questions and with field-level annotations that include the Yorùbá question, short and long answers, a rewrite flag, paragraph location, and a literal-semantic alignment flag. The evaluation machinery is an open-book prompting protocol plus ROUGE-1/2/L scoring, with document-length bucketing and a six-document matched-length subset to control for the confound that Yorùbá documents are shorter. The comparison logic is the load-bearing mechanism: since shorter documents should make answering easier, consistently lower Yorùbá scores under the same protocol constitute evidence of a genuine capability gap rather than task difficulty.","core_discovery":"Y-NQ contains 358 question–answer pairs covering 338 English and 208 Yorùbá Wikipedia articles, with English documents averaging 10,363 words and Yorùbá documents 430 words. The authors compare three large language models under a uniform open-book prompt: read the passage and answer the question in a single paragraph using only the passage. ROUGE evaluation shows Yorùbá behind English for every model and metric, for example GPT-4o ROUGE-1 0.39 versus 0.34, ROUGE-2 0.23 versus 0.19, and ROUGE-L 0.30 versus 0.27; Llama-3.1-8b scores 0.31 versus 0.20 on ROUGE-1. Because Yorùbá documents are much shorter, the task is easier for Yorùbá, so the deficit is evidence that English-oriented reading comprehension does not extend to Yorùbá. Length-bucket analysis shows Yorùbá performance drops sharply for documents around 1,500 words while English stays nearly flat, and on six matched topics with near-equal average length, English ROUGE-2 is more than 2.5 times higher (0.23 versus 0.09). A by-product of annotation was the discovery of 26 incorrect English answers in the source Wikipedia material, underscoring that cross-lingual reference quality matters.","pith_inferences":["The sharp Yorùbá drop beyond 1,500 words may stem from context-window degradation interacting with sparser token representations; this could be tested by measuring accuracy on Yorùbá passages of increasing length against English passages of the same length, and by probing where in the passage the needed answer sits.","ROUGE's n-gram overlap could systematically under-count correct Yorùbá answers, since Yorùbá is morphologically rich and permits flexible word order; a human rating pass or a semantic-similarity metric on the same outputs could change the size of the reported gap even if not its direction.","If the gap is confirmed by human evaluation, the likely mechanism is pretraining data imbalance, which would predict that the gap narrows for languages with more pretraining tokens and widens for morphologically richer low-resource languages; Y-NQ-style datasets for other African languages could test this.","The 26 wrong English Wikipedia answers imply that part of the English–Yorùbá gap might be due to reference errors; extending the same annotation pass to Yorùbá references could reveal whether some Yorùbá failures are actually correct answers to bad or incomplete references."],"forward_implications":["Open-book benchmarks in low-resource languages need controlled or matched document lengths; otherwise shorter documents can mask or amplify the true capability gap.","Developers evaluating multilingual LLMs should expect that high-resource-language reading comprehension does not automatically transfer to Yorùbá, and should test with generative open-book tasks rather than span or multiple-choice only.","The 1,500-word degradation point gives a concrete target: improving long-context handling for Yorùbá may be a separate problem from overall language capability.","Y-NQ provides a reusable test set for future models; a model that closes the Yorùbá–English gap despite Yorùbá's shorter documents would demonstrate genuine cross-lingual reading comprehension.","The annotation finding that 26 English Wikipedia answers were incorrect for Yorùbá-related content indicates that cross-lingual reference sets need verification, and that model errors are not the only source of mismatch."],"supporting_citations":[{"why":"Supplies the source questions and English Wikipedia documents that Y-NQ adapts into a Yorùbá open-book reading-comprehension benchmark.","marker":"Kwiatkowski et al., 2019"},{"why":"Defines the ROUGE metric family used for all automatic comparisons between English and Yorùbá outputs.","marker":"Lin, 2004"},{"why":"Defines the open-book reading-comprehension task format that the paper's prompting and evaluation follow.","marker":"Rajpurkar et al., 2016"},{"why":"The Belebele benchmark is the existing Yorùbá multiple-choice reading-comprehension resource the paper contrasts with generative open-book QA.","marker":"Bandarkar et al., 2024"},{"why":"AfriQA is the prior African-language QA dataset that is open-retrieval rather than open-book, motivating Y-NQ's design.","marker":"Ogundepo et al., 2023"},{"why":"SONAR embeddings were used in the attempted automatic pre-annotation of Yorùbá answer candidates, part of the dataset-creation pipeline.","marker":"Duquenne et al., 2023"},{"why":"The GPT-4o technical report documents one of the three baseline models evaluated on Y-NQ.","marker":"et al., 2024b"},{"why":"The Llama 3 model card documents the open-weight baseline model evaluated on Y-NQ.","marker":"et al., 2024a"}],"fun_headline_variants":["English LLMs fail Yorùbá reading comprehension test","Even short Yorùbá articles stump English LLMs","English LLMs score worse on Yorùbá reading","Yorùbá reading comprehension lags behind English for LLMs","English-centric LLMs flunk Yorùbá comprehension"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes that ROUGE scores against a single reference answer measure reading-comprehension quality in Yorùbá as faithfully as in English, even though Yorùbá's morphology and word order allow many correct answers that share few n-grams with the reference.","fun_headline_variants_meta":{"raw":{"variants":["English LLMs fail Yorùbá reading comprehension test","Even short Yorùbá articles stump English LLMs","English LLMs score worse on Yorùbá reading","Yorùbá reading comprehension lags behind English for LLMs","English-centric LLMs flunk Yorùbá comprehension"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000423,"raw_usage":{"total_tokens":2222,"prompt_tokens":1044,"completion_tokens":1178,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":660,"completion_tokens_details":{"reasoning_tokens":1094}},"tokens_in":660,"tokens_out":1178,"duration_ms":9486,"temperature":1.0,"reasoning_tokens":1094,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:59:06.038886+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample about 50 Y-NQ questions, have bilingual annotators write two or three acceptable Yorùbá answers per question, and recompute ROUGE-L against the best-matching reference; if Yorùbá scores rise to English parity, the reported gap is a single-reference scoring artifact, while if they stay lower, the comprehension deficit is real.","supporting_citations":[],"review_version":1}