{"id":"1c5064fd-e67d-4b05-a123-53c3af1e9014","arxiv_id":"2509.03162","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new 7,044-question native Sinhala exam benchmark shows the best LLM at 67.65% accuracy, with large drops on culturally specific subjects.","lead":"The paper introduces SinhalaMMLU, a benchmark of over 7,000 multiple-choice questions in Sinhala drawn from official Sri Lankan school exams, and tests 26 language models on it. The best model, Claude 3.5 Sonnet, scores 67.65% while all models do worst on culturally specific subjects such as Sinhala history and drama.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Answer-key integrity is the load-bearing assumption: with no human verification of OCR/transcribed answers, even a modest error rate in Sinhala-script official keys could shift every reported accuracy and the claimed Humanities deficit.","rationale":"The reader's weakest_assumption—that the official answer keys survive OCR and manual transcription without systematic errors—is also the most load-bearing assumption in the paper. Every headline number is an accuracy against those keys; every domain comparison (Humanities vs Social Science, negation, suboptions, 5-option vs 4-option, cultural subset) inherits their correctness. The paper provides no verification of labels, and Section 10 admits human evaluation was not conducted. This is not a manufactured concern: Sinhala script OCR and manual transcription of a 7,044-question dataset drawn from PDFs is exactly the kind of pipeline where key-swap, glyph confusion, or misread answer letters can occur, and the paper reports no quality-control statistics (e.g., annotator agreement, re-check rate, spot audits).\n\nI do not think the paper should be rejected: the dataset is public, the pipeline is described, and the concern is empirically addressable. The conditional verdict is appropriate because a single audit could either confirm the labels or show that the benchmark needs correction. I agree with the reader's identification of the weak point, and I see no reason to move the verdict: the concern strengthens the case for CONDITIONAL rather than ACCEPT, which is exactly where the reader landed.\n\nI considered other potential concerns. Contamination from public exam papers is real but less decisive because if extensive memorization had occurred, even open models would likely score much higher than the low-20s they actually achieve; contamination alone would not explain the observed pattern. The naturalness comparison in Section 7 is based on only 200 questions and two annotators with no reported agreement, but that is a secondary analysis, not the central benchmark claim. The ethics offer of co-authorship is unusual but does not bear on the scientific validity of the measurements. The absence of human evaluation remains the central vulnerability, and it is testable.","tokens_in":25133,"tokens_out":2419,"duration_ms":33436,"concrete_test":"Independently audit the answer keys and question text. Stratified random sample of ~300 questions (oversampling Humanities and the 44 geography items with LLM-generated options), have two Sinhala-speaking subject-matter experts independently verify each stored answer against the original source PDF/marking scheme and flag OCR/transcription mismatches or ambiguous/wrong official keys; compute inter-annotator agreement. Then recompute accuracies for at least Claude 3.5 Sonnet, GPT-4o, and Qwen2.5-72B-Chat after excluding or correcting all disputed items. If accuracy changes by more than ~2 points overall or the Humanities gap versus Social Science changes by more than ~5 points, the reported benchmark numbers and the 'culturally rich domains' conclusion need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—that SinhalaMMLU measures genuine capability, that Claude 3.5 Sonnet reaches 67.65%, and that models 'struggle' in Humanities—all presuppose that the stored correct answers are the true answers. Section 3.1 states that four annotators 'manually extracted MCQs from PDF documents (after OCR)' with instructions to include questions with 'an identified correct answer,' but no step verifies that the extracted answer key matches the official key or that the official key itself is correct. Section 10 explicitly concedes: 'Human evaluation was not conducted in this study due to practical limitations.'\n\nOCR on Sinhala script is error-prone, especially for visually similar glyphs and diacritics, and culturally rich Humanities items (history, literature, drama) are precisely where OCR and transcription errors and ambiguous answer keys are most likely. A systematic label-error rate of even 5–10% would materially change model accuracies and could compress or invert the reported Humanities-vs-Social-Science gap (77.55% vs 66.15% for Claude), because the gap itself is only about 11 points. The 44 geography items given a Claude-generated fourth option (Section 3.3) are a smaller, contained instance of the same problem: if that generated option is sometimes correct or duplicates another option, those 44 labels are unreliable.\n\nContamination from public exam papers is also unaddressed, but it is less load-bearing here: the low open-model accuracies (mostly low-20s) are not what one would expect if the models had memorized the source exams, so contamination would most plausibly affect only the highest closed models. Label integrity, by contrast, has no such indirect argument in its favor and is the foundation of every number in Tables 3–6 and 8–9.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SinhalaMMLU, a 7,044-question multiple-choice benchmark constructed from Sri Lankan national and provincial exam papers (Grades 6–13), covering six domains and 30 subjects, with questions written natively in Sinhala rather than translated from English. The authors evaluate 26 open and closed LLMs in zero- and few-shot settings, reporting that Claude 3.5 Sonnet and GPT-4o achieve 67.65% and 62.95% average accuracy, while open models score much lower, and that models perform worst in culturally grounded domains such as Humanities and Language. Additional analyses address difficulty levels, negation, suboptions, the effect of reducing options from five to four, a comparison of native vs. translated STEM questions, and a handpicked cultural subset.","tokens_in":25388,"tokens_out":4140,"duration_ms":49864,"significance":"If the quality controls hold, SinhalaMMLU is a valuable resource for evaluating LLMs in a low-resource, culturally specific setting. The construction from official exam papers with curriculum alignment directly addresses the gap left by translated multilingual benchmarks, and the public release of the dataset plus evaluation of 26 models are clear strengths. The descriptive findings on scaling, negation, and suboption questions are useful. However, the benchmark's central claims rest on the correctness of the answer keys and on the validity of several smaller comparative studies; these need additional validation before the headline numbers can be taken at face value.","major_comments":[{"comment":"Answer-key integrity is the load-bearing assumption. Section 3.1 states that four annotators 'manually extracted MCQs from PDF documents (after OCR)' with no step that verifies the extracted answer against the official key or the official key itself, and Section 10 concedes 'Human evaluation was not conducted.' For Sinhala-script OCR, a label-error rate of even 5–10% would materially shift every accuracy in Table 3 and could compress or invert the reported Humanities-vs-Social-Science gap (66.15 vs. 77.55 for Claude, an 11-point difference). A random-sample human audit of the stored answers, with agreement metrics, is necessary to support the benchmark's validity.","section":"§3.1, §10"},{"comment":"The 44 geography questions whose missing fourth option was generated by Claude 3.7 Sonnet are a contained but unvalidated instance of LLM involvement in benchmark construction. The paper reports no check that the generated option is not actually correct, does not duplicate another option, and does not alter the intended answer. Since these items are part of the released benchmark, the authors should either remove them, re-derive them with a non-LLM heuristic (e.g., from other government papers), or provide a human-verified validation of each generated option.","section":"§3.3"},{"comment":"The naturalness comparison (Table 8) reports scores of 97.30 (SinhalaMMLU) vs. 71.07 (GlobalMMLU-si) based on 100 questions per set rated by two annotators. No inter-annotator agreement metric (e.g., Cohen's kappa), no per-item distribution, and no confidence intervals are reported, so the claim that native content is 'significantly higher' is not statistically supported. The result is used to motivate the entire benchmark's non-translation design, so it needs a more rigorous analysis: agreement, a stratified sampling description, and a test of the difference.","section":"§7"},{"comment":"The cultural-subset analysis (Table 9) claims a 'consistent performance drop' across models, but the subset is a handpicked 1,608 questions (22%) with no formal selection criteria and no matched non-cultural baseline. The drop of −0.01 for LLaMA-3.1-70B-Chat directly contradicts the 'consistent' phrasing, suggesting the effect is model-dependent and partly a selection artifact. The abstract's conclusion that models 'struggle in culturally rich domains such as the Humanities' should be supported by comparing the cultural subset to a matched set of non-cultural questions of similar difficulty and option count, or by a regression that controls for question length and option count.","section":"§8"}],"minor_comments":[{"comment":"Clarify whether the 3-question few-shot set is disjoint from the test set for each subject, and whether the same 3 examples are used for all models and all reported few-shot runs.","section":"§3.4"},{"comment":"The phrase 'one incorrect option was randomly removed' should report whether the removal was done once or averaged over multiple random seeds; a single random choice can add noise to the 4-option results in Table 6.","section":"§6"},{"comment":"Specify the 'linear transformation' that maps the 5-point naturalness scale to a 100-point scale; the current description is ambiguous.","section":"§7"},{"comment":"Model names are inconsistent (e.g., 'CLAUDE-3-5-SONNET' vs. 'Claude 3.5-sonnet', 'GPT4O' vs. 'GPT-4O'). Normalize names across tables and the appendix.","section":"Table 3, Table 5"},{"comment":"Section 10 says 'Human evaluation was not conducted,' but Section 7 describes two annotators rating naturalness. Clarify that the naturalness study is a form of human evaluation and explain why it does not constitute the 'domain-expert' evaluation that is missing.","section":"§10"},{"comment":"The OCR pipeline (tool, post-processing, error correction) is not described. A brief account of how OCR errors were handled during manual extraction would help readers assess the answer-key risk.","section":"§3.1"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially a valuable community resource, and the construction pipeline is described in enough detail to be reproducible. The main risk is the missing answer-key validation: without a sample audit, the headline accuracies and domain comparisons are not fully trustworthy. I would encourage the editors to ask for at least a 100–200 item human verification of the stored answers and a description of disagreements. Also, the authors should double-check their 'first Sinhala MCQ benchmark' claim against MILU (Verma et al., 2025) and Include (Romanou et al., 2024), which may already contain Sinhala subsets; if so, the novelty framing should be adjusted rather than repeated as absolute."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper builds something the Sinhala NLP community actually lacks: a native, curriculum-aligned MCQ benchmark, 7,044 questions across 30 subjects sourced from Sri Lankan national and provincial exams, not translated from English. The dataset is public on GitHub and Hugging Face, and the paper evaluates 26 models with several secondary analyses (negation, suboptions, 5- vs 4-option, native vs translated STEM naturalness, cultural subset). That is a real contribution, and the reproducible artifact is a large part of it: anyone can rerun or audit it. The headline numbers (Claude 3.5 Sonnet 67.65%, GPT-4o 62.95%, open models mostly in the low 20s) are plausible and consistent with the low-resource language literature.\n\nThe load-bearing soft spot is answer-key integrity. Section 3.1 says annotators manually extracted MCQs from OCR'd PDFs, and Section 10 explicitly says no human evaluation was conducted. If OCR or transcription introduced systematic errors in the Sinhala-script keys, every accuracy and domain gap shifts, especially the Humanities vs Social Science gap, which is only about 11 points for Claude. That said, this is a conditional risk, not a demonstrated flaw, and the released data means the field can verify it. The 44 geography questions with a Claude-generated fourth option are a smaller version of the same concern; the paper says the generated option was \"not used during model evaluation,\" which is ambiguous and worth clarifying. The naturalness comparison (Section 7) uses only 100 questions per condition with two annotators and no agreement metric; the 97.3 vs 71.07 difference is suggestive but not solid. The cultural subset (Section 8) is handpicked and the drop is partly an artifact of selection, though the size of the drop for closed models makes the qualitative point.\n\nMinor points: no contamination check against public exam papers is reported, and the ethics section's offer of co-authorship to data-collection volunteers is unusual and should be clarified. None of these are fatal. The central claim, that a native benchmark reveals poor LLM performance on culturally grounded Sinhala content, holds up in outline.\n\nThis paper deserves a serious referee. The resource is valuable, the analysis is honest about many limitations, and the community can check the key assumptions. I would send it to review with a request to either add a small human verification of a sample of answer keys or clearly frame the benchmark as provisionally validated and subject to audit. My own verdict would be conditional accept with revisions.\n\nWho benefits: Sinhala NLP researchers, multilingual benchmark builders, and anyone studying cultural knowledge in LLMs. I would cite it and probably bring it to a reading group.\n\nRecommendation: engage with it; it is a worthwhile contribution with addressable soft spots.","headline":"A genuinely useful native Sinhala MMLU resource with a real caveat: the official answer keys were not independently verified, but the released data makes that check possible.","tokens_in":26041,"tokens_out":2112,"would_cite":true,"duration_ms":22140,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new 7,044-question benchmark built from Sri Lankan national exams shows the best AI models still score under 68 percent on Sinhala, and far lower on cultural questions.","keywords":["Sinhala","LLM benchmark","MMLU","low-resource language","multiple-choice QA","cultural knowledge","Sri Lankan national curriculum","multilingual evaluation"],"falsifier":"A human audit of the answer keys: take a random sample of, say, 300 questions across subjects and difficulty levels, have Sinhala-speaking subject teachers independently re-answer them, and measure disagreement with the dataset keys. Near-zero disagreement supports the reported accuracies; a disagreement rate of a few percent would shift every model ranking and domain comparison. A second check: compare model accuracy on the 44 geography questions with the Claude-generated fourth option against the other geography questions; a large divergence would indicate contamination.","tokens_in":24944,"feed_emoji":"📝","tokens_out":5653,"duration_ms":54736,"temperature":0.7,"pith_summary":"This paper introduces SinhalaMMLU, a benchmark of 7,044 multiple-choice questions drawn from Sri Lankan national and provincial examination papers (grades 6 through A-Level), written natively in Sinhala rather than translated. The authors argue this is the first Sinhala-specific multitask language understanding benchmark and that it measures what a translated benchmark cannot: curriculum-aligned, culturally grounded content such as Sri Lankan history, Sinhala literature, indigenous dance, and oriental music. Across 26 large language models, the best performers—Claude 3.5 Sonnet at 67.65% and GPT-4o at 62.95%—leave substantial room for improvement, and every model loses ground on culturally specific questions. If the benchmark is valid, it gives the Sinhala NLP community a standardized, curriculum-grounded yardstick and concrete evidence that translation-based multilingual benchmarks understate how much cultural and technical knowledge models lack.","feed_headline":"Top AI scores under 68% on new Sinhala curriculum benchmark","feed_subtitle":"A native 7,000-question exam benchmark shows even the best models lag, especially on Sri Lankan cultural knowledge.","key_machinery":"The central object is the dataset itself: 7,044 multiple-choice questions, each with a question stem, four or five choices, and a single correct answer, plus metadata for subject, difficulty (mapped to school grade), source exam and year, and province. Four annotators manually extracted the questions from OCR'd PDFs of government exam papers hosted on the official e-thaksalawa platform, deduplicated them by exact match and cosine similarity, and organized them into six domains and 30 subjects, following the structure of the original English MMLU. The evaluation protocol—highest-probability option selection for open models, first-token regex extraction for closed models, and prompts following","core_discovery":"The paper's central claim is that SinhalaMMLU is the first multiple-choice question answering benchmark designed natively for Sinhala, built from official Sri Lankan exam content rather than translated from English, and that on it state-of-the-art LLMs remain far from competent: Claude 3.5 Sonnet reaches 67.65%, GPT-4o 62.95%, the best open-weight model (Qwen2.5-72B-chat) 41.18%, and many open models hover near 22%. Domain analysis shows models struggle most in culturally rich areas, with accuracy dropping 7.8 to 28.2 percentage points on a 1,608-question cultural subset. The paper further shows that native Sinhala STEM questions are rated far more linguistically natural than a translated gl","pith_inferences":["Because the paper reports no human verification of answer keys, a targeted audit of a few hundred randomly sampled questions by Sri Lankan exam-subject teachers would independently settle whether the reported accuracies are trustworthy; the paper itself flags this gap in its Limitations section.","The 44 geography questions whose fourth option was generated by Claude 3.7 Sonnet create a testable contamination check: if models perform unusually well or poorly on exactly those 44 relative to neighboring geography items, the synthetic options may be leaking or confusing signal.","The benchmark's recipe—harvest official exam papers from a ministry platform, align to curriculum levels, keep content native—transfers directly to other low-resource languages with centralized exam systems, potentially producing a family of curriculum-grounded benchmarks.","The strong negation and suboption performance gaps suggest a concrete extension: prompting or fine-tuning that explicitly targets negation detection and multi-part option structures could be evaluated directly on this benchmark without any new data collection."],"forward_implications":["Sinhala LLM development gets a public, curriculum-aligned yardstick: the paper's grade-level analysis shows that a 40% accuracy—the national passing threshold for the MCQ portion—is barely reached by most open models, so the benchmark can track real exam readiness.","Translated benchmarks mislead for low-resource languages: the naturalness gap (97.3 vs 71.07) implies that scores on translated Sinhala MMLU do not reflect how Sinhala is actually written in academic settings.","Culturally grounded knowledge is a measurable, separable deficit: the consistent drop on the 1,608-question cultural subset across strong models defines an explicit target for culturally aware training and evaluation.","Benchmark format choices materially affect reported capability: converting hard 5-option questions to 4 options improves accuracy by 3.9 to 6.7 points, so cross-benchmark comparisons must control for option count.","Few-shot prompting does not reliably help instruction-tuned models on Sinhala and sometimes hurts, which argues for zero-shot evaluation as the more stable default in this setting."],"supporting_citations":[{"why":"Supplies the MMLU structure and the prompt instruction format the benchmark and evaluation follow.","marker":"(Hendrycks et al., 2021)"},{"why":"The translated GlobalMMLU-Sinhala resource the paper contrasts with and rates for linguistic naturalness.","marker":"(Singh et al., 2025)"},{"why":"Supplies the highest-probability selection method used to score open-source models.","marker":"(Koto et al., 2023)"},{"why":"Supplies the CMMLU evaluation method and prior negation findings that the negation analysis extends.","marker":"(Li et al., 2024)"},{"why":"The MILU benchmark whose few-shot finding on instruction-tuned models the paper's few-shot results corroborate.","marker":"(Verma et al., 2025)"},{"why":"Prior evidence that negation challenges language models, framing the negation analysis.","marker":"(Truong et al., 2023)"},{"why":"MalayMMLU, cited as the evaluation pipeline the paper aligns with.","marker":"(Poh et al., 2024)"}],"fun_headline_variants":["First native Sinhala AI benchmark: best model hits 67%","New 7k-question Sinhala exam stumps top AI models","AI stumbles on Sri Lankan cultural questions in new benchmark","Best LLM scores 67% on first native Sinhala benchmark","Open-sourced models lag far behind on Sinhala exam benchmark"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The dataset's ground-truth answers come from official exam keys transcribed via OCR and manual extraction with no human verification of correctness; if those keys contain systematic errors, every reported accuracy and domain comparison shifts.","fun_headline_variants_meta":{"raw":{"variants":["First native Sinhala AI benchmark: best model hits 67%","New 7k-question Sinhala exam stumps top AI models","AI stumbles on Sri Lankan cultural questions in new benchmark","Best LLM scores 67% on first native Sinhala benchmark","Open-sourced models lag far behind on Sinhala exam benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000628,"raw_usage":{"total_tokens":2750,"prompt_tokens":762,"completion_tokens":1988,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":1895}},"tokens_in":506,"tokens_out":1988,"duration_ms":15374,"temperature":1.0,"reasoning_tokens":1895,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:06:06.677146+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A human audit of the answer keys: take a random sample of, say, 300 questions across subjects and difficulty levels, have Sinhala-speaking subject teachers independently re-answer them, and measure disagreement with the dataset keys. Near-zero disagreement supports the reported accuracies; a disagreement rate of a few percent would shift every model ranking and domain comparison. A second check: compare model accuracy on the 44 geography questions with the Claude-generated fourth option against the other geography questions; a large divergence would indicate contamination.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the MMLU structure and the prompt instruction format the benchmark and evaluation follow."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The MILU benchmark whose few-shot finding on instruction-tuned models the paper's few-shot results corroborate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MalayMMLU, cited as the evaluation pipeline the paper aligns with."}],"review_version":1}