REVIEW 4 major objections 6 minor 38 references
SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala
T0 review · 4 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A new 7,044-question benchmark built from Sri Lankan national exams shows the best AI models still score under 68 percent on Sinhala, and far lower on cultural questions.
desk verdict A genuinely useful native Sinhala MMLU resource with a real caveat: the official answer keys were not independently verified, but the released data makes that check possible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the dataset itself: 7,044 multiple-choice questions, each with a question stem, four or five choices, and a single correct answer, plus metadata for subject, difficulty (mapped to school grade), source exam and year, and province. Four annotators manually extracted the questions from OCR'd PDFs of government exam papers hosted on the official e-thaksalawa platform, deduplicated them by exact match and cosine similarity, and organized them into six domains and 30 subjects, following the structure of the original English MMLU. The evaluation protocol—highest-probability option selection for open models, first-token regex extraction for closed models, and prompts following
What would settle it
A human audit of the answer keys: take a random sample of, say, 300 questions across subjects and difficulty levels, have Sinhala-speaking subject teachers independently re-answer them, and measure disagreement with the dataset keys. Near-zero disagreement supports the reported accuracies; a disagreement rate of a few percent would shift every model ranking and domain comparison. A second check: compare model accuracy on the 44 geography questions with the Claude-generated fourth option against the other geography questions; a large divergence would indicate contamination.
Extended reading notes
Core claim
The paper's central claim is that SinhalaMMLU is the first multiple-choice question answering benchmark designed natively for Sinhala, built from official Sri Lankan exam content rather than translated from English, and that on it state-of-the-art LLMs remain far from competent: Claude 3.5 Sonnet reaches 67.65%, GPT-4o 62.95%, the best open-weight model (Qwen2.5-72B-chat) 41.18%, and many open models hover near 22%. Domain analysis shows models struggle most in culturally rich areas, with accuracy dropping 7.8 to 28.2 percentage points on a 1,608-question cultural subset. The paper further shows that native Sinhala STEM questions are rated far more linguistically natural than a translated gl
Load-bearing premise
The dataset's ground-truth answers come from official exam keys transcribed via OCR and manual extraction with no human verification of correctness; if those keys contain systematic errors, every reported accuracy and domain comparison shifts.
Editorial extensions
If this is right
- Sinhala LLM development gets a public, curriculum-aligned yardstick: the paper's grade-level analysis shows that a 40% accuracy—the national passing threshold for the MCQ portion—is barely reached by most open models, so the benchmark can track real exam readiness.
- Translated benchmarks mislead for low-resource languages: the naturalness gap (97.3 vs 71.07) implies that scores on translated Sinhala MMLU do not reflect how Sinhala is actually written in academic settings.
- Culturally grounded knowledge is a measurable, separable deficit: the consistent drop on the 1,608-question cultural subset across strong models defines an explicit target for culturally aware training and evaluation.
- Benchmark format choices materially affect reported capability: converting hard 5-option questions to 4 options improves accuracy by 3.9 to 6.7 points, so cross-benchmark comparisons must control for option count.
- Few-shot prompting does not reliably help instruction-tuned models on Sinhala and sometimes hurts, which argues for zero-shot evaluation as the more stable default in this setting.
Reading between the lines
- Because the paper reports no human verification of answer keys, a targeted audit of a few hundred randomly sampled questions by Sri Lankan exam-subject teachers would independently settle whether the reported accuracies are trustworthy; the paper itself flags this gap in its Limitations section.
- The 44 geography questions whose fourth option was generated by Claude 3.7 Sonnet create a testable contamination check: if models perform unusually well or poorly on exactly those 44 relative to neighboring geography items, the synthetic options may be leaking or confusing signal.
- The benchmark's recipe—harvest official exam papers from a ministry platform, align to curriculum levels, keep content native—transfers directly to other low-resource languages with centralized exam systems, potentially producing a family of curriculum-grounded benchmarks.
- The strong negation and suboption performance gaps suggest a concrete extension: prompting or fine-tuning that explicitly targets negation detection and multi-part option structures could be evaluated directly on this benchmark without any new data collection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces SinhalaMMLU, a 7,044-question multiple-choice benchmark constructed from Sri Lankan national and provincial exam papers (Grades 6–13), covering six domains and 30 subjects, with questions written natively in Sinhala rather than translated from English. The authors evaluate 26 open and closed LLMs in zero- and few-shot settings, reporting that Claude 3.5 Sonnet and GPT-4o achieve 67.65% and 62.95% average accuracy, while open models score much lower, and that models perform worst in culturally grounded domains such as Humanities and Language. Additional analyses address difficulty levels, negation, suboptions, the effect of reducing options from five to four, a comparison of native vs. translated STEM questions, and a handpicked cultural subset.
Significance. If the quality controls hold, SinhalaMMLU is a valuable resource for evaluating LLMs in a low-resource, culturally specific setting. The construction from official exam papers with curriculum alignment directly addresses the gap left by translated multilingual benchmarks, and the public release of the dataset plus evaluation of 26 models are clear strengths. The descriptive findings on scaling, negation, and suboption questions are useful. However, the benchmark's central claims rest on the correctness of the answer keys and on the validity of several smaller comparative studies; these need additional validation before the headline numbers can be taken at face value.
major comments (4)
- [§3.1, §10] Answer-key integrity is the load-bearing assumption. Section 3.1 states that four annotators 'manually extracted MCQs from PDF documents (after OCR)' with no step that verifies the extracted answer against the official key or the official key itself, and Section 10 concedes 'Human evaluation was not conducted.' For Sinhala-script OCR, a label-error rate of even 5–10% would materially shift every accuracy in Table 3 and could compress or invert the reported Humanities-vs-Social-Science gap (66.15 vs. 77.55 for Claude, an 11-point difference). A random-sample human audit of the stored answers, with agreement metrics, is necessary to support the benchmark's validity.
- [§3.3] The 44 geography questions whose missing fourth option was generated by Claude 3.7 Sonnet are a contained but unvalidated instance of LLM involvement in benchmark construction. The paper reports no check that the generated option is not actually correct, does not duplicate another option, and does not alter the intended answer. Since these items are part of the released benchmark, the authors should either remove them, re-derive them with a non-LLM heuristic (e.g., from other government papers), or provide a human-verified validation of each generated option.
- [§7] The naturalness comparison (Table 8) reports scores of 97.30 (SinhalaMMLU) vs. 71.07 (GlobalMMLU-si) based on 100 questions per set rated by two annotators. No inter-annotator agreement metric (e.g., Cohen's kappa), no per-item distribution, and no confidence intervals are reported, so the claim that native content is 'significantly higher' is not statistically supported. The result is used to motivate the entire benchmark's non-translation design, so it needs a more rigorous analysis: agreement, a stratified sampling description, and a test of the difference.
- [§8] The cultural-subset analysis (Table 9) claims a 'consistent performance drop' across models, but the subset is a handpicked 1,608 questions (22%) with no formal selection criteria and no matched non-cultural baseline. The drop of −0.01 for LLaMA-3.1-70B-Chat directly contradicts the 'consistent' phrasing, suggesting the effect is model-dependent and partly a selection artifact. The abstract's conclusion that models 'struggle in culturally rich domains such as the Humanities' should be supported by comparing the cultural subset to a matched set of non-cultural questions of similar difficulty and option count, or by a regression that controls for question length and option count.
minor comments (6)
- [§3.4] Clarify whether the 3-question few-shot set is disjoint from the test set for each subject, and whether the same 3 examples are used for all models and all reported few-shot runs.
- [§6] The phrase 'one incorrect option was randomly removed' should report whether the removal was done once or averaged over multiple random seeds; a single random choice can add noise to the 4-option results in Table 6.
- [§7] Specify the 'linear transformation' that maps the 5-point naturalness scale to a 100-point scale; the current description is ambiguous.
- [Table 3, Table 5] Model names are inconsistent (e.g., 'CLAUDE-3-5-SONNET' vs. 'Claude 3.5-sonnet', 'GPT4O' vs. 'GPT-4O'). Normalize names across tables and the appendix.
- [§10] Section 10 says 'Human evaluation was not conducted,' but Section 7 describes two annotators rating naturalness. Clarify that the naturalness study is a form of human evaluation and explain why it does not constitute the 'domain-expert' evaluation that is missing.
- [§3.1] The OCR pipeline (tool, post-processing, error correction) is not described. A brief account of how OCR errors were handled during manual extraction would help readers assess the answer-key risk.
Circularity Check
No significant circularity: the paper's results are direct measurements on an external benchmark, not derivations from its own construction choices.
full rationale
SinhalaMMLU is a collected benchmark; its evaluation numbers are produced by running LLMs on externally sourced exam questions with official answer keys. There is no fitted parameter that is later renamed a prediction, no defining equation that makes a claimed result true by construction, and no load-bearing self-citation chain. The paper's main conclusions (Claude/GPT-4o accuracy, domain gaps, difficulty trends, negation/suboption effects) are empirical measurements. The cultural-subset analysis (Section 8) stratifies the benchmark by manually annotated content and then measures performance; although the selection could bias the magnitude of the observed drop, this is an annotation/analysis choice, not a case where the output is equivalent to the input by definition. The 44 geography items with a Claude-generated fourth option and the absence of human verification of OCR/answer keys (Section 10) are data-quality and correctness risks that may affect the accuracy numbers, but they do not make the derivation circular. Minor self-citations (e.g., Sakai et al. 2024) are used only to point to related conclusions and are not load-bearing evidence for the paper's central claim.
Assumptions & free parameters
free parameters (1)
- Cosine similarity deduplication threshold =
0.95
assumptions (4)
- domain assumption Official exam answer keys used to label ground truth are correct and unambiguously applicable to the transcribed MCQs.
- domain assumption The training data of evaluated LLMs does not contain the SinhalaMMLU test questions in a form that inflates accuracy.
- ad hoc to paper The 44 geography questions with a fourth option generated by Claude 3.7 Sonnet retain the same difficulty and correctness as native options.
- domain assumption The handpicked cultural subset is a representative sample of culturally rich questions rather than a selection biased toward the hardest items.
Cite this review
Pith. "Pith review of SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala." pith.science (2026). https://pith.science/paper/ZRNW4YSS
@misc{pith2026250903162,
author = {Pith},
title = {Pith review of: SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZRNW4YSS}},
note = {Machine review of arXiv:2509.03162}
}
read the original abstract
Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. While recent multilingual benchmarks attempt to bridge this gap, many rely on automatic translation, which can introduce errors and misrepresent the original cultural context. To address this, we introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language. The dataset includes over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum, and covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluate 26 LLMs on SinhalaMMLU and observe that, while Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies at 67% and 62% respectively, overall model performance remains limited. In particular, models struggle in culturally rich domains such as the Humanities, revealing substantial room for improvement in adapting LLMs to low-resource and culturally specific contexts.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba Oluwadara Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Ijeoma Chukwuneke, Happy Buzaaba, Blessing Kudzaishe Sibanda, Godson Koffi Kalipe, Jonathan Mukiibi, Salomon Kabongo Kabenamualu, Foutse Yuehgoh, Mmasibidi Setaka, Lolwe...
work page 2025
-
[2]
Anthropic. 2024. https://www.anthropic.com/news/claude-3-family Introducing the next generation of Claude . Anthropic News. Accessed: 2025-05-04
work page 2024
-
[3]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. https://doi.org/10.18653/v1/2024.acl-long.44 The belebele benchmark: a parallel reading comprehension dataset in 122 language variants . In Proceedings of the 62nd Annual Meeting of t...
-
[4]
Julen Etxaniz, Gorka Azkune, Aitor Soroa, Oier Lopez de Lacalle, and Mikel Artetxe. 2024. http://papers.nips.cc/paper\_files/paper/2024/hash/3bb42f6bb1b1ab6809afd6c90865b087-Abstract-Datasets\_and\_Benchmarks\_Track.html Bertaqa: How much do language models know about local culture? In Advances in Neural Information Processing Systems 38: Annual Conferenc...
work page 2024
-
[5]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, and 5 others. 2024. https://doi.org/10.5281/zenodo.12608602 The languag...
-
[6]
Omid Ghahroodi, Marzia Nouri, Mohammad V. Sanian, Alireza Sahebi, Doratossadat Dastgheib, Ehsaneddin Asgari, Mahdieh Soleymani Baghshah, and Mohammad Hossein Rohban. 2024. https://api.semanticscholar.org/CorpusID:269033069 Khayyam challenge (persianmmlu): Is your llm truly wise to the persian language? ArXiv, abs/2404.06644
arXiv 2024
-
[7]
Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. https://doi.org/10.18653/v1/2021.findings-acl.413 XL -sum: Large-scale multilingual abstractive summarization for 44 languages . In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 4...
-
[8]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . Proceedings of the International Conference on Learning Representations (ICLR)
work page 2021
Show all 38 references
-
[9]
Hansi Hettiarachchi, Damith Premasiri, Lasitha Randunu Chandrakantha Uyangodage, and Tharindu Ranasinghe. 2024. https://aclanthology.org/2024.lrec-main.1076/ NS ina: A news corpus for S inhala . In Proceedings of the 2024 Joint International Conference on Computational Linguis...
2024
-
[10]
Athuraliya
Kushan Hewapathirana, Nisansa de Silva, and C.D. Athuraliya. 2024. https://link.springer.com/chapter/10.1007/978-3-031-70248-8_17 M2ds: Multilingual dataset for multi-document summarisation . In Advances in Computational Collective Intelligence, pages 219--231, Cham. Springer...
2024 doi
-
[11]
Meng Ji, Meng Ji, Pierrette Bouillon, and Mark Seligman. 2023. https://www.cambridge.org/core/books/translation-technology-in-accessible-health-communication/cultural-and-linguistic-bias-of-neural-machine-translation-technology/0635E9793BA9B604B6BC8D527B91CF25?utm_campaign=sha...
2023
-
[12]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...
2024 arXiv
-
[13]
Fajri Koto, Nurul Aisyah, Haonan Li, and Timothy Baldwin. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.760 Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU . In Proceedings of the 2023 Conference on Empirical Methods i...
2023 doi
-
[14]
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin. 2024. https://doi.org/10.18653/v1/2024.findings-acl.334 A rabic MMLU : Ass...
2024 doi
-
[15]
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2024. https://doi.org/10.18653/v1/2024.findings-acl.671 CMMLU : Measuring massive multitask language understanding in C hinese . In Findings of the Association for Computation...
2024 doi
-
[16]
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Ruofei Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, and 5 others. 2020. https://doi.org/10....
2020 doi
-
[17]
Yang Liu, Meng Xu, Shuo Wang, Liner Yang, Haoyu Wang, Zhenghao Liu, Cunliang Kong, Yun Chen, Yang Liu, Maosong Sun, and Erhong Yang. 2024. https://doi.org/10.48550/ARXIV.2402.13524 Omgeval: An open multilingual generative evaluation benchmark for large language models . CoRR, ...
2024 doi
-
[18]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[19]
Soon Chang Poh, Sze Jue Yang, Jeraelyn Ming Li Tan, Lawrence Leroy Tze Yao Chieng, Jia Xuan Tan, Zhenyu Yu, Foong Chee Mun, and Chee Seng Chan. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.36 M alay MMLU : A multitask benchmark for the low-resource M alay language . I...
2024 doi
-
[20]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025. https://arxiv.org/abs/2412.15115 Qwe...
2025 arXiv
-
[21]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. https://doi.org/10.18653/v1/D16-1264 SQ u AD : 100,000+ questions for machine comprehension of text . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383-...
2016 doi
-
[22]
Tharindu Ranasinghe, Isuri Anuradha, Damith Premasiri, Kanishka Silva, Hansi Hettiarachchi, Lasitha Uyangodage, and Marcos Zampieri. 2024. https://doi.org/10.1007/s10579-024-09723-1 Sold: Sinhala offensive language dataset . Language Resources and Evaluation
2024 doi
-
[23]
Haggag, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Michael Chang, Jenny Chim, Gal Cohen, Aditya Kumar Dalmia, Abraham Diress, and 40 others
Angelika Romanou, Negar Foroutan, Anna Sotnikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Altomare, Mohamed A. Haggag, Snegha A, Alfonso Amayuelas, Azril Hafizi Amirudin, Viraat Aryabumi, Danylo Boiko, Michael Chang, Jenny Chim, Gal Cohen,...
2024 arXiv
-
[24]
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.802 XTREME - R : Towards more challenging and nuanced multilingual eva...
2021 doi
-
[25]
Yusuke Sakai, Hidetaka Kamigaito, and Taro Watanabe. 2024. https://doi.org/10.18653/v1/2024.findings-acl.844 m CSQA : Multilingual commonsense reasoning dataset with unified creation strategy by language models and humans . In Findings of the Association for Computational Ling...
2024 doi
-
[26]
Shivalika Singh, Angelika Romanou, Cl \'e mentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, ...
2025 doi
-
[27]
Shivalika Singh, Freddie Vargus, Daniel D ' souza, B \"o rje Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O ' Mahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Moura, Dominik Krzemi \'n ski, Haki...
2024 doi
-
[28]
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2025. https://aclanthology.org/2025.naacl-long.206/ KMMLU : Measuring massive multitask language understanding in K orean . In Proceedings ...
2025
-
[29]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...
2023 arXiv
-
[30]
Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023. https://doi.org/10.18653/v1/2023.starsem-1.10 Language models are not naysayers: an analysis of language models on negation benchmarks . In Proceedings of the 12th Joint Conference on Lexical and Comput...
2023 doi
-
[31]
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, and Jaydeep Sen. 2025. https://aclanthology.org/2025.naacl-long.507/ MILU : A multi-task I ndic language understanding benchmark . In Proceedings of the 2025 Conference of the Nations of the Americas ...
2025
-
[32]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. https://proceedings.neurips.cc/paper_files/paper/2019/file/4496bf24afe7fab6f046bf4923da8de6-Paper.pdf Superglue: A stickier benchmark for general-purp...
2019
-
[33]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/W18-5446 GLUE : A multi-task benchmark and analysis platform for natural language understanding . In Proceedings of the 2018 EMNLP Workshop B lackbox NLP : A...
2018 doi
-
[34]
Disura Warusawithana, Nilmani Kulaweera, Lakshan Weerasinghe, and Buddhika Karunarathne. 2022. https://aclanthology.org/2022.lrec-1.546/ A systematic approach to derive a refined speech corpus for S inhala . In Proceedings of the Thirteenth Language Resources and Evaluation Co...
2022
-
[35]
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yutong, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Chen...
2025
-
[36]
u ksel, Abdullatif K \
Arda Y \"u ksel, Abdullatif K \"o ksal, L \"u tfi Kerem Senel, Anna Korhonen, and Hinrich Schuetze. 2024. https://doi.org/10.18653/v1/2024.findings-emnlp.413 T urkish MMLU : Measuring massive multitask language understanding in T urkish . In Findings of the Association for Com...
2024 doi
-
[37]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[38]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.