REVIEW 4 major objections 5 minor 3 cited by
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Two Korean benchmarks claim to measure whether LLMs can pass official professional licensure exams, with KMMLU-Redux fixing errors in the original KMMLU and KMMLU-Pro built from fresh government exam PDFs
desk verdict Useful Korean professional-knowledge benchmarks, but the headline pass/fail and contamination-free claims outrun the text-only, n-gram-checked evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the professional licensing system itself: Korean National Technical Qualification (KNTQ) exams for KMMLU-Redux and Korean National Professional Licensure (KNPL) exams for KMMLU-Pro. These are high-stakes exams run annually by the government, so the questions come from an institutionally vetted source rather than crawled from the web. The pipeline uses GPT-4o OCR on official PDFs, human review of tables and images, removal of ambiguous or image-dependent questions, and a decontamination step that n-gram matches the released questions against FineWeb2 and KMMLU. The other key mechanism is scoring against real licensure criteria, where candidates usually need at least 40 percent in every subject and 60 percent overall, which turns accuracy into a pass-fail judgment aligned with human certification.
What would settle it
Run each KMMLU-Pro question through 10-gram and paraphrase-level matching against several large web corpora beyond FineWeb2, including Korean-language forums, Q&A sites, and archived government PDFs; if even a few percent of questions appear in pre-cutoff material, the contamination-free claim fails.
Extended reading notes
Core claim
The central claim is that KMMLU-Pro measures whether LLMs can meet Korean professional certification standards, because it is assembled directly from the most recent official licensure exam PDFs released by the Korean government, OCR-parsed with GPT-4o, then manually reviewed so that images, tables, and answer keys are converted faithfully. The authors find that Claude 3.7 Sonnet with thinking passes 12 of 14 licensure exams, o1 passes 10, while DeepSeek R1 passes 7 among open-weight models. Across domains, models perform well on medicine but fall below the 60 percent average and 40 percent per-subject thresholds in law and tax-accounting; no model passes the Certified Judicial Scrivener or Public Accountant exams. The paper also claims that both benchmarks are contamination-free based on n-gram matching against FineWeb2 and KMMLU, and that KMMLU-Redux correlates almost perfectly with the original KMMLU while being harder and cleaner.
Load-bearing premise
The reliability story depends on n-gram overlap with FineWeb2 and KMMLU being enough to prove that no official exam questions leaked into model training data; paraphrased or non-indexed copies would slip through.
Editorial extensions
If this is right
- KMMLU-Redux provides a harder, cleaner subset of KMMLU, so future model scores on it are less likely to be inflated by leaked answers and duplicate train-test overlap.
- KMMLU-Pro can be updated annually with each year's official exams, giving the benchmark a built-in refresh that web-crawled benchmarks lack.
- License pass rates give an interpretable threshold: a model can have high overall accuracy yet fail a license by being weak in one subject, as o1 does relative to Claude 3.7 Sonnet.
- The large gap between KMMLU-Pro and translated MMMLU on law questions implies that translated benchmarks understate what locally grounded evaluation can measure.
- Reasoning budget helps overall accuracy but not uniformly, so licensure-style evaluation reveals domains where more thinking does not buy correctness.
Reading between the lines
- The contamination-free claim is only as strong as the n-gram check; official exam PDFs could leak into training corpora through paraphrased or transcribed versions that no n-gram match would catch, so a broader contamination audit is a testable next step.
- The annual refresh plan only protects KMMLU-Pro if each year's release is actually withheld until after the benchmark is public; otherwise the newest exam questions could enter pretraining data before evaluation.
- The same official-exam construction could be extended to open-ended and multimodal licensure questions, which the authors explicitly exclude; that would test whether pass rates on multiple-choice questions carry over to constructed-response competence.
- The finding that no evaluated model passes the Judicial Scrivener or Public Accountant exams marks those professions as the hardest near-term barrier for LLM deployment in Korean professional services.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces two Korean expert-level benchmarks: KMMLU-Redux, a cleaned subset of KMMLU with 2,587 questions from Korean National Technical Qualification exams, and KMMLU-Pro, a new benchmark of 2,822 multiple-choice questions from 14 Korean National Professional Licensure exams. The authors evaluate a wide range of LLMs, reporting accuracy, license pass counts based on official thresholds, and analyses of reasoning budget and prompt language. They claim both benchmarks are contamination-free and comprehensively represent Korean industrial knowledge, and they release the datasets publicly.
Significance. If the claims hold, these benchmarks address an important gap in Korean professional knowledge evaluation, using official exam sources with human annotation and a planned annual update cycle. The paper's strengths include the detailed construction pipeline (Section 2.2 and 3.2), the explicit error analysis of KMMLU, the transparency about annotator demographics and compensation (Appendix D.2), and the public release of the datasets. The comparative analysis against MMMLU (Section 6.1) is also valuable. However, as detailed below, several load-bearing claims about pass status, contamination, data provenance, and prompt selection need revision before the benchmark's reliability can be fully accepted.
major comments (4)
- [Section 5.3, Table 2, Appendix E.3] The '# of passed KNPLs' metric in Table 2 and Figure 1 is computed on a text-only, multiple-choice-only subset of each exam, with image questions, descriptive/essay questions, and multiple-answer items excluded as stated in Appendix E.3. The official pass thresholds (40% per subject, 60% average, or the special rules for Judicial Scrivener and Lawyer) are applied to this partial subset without any correction for missing items or item weights. For exams with substantial non-multiple-choice or image-based components (e.g., Physician, Dentist, Lawyer), a model's pass/fail on the subset does not necessarily equal its pass/fail on the official exam. The Limitations section acknowledges this, but the main results and abstract present these counts as unqualified pass/fail outcomes. I recommend renaming the metric (e.g., 'pass on text-based MC subset') or providing a corrected/scaled pass criterion, and stating this caveat wherever pass counts appear in the main text and figures.
- [Section 3.3] The assertion that KMMLU-Pro is 'contamination-free' rests on n-gram matching against FineWeb2 and the KMMLU train/validation sets, which found zero matches. This evidence is insufficient to certify the absence of contamination, because official exam PDFs could appear in other training corpora or in paraphrased form that n-gram matching does not detect. I recommend softening the claim to 'no contamination detected in the checked corpora' and adding a discussion of the limitations of n-gram-based decontamination, including the possibility of near-duplicate or translated leakage.
- [Table 1 caption vs. Section 3.2] Table 1's caption states 'We use KorMedMCQA (Kweon et al., 2024) for three licenses in the Medical category, and KBL (Kim et al., 2024b) for the bar exam of lawyer,' whereas Section 3.2 states 'We directly download the PDF files from the government's websites for each license and use GPT-4o for OCR parsing.' These statements are contradictory. If some KMMLU-Pro questions are drawn from existing datasets rather than from direct government downloads, the provenance claim in the Introduction ('we collect data directly from the official source of each license') is inaccurate, and the paper must clarify how those existing datasets were incorporated and whether they underwent the same human annotation and decontamination procedures.
- [Section 4] The prompt language (Korean vs. English) is selected per model based on the highest average score on the test benchmarks themselves. This is a form of test-set overfitting that can inflate the reported performance numbers and may not reflect a model's behavior under a fixed, pre-specified evaluation protocol. I recommend using a held-out validation set for prompt selection or reporting results for both languages without selecting the better one. At minimum, the authors should explicitly discuss this as a limitation in the main text.
minor comments (5)
- [Section 2.2.1] The text says 'seven smaller LLMs' but the footnote lists eight models (Llama 3.2 3B, Qwen 2.5 3B, Gemma 3 4B IT, Kanana Nano 2.1B, EXAONE 3.5 2.4B, DeepSeek-R1-Distill-Qwen-1.5B, EXAONE Deep 2.4B, and Ko-R1-7B-v2.1). Please correct the count or the list.
- [Figure 1 caption] The caption reads 'licensure exam pass status (indicated by )' with the medal icon missing. Please include the icon or describe the indicator in words.
- [Abstract and title] The name 'KMMLU-R EDUX' contains a space; it should be 'KMMLU-Redux' throughout for consistency.
- [Appendix E.3] The appendix would benefit from a table summarizing the exclusion criteria and the number of questions removed per license, so readers can judge how representative the text-only MC subset is of each official exam.
- [Section 4 and Table 3] Table 3 reports relative differences in scores between English and Korean prompts, but the direction of the difference is not always clear from the 'diff(%)' column; consider adding a note that positive values favor English and negative values favor Korean.
Circularity Check
No significant circularity: KMMLU-Pro is anchored in external official exam PDFs, and the only self-referential element (small-model filtering of KMMLU-Redux) is explicitly disclosed as biased rather than used as a prediction.
full rationale
The paper's derivation chain is not circular. KMMLU-Pro is built by downloading official KNPL exam PDFs from government websites, OCR-parsing them with GPT-4o, and human-reviewing the result (Section 3.2); the pass criteria in Appendix E.3 are taken from official licensing rules, not fitted to the evaluated models. The contamination check (Section 3.3) compares KMMLU-Pro against FineWeb2 and KMMLU by n-gram matching; a zero match is an external empirical result, not an equation that reduces to the dataset's own definition, and the limitation that n-gram matching is not exhaustive is a coverage caveat rather than circularity. The KMMLU-Redux construction does use seven small LLMs to remove 38.6% of 'easy' items (Section 2.2.1), which makes those same models' Redux scores in Table 7 mechanically depressed; however, the paper flags exactly this with a star and states the scores are biased because the models were used for filtration, and the main Table 2 omits those filter models. No hidden prediction is derived from that selection. The Limitations passage also concedes that the text-only MC format cannot fully assess professional competence, so the '# of passed KNPLs' results are explicitly scoped to the benchmark subset rather than a disguised restatement of official outcomes. Self-citations to KMMLU, EXAONE reports, and related works are contextual and not load-bearing for any derivation, so no circularity score is warranted.
Assumptions & free parameters
free parameters (2)
- small_model_easy_threshold =
4 or more of 7 small LLMs
- n_gram_contamination_threshold =
not reported
assumptions (4)
- domain assumption Official government exam PDFs and answer keys provide the correct ground truth for all questions in KMMLU-Pro.
- ad hoc to paper N-gram matching against FineWeb2 and KMMLU is sufficient to certify that KMMLU-Pro is contamination-free.
- domain assumption Pass/fail determinations can be validly computed from the retained text-only multiple-choice subset of each licensure exam.
- ad hoc to paper Agreement by four of seven small LLMs is a valid signal that a question is easy and should be removed from KMMLU-Redux.
Cite this review
Pith. "Pith review of From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation." pith.science (2026). https://pith.science/paper/ACYLMDBJ
@misc{pith2026250708924,
author = {Pith},
title = {Pith review of: From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ACYLMDBJ}},
note = {Machine review of arXiv:2507.08924}
}
read the original abstract
The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applicability in real-world scenarios. In this paper, we introduce two Korean expert-level benchmarks. KMMLU-Redux, reconstructed from the existing KMMLU, consists of questions from the Korean National Technical Qualification exams, with critical errors removed to enhance reliability. KMMLU-Pro is based on Korean National Professional Licensure exams to reflect professional knowledge in Korea. Our experiments demonstrate that these benchmarks comprehensively represent industrial knowledge in Korea. We release our dataset publicly available.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 3 Pith papers
-
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
CAD generation agents are augmented with FEA feedback plus text blueprint and 21-view image signals, raising Box-IoU on S2O and Fusion360 while showing that base models produce no strict-passing FEA artifacts.
-
Raon-Speech Technical Report
A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.
-
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
CAD agents using finite element analysis feedback plus new text blueprint and multi-view image signals improve geometric accuracy on S2O and Fusion360 benchmarks while addressing physical validity gaps in prior genera...
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Hewett, Mojan Javaheripi, Piero Kauffmann, James R
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu,...
arXiv 2024
-
[4]
Anthropic. 2025. https://www.anthropic.com/news/claude-3-7-sonnet Claude 3.7 sonnet and claude code
2025
-
[5]
Zina Chkirbene, Ridha Hamila, Ala Gouissem, and Unal Devrim. 2024. https://doi.org/10.1109/HONET63146.2024.10822885 Large language models (llm) in industry: A survey of applications, challenges, and trends . In 2024 IEEE 21st International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET), pages 229--234
-
[6]
Cohere. 2025. https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 Model card for c4ai command a
work page 2025
-
[7]
John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, Acyr Locatelli, Hangyu Lin, Dwarak Talupuru, Bharat Venk...
arXiv 2024
-
[8]
Google Deepmind. 2024. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/ Introducing gemini 2.0: our new ai model for the agentic era
work page 2024
Show all 68 references
-
[9]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...
2025 arXiv
-
[10]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2025 arXiv
-
[11]
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini...
2025 arXiv
-
[12]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[13]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . In International Conference on Learning Representations
2021
-
[14]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. https://openreview.net/forum?id=rygGQyrFvH The curious case of neural text degeneration . In International Conference on Learning Representations
2020
-
[15]
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. https://openreview.net/forum?id=chfJJYC3iL Livecodebench: Holistic and contamination free evaluation of large language models for code . I...
2025
-
[16]
Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q
Mehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch, Chrysovalantis Anastasiou, Sanket Vaibhav Mehta, Lalit K. Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q. Tran, Q...
2025 arXiv
-
[17]
Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, and Alice Oh. 2024 a . https://aclanthology.org/2024.lrec-main.296/ CLI c K : A benchmark dataset of cultural and linguistic intelligence in K orean . In Proceedings of the 2024 Joint International Conference on Co...
2024
-
[18]
Hyeonwoo Kim, Dahyun Kim, Jihoo Kim, Sukyung Lee, Yungi Kim, and Chanjun Park. 2025. https://arxiv.org/abs/2410.12445 Open ko-llm leaderboard2: Bridging foundational and practical evaluation for korean llms . Preprint, arXiv:2410.12445
2025 arXiv
-
[19]
Yeeun Kim, Youngrok Choi, Eunkyung Choi, JinHwan Choi, Hai Jin Park, and Wonseok Hwang. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.319 Developing a pragmatic benchmark for assessing K orean legal language understanding in large language models . In Findings of th...
2024 doi
-
[20]
Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, volume 35, pages 22199--22213
2022
-
[21]
Sunjun Kweon, Byungjin Choi, Gyouk Chu, Junyeong Song, Daeun Hyeon, Sujin Gan, Jueon Kim, Minkyu Kim, Rae Woong Park, and Edward Choi. 2024. https://arxiv.org/abs/2403.01469 Kormedmcqa: Multi-choice question answering benchmark for korean healthcare professional licensing exam...
2024 arXiv
-
[22]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...
2023
-
[23]
Huiyuan Lai and Malvina Nissim. 2024. https://doi.org/10.18653/v1/2024.acl-long.649 m C o T : Multilingual instruction tuning for reasoning consistency in language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...
2024 doi
-
[24]
Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D
Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...
2025 arXiv
-
[25]
Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, and Jung-Woo Ha. 2023. https://doi.org/10.18653/v1/2023.acl-long.370 SQ u AR e: A large-scale dataset of sensitive question...
2023 doi
-
[26]
Meta. 2024 a . https://ai.meta.com/blog/future-of-ai-built-with-llama/ The future of ai: Built with llama
2024
-
[27]
Meta. 2024 b . https://ai.meta.com/blog/meta-llama-3-1/ Introducing llama 3.1: Our most capable models to date
2024
-
[28]
Meta. 2024 c . https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
2024
-
[29]
Meta. 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ The llama 4 herd: The beginning of a new era of natively multimodal ai innovation
2025
-
[30]
Mistral. 2025. https://mistral.ai/news/mistral-small-3-1 Mistral small 3.1
2025
-
[31]
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. https://arxiv.org/abs/2501.19393 s1: Simple test-time scaling . Preprint, arXiv:2501.19393
2025 arXiv
-
[32]
naver hyperclovax. 2025. https://huggingface.co/naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B Model card for hyperclovax-seed-text-instruct-1.5b
2025
-
[33]
OneLineAI. 2025. Ko-r1-7b-v2.1. https://huggingface.co/OLAIR/ko-r1-7b-v2.1. Accessed: 26 March 2025
2025
-
[34]
OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Pai...
2024 arXiv
-
[35]
OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennet...
2024 arXiv
-
[36]
OpenAI. 2024. https://huggingface.co/datasets/openai/MMMLU Multilingual massive multitask language understanding (mmmlu)
2024
-
[37]
OpenAI. 2025 a . https://openai.com/index/gpt-4-1/ Introducing gpt-4.1 in the api
2025
-
[38]
OpenAI. 2025 b . https://openai.com/index/introducing-o3-and-o4-mini/ Introducing openai o3 and o4-mini
2025
-
[39]
OpenAI. 2025 c . https://openai.com/index/openai-o3-mini/ Openai o3-mini
2025
-
[40]
Chanjun Park, Hyeonwoo Kim, Dahyun Kim, SeongHwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, and Hwalsuk Lee. 2024. https://doi.org/10.18653/v1/2024.acl-long.177 Open K o- LLM leaderboard: Evaluating large language models in K orean with K o-h5 benchmark . In Proceedings of th...
2024 doi
-
[41]
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. 2025. https://arxiv.org/abs/2506.20920 Fineweb2: One pipeline to scale them all -- adapting pre-training da...
2025 arXiv
-
[42]
Wang, Robert Gerbicz, John-Clark Levin, Serguei Popov, Fiona Feng, Steven Y
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Tung Nguyen, Daro...
2025 arXiv
-
[43]
Irene Plaza, Nina Melero, Cristina del Pozo, Javier Conde, Pedro Reviriego, Marina Mayor-Rocher, and María Grandury. 2024. https://arxiv.org/abs/2406.17789 Spanish and llm benchmarks: is mmlu lost in translation? Preprint, arXiv:2406.17789
2024 arXiv
-
[44]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...
2025 arXiv
-
[45]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. https://openreview.net/forum?id=Ti67584b98 GPQA : A graduate-level google-proof q&a benchmark . In First Conference on Language Modeling
2024
-
[46]
LG AI Research, :, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Kyubeen Han, Seokhee Hong, Junwon Hwang, Taewan Hwang, Joonwon Jang, Hyojin Jeon, Kijeong Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Euisoon Kim, Hyosang Kim, Jihoon Kim, Joon...
2025
-
[47]
LG AI Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Yountae Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil...
2024
-
[48]
LG AI Research, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Kijeong Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil...
2025
-
[49]
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. 2024. https://openreview.net/forum?id=m2NVG4Htxs To the cutoff... and beyond? a longitudinal perspective on LLM data contamination . In The Twelfth International Conference on Learning Representations
2024
-
[51]
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Madeline Smith, Antoine Bosselut, Alice Oh, Andre F. T....
2024 arXiv
-
[52]
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024 a . https://arxiv.org/abs/2402.11548 Kmmlu: Measuring massive multitask language understanding in korean . Preprint, arXiv:2402.11548
2024 arXiv
-
[53]
Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jae cheol Lee, Je Won Yeom, Jihyu Jung, Jung woo Kim, and Songseong Kim. 2024 b . https://aclanthology.org/2024.lrec-main.704/ HAE - RAE bench: Evaluation of K orean knowledge in language models . In Proceedings of the 2024 Joint...
2024
-
[54]
Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...
2023
-
[55]
Google Team. 2025 a . https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf Gemma 3 technical report
2025
-
[56]
Kanana LLM Team, Yunju Bak, Hojin Lee, Minho Ryu, Jiyeon Ham, Seungjae Jung, Daniel Wontae Nam, Taegyeong Eo, Donghun Lee, Doohae Jung, Boseop Kim, Nayeon Kim, Jaesun Park, Hyunho Kim, Hyunwoong Ko, Changmin Lee, Kyoung-Woon On, Seulye Baeg, Junrae Cho, Sunghee Jung, Jieun Kan...
2025 arXiv
-
[57]
M-A-P Team, Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, Kang Zhu, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, Kaixing Deng, Shuyue Guo, Shian Jia, Sichao Jiang, Yiyan Liao, Rui Li, Qinrui Li, Sirun Li, Yizhi Li, Yunwen Li, Dehua Ma, Yua...
2025 arXiv
-
[58]
Qwen Team. 2025 b . https://qwenlm.github.io/blog/qwq-32b/ Qwq-32b: Embracing the power of reinforcement learning
2025
-
[59]
Joshua Vendrow, Edward Vendrow, Sara Beery, and Aleksander Madry. 2025. https://arxiv.org/abs/2502.03461 Do large language model benchmarks test reliability? Preprint, arXiv:2502.03461
2025 arXiv
-
[60]
Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. https://arxiv.org/abs/2504.18428 Polymath: Evaluating mathematica...
2025
-
[61]
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024. https://openreview.net/forum?id=y10DM6R2r3 MMLU -pro: A more...
2024
-
[62]
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and M...
2025
-
[63]
xAI. 2025. https://x.ai/news/grok-3 Grok 3 beta — the age of reasoning agents
2025
-
[64]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...
2025 arXiv
-
[65]
Kang Min Yoo, Jaegeun Han, Sookyo In, Heewon Jeon, Jisu Jeong, Jaewook Kang, Hyunwook Kim, Kyung-Min Kim, Munhyong Kim, Sungju Kim, Donghyun Kwak, Hanock Kwak, Se Jung Kwon, Bado Lee, Dongsoo Lee, Gichang Lee, Jooho Lee, Baeseong Park, Seongjin Shin, Joonsang Yu, Seolki Baek, ...
2024 arXiv
-
[66]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...
2019 doi
-
[67]
Yidan Zhang, Yu Wan, Boyi Deng, Baosong Yang, Haoran Wei, Fei Huang, Bowen Yu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. https://arxiv.org/abs/2411.09116 P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms . Preprint, arXiv:2411.09116
2025
-
[68]
Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui, Qinzheng Sun, Shaoguang Mao, Xin Zhang, Ying Xin, Qiufeng Yin, Scarlett Li, and Furu Wei. 2024. https://arxiv.org/abs/2412.15194 Mmlu-cf: A contamination-free multi-task language understanding benchmark . Preprint, arXiv:2412.15194
2024 arXiv
-
[69]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. https://openreview.net/forum?id=VqkAKQibpq SGL ang: Efficient execution of structured language ...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.