Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Two Korean benchmarks claim to measure whether LLMs can pass official professional licensure exams, with KMMLU-Redux fixing errors in the original KMMLU and KMMLU-Pro built from fresh government exam PDFs

desk verdict Useful Korean professional-knowledge benchmarks, but the headline pass/fail and contamination-free claims outrun the text-only, n-gram-checked evidence. read the letter →

arxiv 2507.08924 v2 pith:ACYLMDBJ submitted 2025-07-11 cs.CL cs.AI

classification cs.CLcs.AI
keywords LLMevaluationKoreanbenchmarkprofessionallicensureexamsKMMLU-ProKMMLU-Reduxdatacontaminationmultiple-choicequestionansweringNLP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper introduces two Korean-language benchmarks that test whether LLMs hold professionally certified knowledge that matters on the job, not just academic knowledge. KMMLU-Redux is a cleaned subset of the existing KMMLU benchmark, restricted to Korean National Technical Qualification exams, with leaked answers, ill-posed questions, notation errors, and clarity errors removed. KMMLU-Pro is built from official Korean National Professional Licensure exams in 14 regulated professions, including lawyer, CPA, and physician, with evaluation scored against real pass criteria. The authors report that leading models pass most medical licensure exams, struggle badly on law and tax-accounting licenses, and that no evaluated model passes the Judicial Scrivener or Public Accountant exams. If the benchmarks are sound, they offer a practical, refreshable, and contamination-resistant way to measure whether a model could support certified professional work in Korea.

What carries the argument

The load-bearing mechanism is the professional licensing system itself: Korean National Technical Qualification (KNTQ) exams for KMMLU-Redux and Korean National Professional Licensure (KNPL) exams for KMMLU-Pro. These are high-stakes exams run annually by the government, so the questions come from an institutionally vetted source rather than crawled from the web. The pipeline uses GPT-4o OCR on official PDFs, human review of tables and images, removal of ambiguous or image-dependent questions, and a decontamination step that n-gram matches the released questions against FineWeb2 and KMMLU. The other key mechanism is scoring against real licensure criteria, where candidates usually need at least 40 percent in every subject and 60 percent overall, which turns accuracy into a pass-fail judgment aligned with human certification.

What would settle it

Run each KMMLU-Pro question through 10-gram and paraphrase-level matching against several large web corpora beyond FineWeb2, including Korean-language forums, Q&A sites, and archived government PDFs; if even a few percent of questions appear in pre-cutoff material, the contamination-free claim fails.

Watch

Extended reading notes

Core claim

The central claim is that KMMLU-Pro measures whether LLMs can meet Korean professional certification standards, because it is assembled directly from the most recent official licensure exam PDFs released by the Korean government, OCR-parsed with GPT-4o, then manually reviewed so that images, tables, and answer keys are converted faithfully. The authors find that Claude 3.7 Sonnet with thinking passes 12 of 14 licensure exams, o1 passes 10, while DeepSeek R1 passes 7 among open-weight models. Across domains, models perform well on medicine but fall below the 60 percent average and 40 percent per-subject thresholds in law and tax-accounting; no model passes the Certified Judicial Scrivener or Public Accountant exams. The paper also claims that both benchmarks are contamination-free based on n-gram matching against FineWeb2 and KMMLU, and that KMMLU-Redux correlates almost perfectly with the original KMMLU while being harder and cleaner.

Load-bearing premise

The reliability story depends on n-gram overlap with FineWeb2 and KMMLU being enough to prove that no official exam questions leaked into model training data; paraphrased or non-indexed copies would slip through.

Editorial extensions

If this is right

  • KMMLU-Redux provides a harder, cleaner subset of KMMLU, so future model scores on it are less likely to be inflated by leaked answers and duplicate train-test overlap.
  • KMMLU-Pro can be updated annually with each year's official exams, giving the benchmark a built-in refresh that web-crawled benchmarks lack.
  • License pass rates give an interpretable threshold: a model can have high overall accuracy yet fail a license by being weak in one subject, as o1 does relative to Claude 3.7 Sonnet.
  • The large gap between KMMLU-Pro and translated MMMLU on law questions implies that translated benchmarks understate what locally grounded evaluation can measure.
  • Reasoning budget helps overall accuracy but not uniformly, so licensure-style evaluation reveals domains where more thinking does not buy correctness.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The contamination-free claim is only as strong as the n-gram check; official exam PDFs could leak into training corpora through paraphrased or transcribed versions that no n-gram match would catch, so a broader contamination audit is a testable next step.
  • The annual refresh plan only protects KMMLU-Pro if each year's release is actually withheld until after the benchmark is public; otherwise the newest exam questions could enter pretraining data before evaluation.
  • The same official-exam construction could be extended to open-ended and multimodal licensure questions, which the authors explicitly exclude; that would test whether pass rates on multiple-choice questions carry over to constructed-response competence.
  • The finding that no evaluated model passes the Judicial Scrivener or Public Accountant exams marks those professions as the hardest near-term barrier for LLM deployment in Korean professional services.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces two Korean expert-level benchmarks: KMMLU-Redux, a cleaned subset of KMMLU with 2,587 questions from Korean National Technical Qualification exams, and KMMLU-Pro, a new benchmark of 2,822 multiple-choice questions from 14 Korean National Professional Licensure exams. The authors evaluate a wide range of LLMs, reporting accuracy, license pass counts based on official thresholds, and analyses of reasoning budget and prompt language. They claim both benchmarks are contamination-free and comprehensively represent Korean industrial knowledge, and they release the datasets publicly.

Significance. If the claims hold, these benchmarks address an important gap in Korean professional knowledge evaluation, using official exam sources with human annotation and a planned annual update cycle. The paper's strengths include the detailed construction pipeline (Section 2.2 and 3.2), the explicit error analysis of KMMLU, the transparency about annotator demographics and compensation (Appendix D.2), and the public release of the datasets. The comparative analysis against MMMLU (Section 6.1) is also valuable. However, as detailed below, several load-bearing claims about pass status, contamination, data provenance, and prompt selection need revision before the benchmark's reliability can be fully accepted.

major comments (4)
  1. [Section 5.3, Table 2, Appendix E.3] The '# of passed KNPLs' metric in Table 2 and Figure 1 is computed on a text-only, multiple-choice-only subset of each exam, with image questions, descriptive/essay questions, and multiple-answer items excluded as stated in Appendix E.3. The official pass thresholds (40% per subject, 60% average, or the special rules for Judicial Scrivener and Lawyer) are applied to this partial subset without any correction for missing items or item weights. For exams with substantial non-multiple-choice or image-based components (e.g., Physician, Dentist, Lawyer), a model's pass/fail on the subset does not necessarily equal its pass/fail on the official exam. The Limitations section acknowledges this, but the main results and abstract present these counts as unqualified pass/fail outcomes. I recommend renaming the metric (e.g., 'pass on text-based MC subset') or providing a corrected/scaled pass criterion, and stating this caveat wherever pass counts appear in the main text and figures.
  2. [Section 3.3] The assertion that KMMLU-Pro is 'contamination-free' rests on n-gram matching against FineWeb2 and the KMMLU train/validation sets, which found zero matches. This evidence is insufficient to certify the absence of contamination, because official exam PDFs could appear in other training corpora or in paraphrased form that n-gram matching does not detect. I recommend softening the claim to 'no contamination detected in the checked corpora' and adding a discussion of the limitations of n-gram-based decontamination, including the possibility of near-duplicate or translated leakage.
  3. [Table 1 caption vs. Section 3.2] Table 1's caption states 'We use KorMedMCQA (Kweon et al., 2024) for three licenses in the Medical category, and KBL (Kim et al., 2024b) for the bar exam of lawyer,' whereas Section 3.2 states 'We directly download the PDF files from the government's websites for each license and use GPT-4o for OCR parsing.' These statements are contradictory. If some KMMLU-Pro questions are drawn from existing datasets rather than from direct government downloads, the provenance claim in the Introduction ('we collect data directly from the official source of each license') is inaccurate, and the paper must clarify how those existing datasets were incorporated and whether they underwent the same human annotation and decontamination procedures.
  4. [Section 4] The prompt language (Korean vs. English) is selected per model based on the highest average score on the test benchmarks themselves. This is a form of test-set overfitting that can inflate the reported performance numbers and may not reflect a model's behavior under a fixed, pre-specified evaluation protocol. I recommend using a held-out validation set for prompt selection or reporting results for both languages without selecting the better one. At minimum, the authors should explicitly discuss this as a limitation in the main text.
minor comments (5)
  1. [Section 2.2.1] The text says 'seven smaller LLMs' but the footnote lists eight models (Llama 3.2 3B, Qwen 2.5 3B, Gemma 3 4B IT, Kanana Nano 2.1B, EXAONE 3.5 2.4B, DeepSeek-R1-Distill-Qwen-1.5B, EXAONE Deep 2.4B, and Ko-R1-7B-v2.1). Please correct the count or the list.
  2. [Figure 1 caption] The caption reads 'licensure exam pass status (indicated by )' with the medal icon missing. Please include the icon or describe the indicator in words.
  3. [Abstract and title] The name 'KMMLU-R EDUX' contains a space; it should be 'KMMLU-Redux' throughout for consistency.
  4. [Appendix E.3] The appendix would benefit from a table summarizing the exclusion criteria and the number of questions removed per license, so readers can judge how representative the text-only MC subset is of each official exam.
  5. [Section 4 and Table 3] Table 3 reports relative differences in scores between English and Korean prompts, but the direction of the difference is not always clear from the 'diff(%)' column; consider adding a note that positive values favor English and negative values favor Korean.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: KMMLU-Pro is anchored in external official exam PDFs, and the only self-referential element (small-model filtering of KMMLU-Redux) is explicitly disclosed as biased rather than used as a prediction.

full rationale

The paper's derivation chain is not circular. KMMLU-Pro is built by downloading official KNPL exam PDFs from government websites, OCR-parsing them with GPT-4o, and human-reviewing the result (Section 3.2); the pass criteria in Appendix E.3 are taken from official licensing rules, not fitted to the evaluated models. The contamination check (Section 3.3) compares KMMLU-Pro against FineWeb2 and KMMLU by n-gram matching; a zero match is an external empirical result, not an equation that reduces to the dataset's own definition, and the limitation that n-gram matching is not exhaustive is a coverage caveat rather than circularity. The KMMLU-Redux construction does use seven small LLMs to remove 38.6% of 'easy' items (Section 2.2.1), which makes those same models' Redux scores in Table 7 mechanically depressed; however, the paper flags exactly this with a star and states the scores are biased because the models were used for filtration, and the main Table 2 omits those filter models. No hidden prediction is derived from that selection. The Limitations passage also concedes that the text-only MC format cannot fully assess professional competence, so the '# of passed KNPLs' results are explicitly scoped to the benchmark subset rather than a disguised restatement of official outcomes. Self-citations to KMMLU, EXAONE reports, and related works are contextual and not load-bearing for any derivation, so no circularity score is warranted.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central benchmark-quality claims rely on assumptions about official source reliability, the sufficiency of the contamination check, the validity of pass thresholds applied to a filtered subset, and the difficulty-filtering rule. KMMLU-Redux's construction also uses a hand-chosen model-agreement threshold, and the contamination check does not disclose its n-gram settings.

free parameters (2)
  • small_model_easy_threshold = 4 or more of 7 small LLMs
    Section 2.2.1 marks a question as easy if at least four of seven small LLMs answer correctly, removing 38.6% of KMMLU when building Redux. This hand-chosen threshold directly shapes the difficulty and composition of KMMLU-Redux.
  • n_gram_contamination_threshold = not reported
    Sections 2.1 and 3.3 apply n-gram contamination detection but do not report the n value or match threshold, so the zero-contamination result cannot be independently scoped or reproduced.
assumptions (4)
  • domain assumption Official government exam PDFs and answer keys provide the correct ground truth for all questions in KMMLU-Pro.
    Section 3.2 states the official PDFs let the authors guarantee the correctness of answer labels. If official keys are wrong or human or OCR edits introduce errors, labels can be incorrect.
  • ad hoc to paper N-gram matching against FineWeb2 and KMMLU is sufficient to certify that KMMLU-Pro is contamination-free.
    Section 3.3 reports zero contaminated examples and concludes contamination-free integrity. This treats one detection method and one corpus as a complete guarantee.
  • domain assumption Pass/fail determinations can be validly computed from the retained text-only multiple-choice subset of each licensure exam.
    Appendix E.3 computes license pass status using official 40/60 thresholds after excluding image-based, descriptive, and multi-answer questions; this assumes the excluded questions do not change the pass outcome.
  • ad hoc to paper Agreement by four of seven small LLMs is a valid signal that a question is easy and should be removed from KMMLU-Redux.
    Section 2.2.1 uses this to remove 38.6% of the dataset, treating small-model accuracy as a difficulty measure rather than an independent expert standard.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation." pith.science (2026). https://pith.science/paper/ACYLMDBJ

@misc{pith2026250708924,
  author       = {Pith},
  title        = {Pith review of: From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ACYLMDBJ}},
  note         = {Machine review of arXiv:2507.08924}
}
read the original abstract

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applicability in real-world scenarios. In this paper, we introduce two Korean expert-level benchmarks. KMMLU-Redux, reconstructed from the existing KMMLU, consists of questions from the Korean National Technical Qualification exams, with critical errors removed to enhance reliability. KMMLU-Pro is based on Korean National Professional Licensure exams to reflect professional knowledge in Korea. Our experiments demonstrate that these benchmarks comprehensively represent industrial knowledge in Korea. We release our dataset publicly available.

Figures

Figures reproduced from arXiv: 2507.08924 by the authors.

Figure 1
Figure 1. Performance of leading reasoning models developed by diverse groups on industrial knowledge for [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Performance differences in LLMs on the erro [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Performance of four LLMs on {Medical(left), Accounting(center), Law(right)}-relevant subsets from the MMMLU (Korean) (OpenAI, 2024) and KMMLU-PRO. While the discrepancies in scores are narrow in the medicine domain, they are wider in law-related problems, emphasizing the need for datasets that reflecting real professional knowledge in Korea. that relatively smaller models (<20B) are able to pass in the medicine doma… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Performance of LLMs on KMMLU and KMMLU-REDUX. A high ρ value indicates a strong correlation between the results of the two benchmarks. 6.2 KMMLU vs KMMLU-REDUX [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Reasoning budget results of Qwen3-32B and Claude 3.7 Sonnet on [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: The prompt is used for error type annotation. Each sample is annotated as an error if the respective field [PITH_FULL_IMAGE:figures/full_fig_p023_6.png]
Figure 7
Figure 7. Figure 7: Domain distribution of problems in KMMLU￾REDUX. The total size of the dataset is 2,587. to save costs because we do not need experts for annotations nor the answer relabeling to avoid risk of data from online (Gema et al., 2025; Team et al., 2025b). Before annotation, …
Figure 8
Figure 8. Figure 8: Excerpt from the translated annotation guidelines for converting PDF documents into structured text. We [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: The English prompt used for evaluating LLMs on our [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]
Figure 10
Figure 10. Figure 10: The Korean prompt used for evaluating LLMs on our [PITH_FULL_IMAGE:figures/full_fig_p025_10.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback

    cs.GR 2026-05 unverdicted novelty 6.0 of 10

    CAD generation agents are augmented with FEA feedback plus text blueprint and 21-view image signals, raising Box-IoU on S2O and Fusion360 while showing that base models produce no strict-passing FEA artifacts.

  2. Raon-Speech Technical Report

    cs.CL 2026-04 conditional novelty 5.5 of 10

    A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.

  3. Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback

    cs.GR 2026-05 unverdicted novelty 5.0 of 10

    CAD agents using finite element analysis feedback plus new text blueprint and multi-view image signals improve geometric accuracy on S2O and Fusion360 benchmarks while addressing physical validity gaps in prior genera...

Reference graph

Works this paper leans on

68 extracted references · 25 canonical work pages · cited by 2 Pith papers

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Hewett, Mojan Javaheripi, Piero Kauffmann, James R

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee, Yuanzhi Li, Weishung Liu, Caio C. T. Mendes, Anh Nguyen, Eric Price, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Xin Wang, Rachel Ward, Yue Wu, Dingli Yu,...

  4. [4]

    Anthropic. 2025. https://www.anthropic.com/news/claude-3-7-sonnet Claude 3.7 sonnet and claude code

  5. [5]

    Zina Chkirbene, Ridha Hamila, Ala Gouissem, and Unal Devrim. 2024. https://doi.org/10.1109/HONET63146.2024.10822885 Large language models (llm) in industry: A survey of applications, challenges, and trends . In 2024 IEEE 21st International Conference on Smart Communities: Improving Quality of Life using AI, Robotics and IoT (HONET), pages 229--234

  6. [6]

    Cohere. 2025. https://huggingface.co/CohereForAI/c4ai-command-a-03-2025 Model card for c4ai command a

  7. [7]

    John Dang, Shivalika Singh, Daniel D'souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, Sandra Kublik, Meor Amer, Viraat Aryabumi, Jon Ander Campos, Yi-Chern Tan, Tom Kocmi, Florian Strub, Nathan Grinsztajn, Yannis Flet-Berliac, Acyr Locatelli, Hangyu Lin, Dwarak Talupuru, Bharat Venk...

  8. [8]

    Google Deepmind. 2024. https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/ Introducing gemini 2.0: our new ai model for the agentic era

Show all 68 references
  1. [9]

    DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...

  2. [10]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  3. [11]

    Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile van Krieken, and Pasquale Minervini...

  4. [12]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  5. [13]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=d7KBjmI3GmQ Measuring massive multitask language understanding . In International Conference on Learning Representations

  6. [14]

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. https://openreview.net/forum?id=rygGQyrFvH The curious case of neural text degeneration . In International Conference on Learning Representations

  7. [15]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. https://openreview.net/forum?id=chfJJYC3iL Livecodebench: Holistic and contamination free evaluation of large language models for code . I...

  8. [16]

    Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q

    Mehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch, Chrysovalantis Anastasiou, Sanket Vaibhav Mehta, Lalit K. Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q. Tran, Q...

  9. [17]

    Eunsu Kim, Juyoung Suk, Philhoon Oh, Haneul Yoo, James Thorne, and Alice Oh. 2024 a . https://aclanthology.org/2024.lrec-main.296/ CLI c K : A benchmark dataset of cultural and linguistic intelligence in K orean . In Proceedings of the 2024 Joint International Conference on Co...

  10. [18]

    Hyeonwoo Kim, Dahyun Kim, Jihoo Kim, Sukyung Lee, Yungi Kim, and Chanjun Park. 2025. https://arxiv.org/abs/2410.12445 Open ko-llm leaderboard2: Bridging foundational and practical evaluation for korean llms . Preprint, arXiv:2410.12445

  11. [19]

    Yeeun Kim, Youngrok Choi, Eunkyung Choi, JinHwan Choi, Hai Jin Park, and Wonseok Hwang. 2024 b . https://doi.org/10.18653/v1/2024.findings-emnlp.319 Developing a pragmatic benchmark for assessing K orean legal language understanding in large language models . In Findings of th...

  12. [20]

    Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In Advances in Neural Information Processing Systems, volume 35, pages 22199--22213

  13. [21]

    Sunjun Kweon, Byungjin Choi, Gyouk Chu, Junyeong Song, Daeun Hyeon, Sujin Gan, Jueon Kim, Minkyu Kim, Rae Woong Park, and Edward Choi. 2024. https://arxiv.org/abs/2403.01469 Kormedmcqa: Multi-choice question answering benchmark for korean healthcare professional licensing exam...

  14. [22]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating S...

  15. [23]

    Huiyuan Lai and Malvina Nissim. 2024. https://doi.org/10.18653/v1/2024.acl-long.649 m C o T : Multilingual instruction tuning for reasoning consistency in language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...

  16. [24]

    Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D

    Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luc...

  17. [25]

    Hwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, and Jung-Woo Ha. 2023. https://doi.org/10.18653/v1/2023.acl-long.370 SQ u AR e: A large-scale dataset of sensitive question...

  18. [26]

    Meta. 2024 a . https://ai.meta.com/blog/future-of-ai-built-with-llama/ The future of ai: Built with llama

  19. [27]

    Meta. 2024 b . https://ai.meta.com/blog/meta-llama-3-1/ Introducing llama 3.1: Our most capable models to date

  20. [28]

    Meta. 2024 c . https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

  21. [29]

    Meta. 2025. https://ai.meta.com/blog/llama-4-multimodal-intelligence/ The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

  22. [30]

    Mistral. 2025. https://mistral.ai/news/mistral-small-3-1 Mistral small 3.1

  23. [31]

    Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025. https://arxiv.org/abs/2501.19393 s1: Simple test-time scaling . Preprint, arXiv:2501.19393

  24. [32]

    naver hyperclovax. 2025. https://huggingface.co/naver-hyperclovax/HyperCLOVAX-SEED-Text-Instruct-1.5B Model card for hyperclovax-seed-text-instruct-1.5b

  25. [33]

    OneLineAI. 2025. Ko-r1-7b-v2.1. https://huggingface.co/OLAIR/ko-r1-7b-v2.1. Accessed: 26 March 2025

  26. [34]

    OpenAI, :, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, Alex Kirillov, Alex Nichol, Alex Pai...

  27. [35]

    OpenAI, :, Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, Alex Iftimie, Alex Karpenko, Alex Tachard Passos, Alexander Neitz, Alexander Prokofiev, Alexander Wei, Allison Tam, Ally Bennet...

  28. [36]

    OpenAI. 2024. https://huggingface.co/datasets/openai/MMMLU Multilingual massive multitask language understanding (mmmlu)

  29. [37]

    OpenAI. 2025 a . https://openai.com/index/gpt-4-1/ Introducing gpt-4.1 in the api

  30. [38]

    OpenAI. 2025 b . https://openai.com/index/introducing-o3-and-o4-mini/ Introducing openai o3 and o4-mini

  31. [39]

    OpenAI. 2025 c . https://openai.com/index/openai-o3-mini/ Openai o3-mini

  32. [40]

    Chanjun Park, Hyeonwoo Kim, Dahyun Kim, SeongHwan Cho, Sanghoon Kim, Sukyung Lee, Yungi Kim, and Hwalsuk Lee. 2024. https://doi.org/10.18653/v1/2024.acl-long.177 Open K o- LLM leaderboard: Evaluating large language models in K orean with K o-h5 benchmark . In Proceedings of th...

  33. [41]

    Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. 2025. https://arxiv.org/abs/2506.20920 Fineweb2: One pipeline to scale them all -- adapting pre-training da...

  34. [42]

    Wang, Robert Gerbicz, John-Clark Levin, Serguei Popov, Fiona Feng, Steven Y

    Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Tung Nguyen, Daro...

  35. [43]

    Irene Plaza, Nina Melero, Cristina del Pozo, Javier Conde, Pedro Reviriego, Marina Mayor-Rocher, and María Grandury. 2024. https://arxiv.org/abs/2406.17789 Spanish and llm benchmarks: is mmlu lost in translation? Preprint, arXiv:2406.17789

  36. [44]

    Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...

  37. [45]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. https://openreview.net/forum?id=Ti67584b98 GPQA : A graduate-level google-proof q&a benchmark . In First Conference on Language Modeling

  38. [46]

    LG AI Research, :, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Kyubeen Han, Seokhee Hong, Junwon Hwang, Taewan Hwang, Joonwon Jang, Hyojin Jeon, Kijeong Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Euisoon Kim, Hyosang Kim, Jihoon Kim, Joon...

  39. [47]

    LG AI Research, Soyoung An, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Yountae Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil...

  40. [48]

    LG AI Research, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Seokhee Hong, Junwon Hwang, Hyojin Jeon, Kijeong Jeon, Gerrard Jeongwon Jo, Hyunjik Jo, Jiyeon Jung, Hyosang Kim, Joonkee Kim, Seonghwan Kim, Soyeon Kim, Sunkyoung Kim, Yireun Kim, Yongil...

  41. [49]

    Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. 2024. https://openreview.net/forum?id=m2NVG4Htxs To the cutoff... and beyond? a longitudinal perspective on LLM data contamination . In The Twelfth International Conference on Learning Representations

  42. [51]

    Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I. Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Wei-Yin Ko, Madeline Smith, Antoine Bosselut, Alice Oh, Andre F. T....

  43. [52]

    Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2024 a . https://arxiv.org/abs/2402.11548 Kmmlu: Measuring massive multitask language understanding in korean . Preprint, arXiv:2402.11548

  44. [53]

    Guijin Son, Hanwool Lee, Suwan Kim, Huiseo Kim, Jae cheol Lee, Je Won Yeom, Jihyu Jung, Jung woo Kim, and Songseong Kim. 2024 b . https://aclanthology.org/2024.lrec-main.704/ HAE - RAE bench: Evaluation of K orean knowledge in language models . In Proceedings of the 2024 Joint...

  45. [54]

    Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri \`a Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, ...

  46. [55]

    Google Team. 2025 a . https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf Gemma 3 technical report

  47. [56]

    Kanana LLM Team, Yunju Bak, Hojin Lee, Minho Ryu, Jiyeon Ham, Seungjae Jung, Daniel Wontae Nam, Taegyeong Eo, Donghun Lee, Doohae Jung, Boseop Kim, Nayeon Kim, Jaesun Park, Hyunho Kim, Hyunwoong Ko, Changmin Lee, Kyoung-Woon On, Seulye Baeg, Junrae Cho, Sunghee Jung, Jieun Kan...

  48. [57]

    M-A-P Team, Xinrun Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, Kang Zhu, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, Kaixing Deng, Shuyue Guo, Shian Jia, Sichao Jiang, Yiyan Liao, Rui Li, Qinrui Li, Sirun Li, Yizhi Li, Yunwen Li, Dehua Ma, Yua...

  49. [58]

    Qwen Team. 2025 b . https://qwenlm.github.io/blog/qwq-32b/ Qwq-32b: Embracing the power of reinforcement learning

  50. [59]

    Joshua Vendrow, Edward Vendrow, Sara Beery, and Aleksander Madry. 2025. https://arxiv.org/abs/2502.03461 Do large language model benchmarks test reliability? Preprint, arXiv:2502.03461

  51. [60]

    Yiming Wang, Pei Zhang, Jialong Tang, Haoran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. https://arxiv.org/abs/2504.18428 Polymath: Evaluating mathematica...

  52. [61]

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, Tianle Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and Wenhu Chen. 2024. https://openreview.net/forum?id=y10DM6R2r3 MMLU -pro: A more...

  53. [62]

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and M...

  54. [63]

    xAI. 2025. https://x.ai/news/grok-3 Grok 3 beta — the age of reasoning agents

  55. [64]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jia...

  56. [65]

    Kang Min Yoo, Jaegeun Han, Sookyo In, Heewon Jeon, Jisu Jeong, Jaewook Kang, Hyunwook Kim, Kyung-Min Kim, Munhyong Kim, Sungju Kim, Donghyun Kwak, Hanock Kwak, Se Jung Kwon, Bado Lee, Dongsoo Lee, Gichang Lee, Jooho Lee, Baeseong Park, Seongjin Shin, Joonsang Yu, Seolki Baek, ...

  57. [66]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. https://doi.org/10.18653/v1/P19-1472 H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791--4...

  58. [67]

    Yidan Zhang, Yu Wan, Boyi Deng, Baosong Yang, Haoran Wei, Fei Huang, Bowen Yu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. https://arxiv.org/abs/2411.09116 P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms . Preprint, arXiv:2411.09116

  59. [68]

    Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui, Qinzheng Sun, Shaoguang Mao, Xin Zhang, Ying Xin, Qiufeng Yin, Scarlett Li, and Furu Wei. 2024. https://arxiv.org/abs/2412.15194 Mmlu-cf: A contamination-free multi-task language understanding benchmark . Preprint, arXiv:2412.15194

  60. [69]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. https://openreview.net/forum?id=VqkAKQibpq SGL ang: Efficient execution of structured language ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.