REVIEW 10 cited by
KMMLU: Measuring Massive Multitask Language Understanding in Korean
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language. We test 27 public and proprietary LLMs and observe the best public model to score 50.5%, leaving significant room for improvement. This model was primarily trained for English and Chinese, not Korean. Current LLMs tailored to Korean, such as Polyglot-Ko, perform far worse. Surprisingly, even the most capable proprietary LLMs, e.g., GPT-4 and HyperCLOVA X do not exceed 60%. This suggests that further work is needed to improve LLMs for Korean, and we believe KMMLU offers the appropriate tool to track this progress. We make our dataset publicly available on the Hugging Face Hub and integrate the benchmark into EleutherAI's Language Model Evaluation Harness.
Forward citations
Cited by 10 Pith papers
-
GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus
A 303,581-row Korean instruction corpus generated seedlessly from a 1,084-discipline taxonomy, with near-zero duplicates and low measured overlap with KMMLU, KoBEST, and HAE-RAE-Bench.
-
KoBLEX: Open Legal Question Answering with Multi-hop Reasoning
KoBLEX is a bilingual 226-instance provision-grounded legal QA benchmark, and its ParSeR pipeline (generate pseudo-statutes, then retrieve-rerank-select real ones) beats baselines across five LLMs, graded by a new hum...
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.
-
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
A cost-effective recipe consisting of tokenizer extension, continual pretraining, FP8 training, and SFT/DPO post-training yields Korean-English bilingual 8B models with top Korean benchmark scores.
-
MultiHoax: A Dataset of Multi-hop False-Premise Questions
A new multi-hop false-premise benchmark shows that leading large language models detect embedded falsehoods in only a minority of cases, with the best model reaching about 23% on the full two-stage protocol.
-
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
A 102B Korean-English model, expanded from Llama 3 70B with LlamaPro and Masked Structure Growth and trained on 194B tokens, scores 64.74 on KMMLU and 83.34 on KorMedMCQA, roughly matching GPT-4 on Korean benchmarks.
-
Making Sense of Korean Sentences: A Comprehensive Evaluation of LLMs through KoSEnd Dataset
A new Korean benchmark, KoSEnd, shows LLMs have limited grasp of Korean sentence endings, and warning them about potentially missing endings improves their choices.
-
Opt.Gear Technical Report
Opt.Gear is a family of efficient on-device language models using a ConvKV-gated mixer with sparse attention, trained on 0.5T tokens without distillation, claiming up to 4.9x NPU speedups and 20 TPS on a Cortex-M7.
-
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
A backward-propagation scoring scheme over a signed temporal DAG can identify malicious agents in LLM multi-agent systems and cut their communications, improving defended accuracy by 3–7 percentage points in the autho...
-
Smoothie-Qwen: Post-Hoc Smoothing to Reduce Language Bias in Multilingual LLMs
Smoothie-Qwen reduces Chinese-language output in Qwen2.5-Coder by scaling down lm head weights of Chinese tokens, reaching 95% suppression on a synthetic elicitation set while keeping Korean MMLU accuracy roughly stable.
Discussion (0). Sign in to comment.