Pith. sign in

REVIEW 23 cited by

Are We Done with MMLU?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04127 v3 pith:GDS7MHYH submitted 2024-06-06 cs.CL cs.AI

classification cs.CLcs.AI
keywords mmluerrorsquestionsbenchmarkcontainmmlu-reduxsubsetacross
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Maybe not. We identify and analyse errors in the popular Massive Multitask Language Understanding (MMLU) benchmark. Even though MMLU is widely adopted, our analysis demonstrates numerous ground truth errors that obscure the true capabilities of LLMs. For example, we find that 57% of the analysed questions in the Virology subset contain errors. To address this issue, we introduce a comprehensive framework for identifying dataset errors using a novel error annotation protocol. Then, we create MMLU-Redux, which is a subset of 5,700 manually re-annotated questions across all 57 MMLU subjects. We estimate that 6.49% of MMLU questions contain errors. Using MMLU-Redux, we demonstrate significant discrepancies with the model performance metrics that were originally reported. Our results strongly advocate for revising MMLU's error-ridden questions to enhance its future utility and reliability as a benchmark. https://huggingface.co/datasets/edinburgh-dawg/mmlu-redux-2.0.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 23 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fluid Language Model Benchmarking

    cs.CL 2025-09 conditional novelty 8.0 of 10

    Fluid Benchmarking, combining IRT-based ability estimation with Fisher-information-based adaptive item selection, improves LM evaluation across efficiency, validity, variance, and saturation in pretraining settings.

  2. Out-of-Distribution Generalization of Risk Aversion in Language Models

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Risk aversion trained on ≤$100 gambles partially generalizes across 98 orders of magnitude in LMs, raising astronomical-stakes Cooperate rates from ~2% to ~39–70% depending on method.

  3. Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new benchmark shows LLMs' financial calculation accuracy collapses without explicit formulas and degrades further when they must generate multi-metric tables.

  4. From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.

  5. Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TeaR uses GRPO reinforcement learning on test-case output prediction for algorithmic problems, with no code shown, and reports broad reasoning gains across 17 benchmarks.

  6. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  7. Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Under simulated leakage, n-gram-based detection beats permutation and truncation methods, and cleaning flag-prone MMLU instances changes model rankings only slightly.

  8. ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark converts text-only QA datasets into text images and shows that vision-language models degrade sharply on long visually presented contexts.

  9. LLMzSz{\L}: a comprehensive LLM benchmark for Polish

    cs.CL 2025-01 conditional novelty 6.0 of 10

    LLMzSzŁ is a new benchmark of almost 19,000 Polish national exam questions with evaluations of 38 language models and comparisons to human results.

  10. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  11. Raon-Speech Technical Report

    cs.CL 2026-04 conditional novelty 5.5 of 10

    A 9B SpeechLM trained on 1.38M hours of English/Korean data plus a full-duplex chat extension trained on 119K hours of time-aligned dialogue outperform same-size audio models on speech tasks and FDB turn-taking metrics.

  12. The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Small open-source LLMs give inconsistent answers on 20% to 50% of multiple-choice questions even at low temperature, while 50B-80B models are consistent on over 90% of questions.

  13. "Check My Work?": Measuring Sycophancy in a Simulated Educational Context

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across five OpenAI models, mentioning a correct answer in a query boosts LLM accuracy by up to 15 points, while mentioning an incorrect answer lowers it by a similar amount.

  14. dots.llm1 Technical Report

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A 14B-active MoE model roughly matches Qwen2.5-72B on a broad benchmark suite while reporting about a 4x reduction in training GPU-hours.

  15. PC-MoE: Memory-Efficient and Privacy-Preserving Collaborative Training for Mixture-of-Experts LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    PC-MoE shards the expert layers of an MoE LLM across parties and routes only sparse top-k activations between them, achieving near-centralized accuracy with about 70% memory savings and resistance to one partial-gradi...

  16. Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?

    cs.CL 2025-02 conditional novelty 5.0 of 10

    An ensemble built from repeated samples of a single strong LLM outperforms the standard multi-model Mixture-of-Agents on several benchmarks.

  17. MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

    cs.CL 2024-12 conditional novelty 5.0 of 10

    MMLU-CF is a new closed-test-set MCQ benchmark that uses rephrasing, choice shuffling, and 'None of the other choices' distractors to reduce data contamination, where GPT-4o scores 73.4% 5-shot.

  18. GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs

    cs.CR 2024-11 conditional novelty 5.0 of 10

    A three-bit targeted memory corruption, located by a genetic search, collapses an 8B-parameter quantized LLM's benchmark performance to zero.

  19. Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment

    cs.HC 2026-07 accept novelty 4.0 of 10

    AI has rapidly crossed expert baselines on bounded tasks on a still-jagged frontier, so human work must shift from production to specification, verification, and oversight.

  20. AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    AgentScope 1.0 packages the components needed to build, evaluate, and deploy LLM agent applications into one developer framework.

  21. Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Pangu Embedded, a 7B reasoner trained with iterative distillation, RL, and an adaptive fast/slow thinking scheme, reports superior benchmark scores to similarly sized Qwen3-8B and GLM-4-9B.

  22. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

    cs.AI 2025-02 conditional novelty 4.0 of 10

    A meta-review of about 110 critical studies finds nine systemic weaknesses in AI benchmarking and concludes that benchmarks are receiving disproportionate trust in AI governance.

  23. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools