Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T15:51:04.674346Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 100 inbound Pith citation observations for arXiv:2406.01574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T15:51:04.674346Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:55:01.947238Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
54 of 54 outbound references displayed
External citation measurements
13
pith, observed 2026-08-05T02:28:24.338817Z
Observation 8234b9d2-1c13-457f-aebe-7d9712606e39 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ab3c5ce-f484-4c17-abef-3279fbbc686f · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a93ec630-105a-4a4c-a428-285fb95e7f65 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e33005bd-4ea5-4b2a-8f8e-67a422c7979b · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Llemma: An Open Language Model For Mathematics
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ad7b15b-7468-4c6e-87d2-0c0b3c72b908 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Qwen Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 07b26437-c504-4e4c-b354-436e5de5ce8d · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Constitutional AI: Harmlessness from AI Feedback
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dae13852-ec92-4270-b9e0-c5c9fd915394 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Language models are few-shot learners
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5b68c9b-5794-42fe-a950-33507f13d739 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark C4ai-command-r-v01
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 388eeb6b-0750-4eaa-bcb5-f5df8ef21cd4 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark A survey on evaluation of large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c29c35f-b05c-4344-904a-b9625bb5bc31 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Theoremqa: A theorem-driven question answering dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a07d66b-ea87-4efa-9588-0d8643c0835f · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 492b83d4-6c0d-4287-b9c8-94e1c1b41563 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d76b1645-9be5-4808-a115-c7e196f71b2f · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Introducing the next generation of Claude https://www.anthropic.com/news/claude-3- family
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 55ef595b-430f-4c61-b692-a5d75feb2c78 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Opencompass: A universal evaluation platform for foundation models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 57632eba-9ea4-41e9-a3b7-291c12954239 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc57fbf5-dd0d-4f2f-b7e5-f2439d25b8ec · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08d07f93-b368-442c-9580-e75518cd6ca7 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Hello gpt4-o
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 138b782d-5d60-4b69-8591-73fd4c777ae2 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Measuring Massive Multitask Language Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e186c4b4-df80-4dc1-b526-5b50baea891d · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Measuring Mathematical Problem Solving With the MATH Dataset
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a66e1811-fd49-41e9-b8fd-45cfc6627eb9 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Mistral 7B
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc3eb58a-ed94-42c8-a088-2ca4bbe6a028 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Mixtral of Experts
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8a3635d8-5b20-4c17-9a12-330413744d0b · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Holistic Evaluation of Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f640b04-8a51-4717-944e-9497af619c71 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Lingyiwanwu, yi-large
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6f2251c1-4efa-4ced-a393-b3b2eba62bcc · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Build the future of ai with meta llama 3 - https://llama.meta.com/llama3/
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ea39cfd5-e27a-4d09-ba65-aaca128c59da · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Cross-Task Generalization via Natural Language Crowdsourcing Instructions
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14d1b2fb-8384-495c-afbc-8b036362b99c · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Levels of agi: Opera- tionalizing progress on the path to agi
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45418f88-7c29-4464-a544-f49336cdac0f · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 891b46ca-3b5b-424d-b55f-257ac0ee9273 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7a9afcc-b632-4e4e-8d02-8dd740516664 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6e4ec4d-23d5-4def-9ea6-9d2dfe73c24e · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 215a3e30-542b-4e41-855b-8aa9356a55d2 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Multitask Prompted Training Enables Zero-Shot Task Generalization
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6dc02a4e-23d8-4905-8760-e3009390288a · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 825355f2-d67c-41ec-a68a-825ff5a528a9 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 42704a8b-0d90-46d6-ac73-82d3f261e5e2 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Gemma: Open Models Based on Gemini Research and Technology
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e9b7867-776e-40cf-9c5d-817aec9f4813 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a54f8831-b313-4f41-8859-7733f6944d57 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Zephyr: Direct Distillation of LM Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9e9eebd7-178e-4594-85da-abfeecc9fe4e · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d3cd9924-9ccf-47ac-b75b-2bdc04dfca44 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Superglue: A stickier benchmark for general-purpose language understanding systems
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9e16c97-0023-415e-b086-5c13d04767b6 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 84c050a8-5ce5-40d8-ba90-e005cc210ad3 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 709bf6dd-02a4-460e-ab59-0ad70045c6eb · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chain-of-thought prompting elicits reasoning in large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d25fc4c-3820-4026-bafc-a91cb5f3456a · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Internlm-math: Open math large language models toward verifiable reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9247c9ee-d996-4f1e-b553-774a98c2fdfa · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Yi: Open Foundation Models by 01.AI
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2b5358e2-adf0-4273-b65d-bdc78f894573 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark MAmmoTH2: Scaling Instructions from the Web
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b8c5ff66-d693-44cc-b20b-0cb6569e2320 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7e6bb69-0747-4fa8-858c-abb009fa0f6b · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Map-neo: Highly capable and transparent bilingual large language model series
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3e6af77f-8187-4e43-9346-dcd1be80e88d · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Large language models are not robust multiple choice selectors
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8fc1faa6-3370-47a2-a4fd-c42d34a39ac0 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark image_question
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85e82a9f-568c-42d5-8a63-0f8a970a3490 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Strain II has an average weight of 15 grams
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c7efa464-4fb1-41be-bd4d-551bbc153f6a · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Each lowercase letter gene (a, b, d) contributes 2.5 grams
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d53d232d-80ba-4113-8ee5-618e0d3f717e · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - For Strain II (15 grams): - 15 = ( a + a + b + b + d + d) - Each lowercase letter contributes 2.5 grams, so: - 15 = 6 × 2.5 - Therefore, Strain II must have the genotype aabbdd
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96060590-f1cb-401b-bfe7-45e8157c048f · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b440d32-79a4-4cc9-9cf6-414074ad9ce3 · outbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Since there are three pairs, the total weight is 11.25 × 3 = 33 .75 grams
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5589a9fd-58b4-44af-ab06-1c4b4e693fbf · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cbd8afec-9ea4-41ba-b087-9325b32ea148 · inbound
UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d5817a28-27f2-4682-ba99-49f6ee4d2de2 · inbound
SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb473fcc-da64-407d-becc-3038ddc8d4ef · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 31450546-1025-41f3-a1d7-210457e2cc3b · inbound
VoiceBench: Benchmarking LLM-Based Voice Assistants MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 14bf5fc2-f624-4d6f-8bad-89df8cbd4141 · inbound
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8718085d-9a5e-4b4f-a63a-211e2601203f · inbound
Qwen2.5 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b9ffacf9-c1fd-4cf2-962c-989fdc273cb8 · inbound
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9a943b6-6e23-42b6-9f29-1c77c448da02 · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 156
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 95cf7e11-30a1-4de7-bf0b-3c86ec5b9bdc · inbound
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06000166-b00c-4927-90f0-36eb17c4098f · inbound
Humanity's Last Exam MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb0a87d3-9c22-46de-8c26-ea71076eb1c0 · inbound
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 237
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a9aba33-b748-488a-9882-52af28f6a352 · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 945cac13-4dbf-4c8b-bafd-ca751ce6d86b · inbound
Learning to Reason under Off-Policy Guidance MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 40de151a-8bb8-4919-b6aa-c15a3dee4675 · inbound
PRIMETIME : Limits of LLMs in Temporal Primitives MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 21a94ca9-336c-4565-b90e-f089ea37fe87 · inbound
Phi-4-reasoning Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation add70447-8e26-4327-81e8-9795bd1463e3 · inbound
Qwen3 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5874fde9-dc86-446d-8155-58c5b371f3d9 · inbound
Kimi K2: Open Agentic Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 33b53e28-96ef-4d18-8419-7549eb4b4098 · inbound
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ec688596-32e5-417f-8af7-837f217f0b75 · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3c90f828-096b-480c-bf77-958bc06d5451 · inbound
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a51aa7-b234-4598-a0dd-fdd48ef86b7f · inbound
DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41752cb9-bf29-4457-9506-43f60f78cadd · inbound
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 853481f2-d296-431e-9a06-221d526c880a · inbound
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d3ceef-5195-4d7c-b4ee-8e0cb89b9aca · inbound
Scaling Latent Reasoning via Looped Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b4bca95f-70eb-4909-bf5a-99a090fc6780 · inbound
Scaling Latent Reasoning via Looped Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e592b64-9b71-43cf-a18a-f13e60f0fc48 · inbound
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d48a48f-0626-4c04-b14f-446aaa904da4 · inbound
Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a4e6419c-84b2-4293-a891-e7e7ae4d1edf · inbound
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e6b570cf-1747-483d-867b-ea5c69fabbf1 · inbound
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ca3ea0aa-fbef-42cf-a952-6f954ac057f4 · inbound
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d56cf372-e911-493c-b755-c56914d404c9 · inbound
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1bb6e91e-f99c-434f-94c0-21869bc5e718 · inbound
Coupled Variational Reinforcement Learning for Language Model General Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37935a84-316a-48d8-b77d-f9260d32168a · inbound
NVIDIA Nemotron 3: Efficient and Open Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 501a36c9-b6a6-42da-83c3-236f8a557516 · inbound
SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a43d7341-ecf4-40a2-9701-f2a8a24dab06 · inbound
Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa417c7-b8ee-4eeb-99ee-f515404fff38 · inbound
Kimi K2.5: Visual Agentic Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 44893d04-a0d4-4ebf-91d0-affb247be254 · inbound
Interfaze: The Future of AI is built on Task-Specific Small Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788abe45-388b-43ba-bbf4-97259f34881d · inbound
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b1f77faf-3a78-4387-a714-1653c3f5934d · inbound
Model soups need only one ingredient MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3d75179-36ee-422b-a25e-1390f51e1c74 · inbound
When LLMs get significantly worse: A statistical approach to detect model degradations MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab9cbd9a-41aa-410b-9fb2-8beda59b1d44 · inbound
SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b72c55e7-121f-469c-a4df-735975774b2a · inbound
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caeb7d52-bc16-4fac-a8f0-5664799f36bc · inbound
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 349
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f43f4cb7-6c5a-45bc-a725-120c585397cc · inbound
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f83f8fb-db11-4b5a-b2bb-7a190450b54c · inbound
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 80efb0bd-9825-4256-92e4-a6d350c0c80c · inbound
Adaptively Robust LLM Monitoring via Activation Watermarking MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2f23db-da70-419b-984e-dde9af5f0f48 · inbound
Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6c763adf-ed5f-4394-9af3-ae316fca3616 · inbound
Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d78aedea-1cda-495f-b03e-bc2ae3a7bcbd · inbound
Can LLMs Learn to Reason Robustly under Noisy Supervision? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1659414-7bad-4690-9125-b7331fa6d383 · inbound
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ca98958-f293-46bc-aeb3-13faeff71956 · inbound
An Algebraic Introduction to Persistence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e138d32-793c-4c50-bbc2-89e73ec0668f · inbound
MARS: Enabling Autoregressive Models Multi-Token Generation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d9c6c63-4821-4781-943b-05804e029e3b · inbound
EXAONE 4.5 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 79092f1e-a2f9-4337-bc51-0762e73a6eab · inbound
Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4dfae951-1f02-4ee3-9604-a503b70e2e7a · inbound
SAGE Celer 2.6 Technical Card MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf083123-c623-418c-9b18-41476d5f7bac · inbound
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3b268b7-6730-49c6-b7a8-8003c6578aac · inbound
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3e64bc-bbab-4053-b9e2-9ba504ee56bf · inbound
Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f0e7dc22-40c6-45a3-be3e-5f64771b33ab · inbound
Super Apriel: One Checkpoint, Many Speeds MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3b5426d-f612-473c-83dc-df844eb8b772 · inbound
Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 400845a3-79bc-45ec-9893-386c1ea14dc4 · inbound
Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2cb20aaf-5490-4d74-aab3-b8f6c07abe71 · inbound
COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 45f9dfce-013a-4c58-8ccf-878cbf61be79 · inbound
Decoupled DiLoCo for Resilient Distributed Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2cfb4e89-8f68-4fcc-bf3b-8e3619686296 · inbound
Large Language Models Decide Early and Explain Later MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6418bdca-2c56-4701-8d70-70b495f8ae2b · inbound
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0f200792-a149-4cb1-9cdc-88d7d63c1b3a · inbound
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e47e59a7-39f5-4d39-9311-06e1dc67e90f · inbound
On the Privacy of LLMs: An Ablation Study MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c3543f14-d254-456a-bc90-e59de6e7622f · inbound
Analysis and Explainability of LLMs Via Evolutionary Methods MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c459372-68c1-4500-ad77-d9fc3986ac3a · inbound
MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1e9628e-5de0-4c35-82db-d6a5ed30e9a3 · inbound
CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 571dec5d-bd09-4e02-b987-0d54233a5938 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5f17757-15a0-4b7f-b6e5-ced88756b3a4 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7005586a-ccad-47d4-b341-0ef14cc2ac73 · inbound
Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2746932e-ad2b-475b-9cf0-99433ca2815b · inbound
An Interpretable and Scalable Framework for Evaluating Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 309a17ee-5789-4d99-bb5d-11047c569070 · inbound
Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3ffc6f3c-aebb-48eb-8c30-b12165231b2a · inbound
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bb546503-eb3a-4299-9ff5-7413e0c653f4 · inbound
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 253b5cf3-edda-4978-82e5-d2ae8b28a274 · inbound
LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1ad04c0f-798c-4c0c-8e84-b0354f2f0ee8 · inbound
Rotation-Preserving Supervised Fine-Tuning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2aa65d82-511a-41e7-9f90-88520182bc1a · inbound
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73b0d0dd-1501-492f-ac0c-8a5e4ccaedbd · inbound
Unsteady Metrics and Benchmarking Cultures of AI Model Builders MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96fa0261-503c-4882-9cc0-c9bbbce191d5 · inbound
TeachArena: Are Language Agents Ready for Realistic Teaching Work? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc718502-e1b7-4c4a-8984-acb4bd8be7db · inbound
TeachArena: Are Language Agents Ready for Realistic Teaching Work? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c8b5fd-c397-4ecc-b0e4-63273f66d950 · inbound
MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7bafee52-7982-48f9-89b3-e066ae5ebfe5 · inbound
MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9787f1d0-ef75-477e-9f84-53dbdec6798f · inbound
Open-World Evaluations for Measuring Frontier AI Capabilities MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f057cf59-39b0-4044-9dd9-d3d5d2de689e · inbound
Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ddaf38a6-2acb-40d3-9c98-4bb54e3b53c1 · inbound
LoMo: Local Modality Substitution for Deeper Vision-Language Fusion MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4627ff3a-a14e-4ad9-87fb-766af108731d · inbound
Mellum2 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7ca2c20d-24b1-4b15-9f6b-18f515c3dd59 · inbound
CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 511d0302-3f36-4d8a-b12e-2ecd0b4189c1 · inbound
SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3f6cd5f6-0b5c-43e1-8c60-b263422dd82b · inbound
Trading Human Curation for Synthetic Augmentation in RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1d580256-1243-4dee-9619-9d1b8345ac96 · inbound
Trading Human Curation for Synthetic Augmentation in RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d6db900-ff05-4272-b511-3f274435a586 · inbound
Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4bbad33-11f8-4f66-96c4-61e9adfad71b · inbound
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26fa7b81-3a4b-462d-b437-bd9f0af0aa83 · inbound
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4669ed3c-54ee-41b2-ba87-bb23aafd33e2 · inbound
Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e69f046e-df83-451e-9bfa-37776fe2ac3d · inbound
The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b741e2a-e8b5-491c-b6d6-8258f34bfc64 · inbound
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.