Pith. sign in

Paper Citation Record · LEDGER

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 100 inbound Pith citation observations for arXiv:2406.01574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01574 v6

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T15:51:04.674346Z

measured 154 of 154 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 144 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:55:01.947238Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact30
  • verified fuzzy22
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

13
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8234b9d2-1c13-457f-aebe-7d9712606e39 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.647897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:d6dde1bfef846114137dc170cfa778d4eafb94b710be4f70812271e8f2f83a44

Observation 8ab3c5ce-f484-4c17-abef-3279fbbc686f · outbound

This paper cites GPT-4 Technical Report.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:06.112897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:d201044bd71e59d2a8d49ef8704c069a7d7076ea2de6bf001326bc19ccaffb29

Observation a93ec630-105a-4a4c-a428-285fb95e7f65 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:06.285386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:c5f7a0c3cc6ee818963c5266a1c43538cd201ae536a6ca0b2221bc4d516bdfeb

Observation e33005bd-4ea5-4b2a-8f8e-67a422c7979b · outbound

This paper cites Llemma: An Open Language Model For Mathematics.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Llemma: An Open Language Model For Mathematics

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:47.048289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:2b65e73979d9be704b2bf385b21f631e1b375d1c44df2c93db3ac1bc708eb2f2

Observation 8ad7b15b-7468-4c6e-87d2-0c0b3c72b908 · outbound

This paper cites Qwen Technical Report.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.512353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:e6da23aa0b2c1c42b38f88f004d86801e30a1e4b6683d350d2145ccf43a8f442

Observation 07b26437-c504-4e4c-b354-436e5de5ce8d · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:09.594468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:0d24402b084d0297df610ae73217b2c856a62135a819a86fc6dd52cfcb9b0e88

Observation dae13852-ec92-4270-b9e0-c5c9fd915394 · outbound

This paper cites Language models are few-shot learners.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Language models are few-shot learners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:10.387353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:bd57c21e23674ce59c1b306efe437e987cbc9675e755de2cf340193143c284a0

Observation c5b68c9b-5794-42fe-a950-33507f13d739 · outbound

This paper cites C4ai-command-r-v01.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark C4ai-command-r-v01

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:10.586258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:6dab6f8c52173dfc51465a1367e783ba3e38fb6b2c4df5bcaabfd42e96d670b4

Observation 388eeb6b-0750-4eaa-bcb5-f5df8ef21cd4 · outbound

This paper cites A survey on evaluation of large language models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark A survey on evaluation of large language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:10.724417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:9da1da7b5f5cef6a17372581fb05f6649e499aa3b81e36555c65a7aa963ff7e4

Observation 3c29c35f-b05c-4344-904a-b9625bb5bc31 · outbound

This paper cites Theoremqa: A theorem-driven question answering dataset.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Theoremqa: A theorem-driven question answering dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:10.934346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:df9a9eea856768b11a8ac087047ccdd53dd9a6c077ebd3714b0432621475edbe

Observation 8a07d66b-ea87-4efa-9588-0d8643c0835f · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:13:25.541930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:3b33ade104382a1e52118a826be6d60dbb67dd5f0698ca1d0bf8e9cba02186ec

Observation 492b83d4-6c0d-4287-b9c8-94e1c1b41563 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:09.355009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:6f791ef59b81e69952642e70d99591fd67199774da0cf0ea39f7b7fe12b38d28

Observation d76b1645-9be5-4808-a115-c7e196f71b2f · outbound

This paper cites Introducing the next generation of Claude https://www.anthropic.com/news/claude-3- family.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Introducing the next generation of Claude https://www.anthropic.com/news/claude-3- family

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:11.065802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:6d29c423fa804ba7585814de2ef4cd35ea7bd6d86838589e0d7c5de88ce0f1b1

Observation 55ef595b-430f-4c61-b692-a5d75feb2c78 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Opencompass: A universal evaluation platform for foundation models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:11.144240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:2264cd937b32d2a8f7811493075f4eef57606980ae0d298793434f36ad004ba1

Observation 57632eba-9ea4-41e9-a3b7-291c12954239 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:11.304857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:3dcfec07b93e49bb41ba61a8bdee20fb3f68eee3363a1b31157bd22fec370fbc

Observation fc57fbf5-dd0d-4f2f-b7e5-f2439d25b8ec · outbound

This paper cites Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:06.479417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:d75f88abac8df5566e82e9f86d0458ab15453b230bc7472ad55637b87725a024

Observation 08d07f93-b368-442c-9580-e75518cd6ca7 · outbound

This paper cites Hello gpt4-o.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Hello gpt4-o

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:11.564521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:8b8c9de50f4f1266afcdd71dfab72bc9426c7e763485c608de6372078fa520f2

Observation 138b782d-5d60-4b69-8591-73fd4c777ae2 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Measuring Massive Multitask Language Understanding

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:07.557358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:d7fee434357fd245a01e6a44a8490805784fc49dccad603cd386e4f71ff238ff

Observation e186c4b4-df80-4dc1-b526-5b50baea891d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Measuring Mathematical Problem Solving With the MATH Dataset

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:07.937927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:ec6835cbc07d85f13525acb5a37a3f0f400ef3b2cf79ddfa6d54586c8745728b

Observation a66e1811-fd49-41e9-b8fd-45cfc6627eb9 · outbound

This paper cites Mistral 7B.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Mistral 7B

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.055161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:31414d8be8788f3b708a005a11c7b2403da5b71e4560ca0cf8fec7746490802d

Observation cc3eb58a-ed94-42c8-a088-2ca4bbe6a028 · outbound

This paper cites Mixtral of Experts.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Mixtral of Experts

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.257362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:15247432a79b593c4ad9bb30dbefc6af4e30c497d89cd199843aa0284c2de99e

Observation 8a3635d8-5b20-4c17-9a12-330413744d0b · outbound

This paper cites Holistic Evaluation of Language Models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Holistic Evaluation of Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:08.467354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:0098c01b409ca863d4f2bc5d1db38cc820958c4fd0e3f013a34a214ea1cf898e

Observation 6f640b04-8a51-4717-944e-9497af619c71 · outbound

This paper cites Lingyiwanwu, yi-large.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Lingyiwanwu, yi-large

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:11.788363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:a112749cc4d12a3501e63d463f426c5db7380dbd167b5c43ca6dff810ab65b8d

Observation 6f2251c1-4efa-4ced-a393-b3b2eba62bcc · outbound

This paper cites Build the future of ai with meta llama 3 - https://llama.meta.com/llama3/.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Build the future of ai with meta llama 3 - https://llama.meta.com/llama3/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:12.017351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:2372b5390bc562807ec069d10cd3a46887eaa65e34335b1dbfa6cc955d889081

Observation ea39cfd5-e27a-4d09-ba65-aaca128c59da · outbound

This paper cites Cross-Task Generalization via Natural Language Crowdsourcing Instructions.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Cross-Task Generalization via Natural Language Crowdsourcing Instructions

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:57:29.846007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:9e99ba96308f9242d1f604e3a77895ee3699205378444f0e7c3472a3bd548b53

Observation 14d1b2fb-8384-495c-afbc-8b036362b99c · outbound

This paper cites Levels of agi: Opera- tionalizing progress on the path to agi.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Levels of agi: Opera- tionalizing progress on the path to agi

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:09.094372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:a6e44be5af1c9b44235fc25cceb84785d66acef6a6b70de95c00928cf62e2cce

Observation 45418f88-7c29-4464-a544-f49336cdac0f · outbound

This paper cites Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:12.336355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:6a5e678717418fd2129d424172ab28a16ce6296992bc6744b223b54dfbd3643f

Observation 891b46ca-3b5b-424d-b55f-257ac0ee9273 · outbound

This paper cites Training language models to follow instructions with human feedback.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Training language models to follow instructions with human feedback

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:12.858846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:4c5216eaf2ee5adc1da8c78f2fd3131062bc0e81de81a3dee058800fc3bb6462

Observation c7a9afcc-b632-4e4e-8d02-8dd740516664 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:04:44.937279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:2dc827b3f0a14ae4671d497aee28dd34b77443e625bb37c514d117c7170c0d37

Observation d6e4ec4d-23d5-4def-9ea6-9d2dfe73c24e · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:09.784439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:721580e20695e08c4afa0d58ee257ef0ce972157ef1858bf54a0749ffb4e8744

Observation 215a3e30-542b-4e41-855b-8aa9356a55d2 · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:59:43.395267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:dc5d2da90b8ab67a374a38d077d88d06df94b38530aa32501e362611ada462f4

Observation 6dc02a4e-23d8-4905-8760-e3009390288a · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.108235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:b6ed6c0adb80a27ea530e5a2b01c3f2dcb034eff79262e6ac1c5e84573defb45

Observation 825355f2-d67c-41ec-a68a-825ff5a528a9 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.307004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:0e867121bf6f9e4d48fd2b4806786c727b00caf9e4062e0fee1d107ee49d9463

Observation 42704a8b-0d90-46d6-ac73-82d3f261e5e2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Gemma: Open Models Based on Gemini Research and Technology

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.511901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:3d0d0d7f2f01eec9ce04664a5b0282e55ad0897e753c7bf183a73361c62baef6

Observation 0e9b7867-776e-40cf-9c5d-817aec9f4813 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.654440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:723d4524da50d3bafdcc848faf422a1cb9d2e66f5e6c79215d9f0ac905136719

Observation a54f8831-b313-4f41-8859-7733f6944d57 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Zephyr: Direct Distillation of LM Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:13:57.632488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:22bb7dfba3e2ab46214060b00f251ac56772f3021d7fe42c22877b5a8f3065f8

Observation 9e9eebd7-178e-4594-85da-abfeecc9fe4e · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:24:15.760169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:45db3d442a575fe07bdbd885f4f5dacfe5833c6431ef9f67276e21c14bacf7f3

Observation d3cd9924-9ccf-47ac-b75b-2bdc04dfca44 · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:13.147828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:a3e0803ede4c3f8df7082f91a6ae8e1861d54c6053313cb27db24d8c632c07be

Observation b9e16c97-0023-415e-b086-5c13d04767b6 · outbound

This paper cites OpenChat: Advancing Open-source Language Models with Mixed-Quality Data.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:06.187905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:75b700bb5c69aff2014b19fdc1a68b7a25ffe3bfc553f470e6bf298d81271327

Observation 84c050a8-5ce5-40d8-ba90-e005cc210ad3 · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:59:45.390593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:b91106ef26dce05d1bbedb82c1077d351070dd3cd7cef51355bb9158d0f35c47

Observation 709bf6dd-02a4-460e-ab59-0ad70045c6eb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Chain-of-thought prompting elicits reasoning in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:13.407362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:05f0263e851eca70116887ddb72364308f6ef7dff7c60bb416f64d5f655c4b12

Observation 5d25fc4c-3820-4026-bafc-a91cb5f3456a · outbound

This paper cites Internlm-math: Open math large language models toward verifiable reasoning.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Internlm-math: Open math large language models toward verifiable reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:13.546626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:cbcf3df6992dc271d653aa704c888ae4c6e140f01fa36feb24065b38b0d90d41

Observation 9247c9ee-d996-4f1e-b553-774a98c2fdfa · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:47:28.044136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:fcbbf7ab9210ddd60df50fbee28e282eaad2ff67a2210c3284282926f46e867a

Observation 2b5358e2-adf0-4273-b65d-bdc78f894573 · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark MAmmoTH2: Scaling Instructions from the Web

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:06.893449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:9c5d2ec3b802471693be86edd6b4dd35f83e44922cb10e102ec1a48cbfbfc4db

Observation b8c5ff66-d693-44cc-b20b-0cb6569e2320 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:07.104450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:acea4efad57aad67b67f2205a12ae9acc2f49420453f93dcd42d281ff2f31f0a

Observation f7e6bb69-0747-4fa8-858c-abb009fa0f6b · outbound

This paper cites Map-neo: Highly capable and transparent bilingual large language model series.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Map-neo: Highly capable and transparent bilingual large language model series

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:13.645427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:17ac5a70a7f30a8fc8d5353ad746d73a31fa41e9411107c839d186d1103972b5

Observation 3e6af77f-8187-4e43-9346-dcd1be80e88d · outbound

This paper cites Large language models are not robust multiple choice selectors.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Large language models are not robust multiple choice selectors

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:13.874479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:95efe148c9f14bda5c6a6c313f1046e223caa747bf789966966e71d29835e008

Observation 9b4d48db-aefb-4c5d-8d47-bff912511c6e · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:03:59.570579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:529a6b655f52edc5a949324f5d0b65648701e4eb5bd97275cc98bc7381a7f623

Observation 8fc1faa6-3370-47a2-a4fd-c42d34a39ac0 · outbound

This paper cites image_question.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark image_question

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:14.015975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:efb4cd5e46e3bf887a4f08f055f3d98ddc33d2ba530e7f88d7e40d818965278a

Observation 85e82a9f-568c-42d5-8a63-0f8a970a3490 · outbound

This paper cites - Strain II has an average weight of 15 grams.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Strain II has an average weight of 15 grams

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:14.224331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:02d0dd44957c6fe8288666a30a13a42d6878e9c3742f85bd1678961e321b23ae

Observation c7efa464-4fb1-41be-bd4d-551bbc153f6a · outbound

This paper cites - Each lowercase letter gene (a, b, d) contributes 2.5 grams.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Each lowercase letter gene (a, b, d) contributes 2.5 grams

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:14.336253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:88359a9931b743b1b827db57f03c3061eee5f45e12064ca392e42f172f8e3ec6

Observation d53d232d-80ba-4113-8ee5-618e0d3f717e · outbound

This paper cites - For Strain II (15 grams): - 15 = ( a + a + b + b + d + d) - Each lowercase letter contributes 2.5 grams, so: - 15 = 6 × 2.5 - Therefore, Strain II must have the genotype aabbdd.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - For Strain II (15 grams): - 15 = ( a + a + b + b + d + d) - Each lowercase letter contributes 2.5 grams, so: - 15 = 6 × 2.5 - Therefore, Strain II must have the genotype aabbdd

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:14.343457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:99d69cf06f32fc6ad6a6194668128ad2d0a6c95e4efd3845417a2638e7b831ff

Observation 96060590-f1cb-401b-bfe7-45e8157c048f · outbound

This paper cites an unresolved cited work.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-11T15:51:14.454366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:c44cdeb00d5e5572dcfb82933a79ae6c7929e67b685b98fd52d91b63e747588b

Observation 9b440d32-79a4-4cc9-9cf6-414074ad9ce3 · outbound

This paper cites - Since there are three pairs, the total weight is 11.25 × 3 = 33 .75 grams.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark - Since there are three pairs, the total weight is 11.25 × 3 = 33 .75 grams

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T15:51:10.096972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:14132fd14739a921dedddb2373f3edae04a9ab588fc0046fc4e4355c06c8c18b

Pith citing papers

Observation 5589a9fd-58b4-44af-ab06-1c4b4e693fbf · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:48:26.363473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:b2510a273a64709173823d1ea8063f6c82b40ef692050f28f8c8ad7767085ea9

Observation cbd8afec-9ea4-41ba-b087-9325b32ea148 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:23:27.390797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:9767ab71b04ffb520a29e50b27b39dfda8ba2f301cd9a7bb61751c3199085fac

Observation d5817a28-27f2-4682-ba99-49f6ee4d2de2 · inbound

SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration cites this paper.

SmileyLlama: Modifying Large Language Models for Directed Chemical Space Exploration MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:18:26.949262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:16:39.979945Z digest=sha256:a5a28103756f96a1b2975b30109461b157a5d9a17076590d73349a623eea1f31

Observation fb473fcc-da64-407d-becc-3038ddc8d4ef · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:51:48.416467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:d39c4cbdac89b17b67ab7c5a0d81d02a4c075f60de0516e461b36e22fabb247c

Observation 31450546-1025-41f3-a1d7-210457e2cc3b · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:50:14.022353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:3f943b23cf468cca4c98744dac5c64dd83b4ab4deff2e7cd8beef12bbe5519a4

Observation 14bf5fc2-f624-4d6f-8bad-89df8cbd4141 · inbound

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design cites this paper.

MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:57:40.280362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:56:51.829741Z digest=sha256:609b7795da0aaefe42a885551032370d76fb24299b8386a01a30c1a36418e622

Observation 8718085d-9a5e-4b4f-a63a-211e2601203f · inbound

Qwen2.5 Technical Report cites this paper.

Qwen2.5 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:25:27.932902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T06:25:00.376073Z digest=sha256:2bfe38ef238501e862eac174f877bc21a21e6ab115859c40739318a816b5bb85

Observation b9ffacf9-c1fd-4cf2-962c-989fdc273cb8 · inbound

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs cites this paper.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:36:50.187431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:36:50.060335Z digest=sha256:e6d2de47dbdf39d228692e8ebb7185c35aedabc391ffe1154826dccd32d11ff1

Observation f9a943b6-6e23-42b6-9f29-1c77c448da02 · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 156

Resolution
verified exact
local_arxiv, observed 2026-05-15T21:20:59.341382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:d2dd5600ff83bb456d137645f38bb78deb2ef417d4dd3a3f93b485aec1c29cf6

Observation 95cf7e11-30a1-4de7-bf0b-3c86ec5b9bdc · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.308387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:f4ae64e22fa2dfbe650b92c54329de068cdcc6514d01443378ce0741dd6498e2

Observation 06000166-b00c-4927-90f0-36eb17c4098f · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:2f1fcf0efd12269782af74a142ea29f61a4f2e945b2ea71580572a926025c7dd

Observation cb0a87d3-9c22-46de-8c26-ea71076eb1c0 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 237

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:30:03.042368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:2c70f0574b76c7f5c8869501b7dd3f5d3eddeb00a6761a5b40bf097bc8fc3011

Observation 5a9aba33-b748-488a-9882-52af28f6a352 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 72

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:46:30.121500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:43deacd16e8d02ee9b97eefc8a60feb74a527a35344451e0e6625b1433ebe1b0

Observation 945cac13-4dbf-4c8b-bafd-ca751ce6d86b · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:17:02.817507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:1124afd9d8b1311af1f1a1d1384cc68fbaf58b03f763f71e6ad0410b28a1c3a3

Observation 40de151a-8bb8-4919-b6aa-c15a3dee4675 · inbound

PRIMETIME : Limits of LLMs in Temporal Primitives cites this paper.

PRIMETIME : Limits of LLMs in Temporal Primitives MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:36:58.819218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T18:36:48.376877Z digest=sha256:23b1f99a09b28f8092e4921661e24f7589c1ed6db99393e3071f92b89c3142ee

Observation 21a94ca9-336c-4565-b90e-f089ea37fe87 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:40:25.847444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:0fe93e43c9203c1150b1488c38c04c93895f1c349623d64018db9393bfafb738

Observation add70447-8e26-4327-81e8-9795bd1463e3 · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:a093e17ff07615f42d69c12682cdc016e1d825496a9b5572125b515817cb3ed9

Observation 5874fde9-dc86-446d-8155-58c5b371f3d9 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:958827e7bff2e1b812bf0b6dd072d8dae7e5ea41386e841f58ecd4db69ceb937

Observation 33b53e28-96ef-4d18-8419-7549eb4b4098 · inbound

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement cites this paper.

Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:52:51.881624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T22:52:39.416575Z digest=sha256:d6d42765ba6d2ade6595d464aa8bf3ceb318f248413373e03908588116ecabc6

Observation ec688596-32e5-417f-8af7-837f217f0b75 · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:1cef00c84b8961dc5c80039cf52f10ea5d9c34cab2096951953bbe051b8300c5

Observation 3c90f828-096b-480c-bf77-958bc06d5451 · inbound

Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning cites this paper.

Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:55:01.947238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:55:01.947238Z digest=sha256:89ef497b3e8cf19f004d50058b144fe60843715a9c44e99894189027cc0bd216

Observation c0a51aa7-b234-4598-a0dd-fdd48ef86b7f · inbound

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models cites this paper.

DP-FedLoRA: Privacy-Enhanced Federated Fine-Tuning for On-Device Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T19:48:48.841252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:48:48.841252Z digest=sha256:66c8d0ca47eea550a08e82c19860036f0816ff6cb365290a22c591e520acaab8

Observation 41752cb9-bf29-4457-9506-43f60f78cadd · inbound

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs cites this paper.

AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T17:05:55.034000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:05:55.034000Z digest=sha256:862dc5cb770382142cbdf0737c48559b1e0a7d03f070f6fcd6398e2a4abb39e0

Observation 853481f2-d296-431e-9a06-221d526c880a · inbound

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation cites this paper.

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:57:25.863201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:57:25.863201Z digest=sha256:30f70e22dc5002a077a98480e7855084092b4b18eb341dd4fd7738d594c382bd

Observation a3d3ceef-5195-4d7c-b4ee-8e0cb89b9aca · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:43:11.776534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:af8c7686cc1e59eab7c51deb7c30c4e7ee4f402e2148e8905803e94542cf9590

Observation b4bca95f-70eb-4909-bf5a-99a090fc6780 · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:49.282466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:49.282466Z digest=sha256:15c3b7514659d9ed238e1eb9599aa51aee58003a0baa3d8edf4e44add743e059

Observation 9e592b64-9b71-43cf-a18a-f13e60f0fc48 · inbound

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI cites this paper.

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T12:21:30.884331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:9f963723b3f1608bd29bcb88df658da4dcd840dd73b11c3e51b6f5465f2b8a26

Observation 2d48a48f-0626-4c04-b14f-446aaa904da4 · inbound

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization cites this paper.

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:52:06.100355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:50:53.673037Z digest=sha256:dffd5c089eb564b4089fda060198015d392bce119dd1f8eb101239b08099b8d9

Observation a4e6419c-84b2-4293-a891-e7e7ae4d1edf · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T23:13:15.540596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:9f2cf02bdb30e0d013a0ffc2598d73eff3eabbe15b5b4d18984cec5e36249edb

Observation e6b570cf-1747-483d-867b-ea5c69fabbf1 · inbound

Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework cites this paper.

Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T04:24:00.338380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:23:03.393566Z digest=sha256:f2bbb7fd5f08c5c678770a54f2bb2bd5665396a8b8c6c7e78a1bcadbb0eb34de

Observation ca3ea0aa-fbef-42cf-a952-6f954ac057f4 · inbound

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models cites this paper.

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:05:26.667750Z digest=sha256:50693ed23faa32487cf62bff7beb011301c061600d6d79352a0daec62a3d0eb6

Observation d56cf372-e911-493c-b755-c56914d404c9 · inbound

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models cites this paper.

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:05:26.667750Z digest=sha256:fe66535ee77a426fd881440fd2bf400fe82cd5d53161df195fc4f7d09bd08409

Observation 1bb6e91e-f99c-434f-94c0-21869bc5e718 · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.348223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.348223Z digest=sha256:79cf4c4bf619ede856071abedd417a6c05d40da050bb439e2c1693cf9ca0d635

Observation 37935a84-316a-48d8-b77d-f9260d32168a · inbound

NVIDIA Nemotron 3: Efficient and Open Intelligence cites this paper.

NVIDIA Nemotron 3: Efficient and Open Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 130

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T01:40:42.626283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T01:40:42.190369Z digest=sha256:252898f1c1b5428df7b5f74f4d9cfcd845aa4900b89e362d6f65e0893e935271

Observation 501a36c9-b6a6-42da-83c3-236f8a557516 · inbound

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning cites this paper.

SCALER:Synthetic Scalable Adaptive Learning Environment for Reasoning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T16:28:05.717152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T16:24:04.132572Z digest=sha256:18e703c1c66e2e16bf3ef04f4115540e51532d68d479b616bf0bcfe01d87571b

Observation a43d7341-ecf4-40a2-9701-f2a8a24dab06 · inbound

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models cites this paper.

Entropy-Preserving Supervised Fine-Tuning via Adaptive Self-Distillation for Large Reasoning Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T05:28:15.282398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:28:15.282398Z digest=sha256:b84dff75065808d85a4978ffac79ece1a42d92c71d530a9a1bec99b569de51e1

Observation afa417c7-b8ee-4eeb-99ee-f515404fff38 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:90b628ad02c18fe9d16bff303e6dc69fe2934e8dcb31391fc7654464853a8051

Observation 44893d04-a0d4-4ebf-91d0-affb247be254 · inbound

Interfaze: The Future of AI is built on Task-Specific Small Models cites this paper.

Interfaze: The Future of AI is built on Task-Specific Small Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:48:03.446902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:48:03.446902Z digest=sha256:b5f937cb9a60257d7a777a88204e38df3517a430140662c7403ec1464061353a

Observation 788abe45-388b-43ba-bbf4-97259f34881d · inbound

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment cites this paper.

f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:50:42.985357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:48:35.723474Z digest=sha256:3361e41741876457b7be8fbd90a4abac96a8baf8058f6ac86025c980db579616

Observation b1f77faf-3a78-4387-a714-1653c3f5934d · inbound

Model soups need only one ingredient cites this paper.

Model soups need only one ingredient MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T02:51:56.936714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:51:56.936714Z digest=sha256:11b17d4ad1114aa1552f85b8d177437159a6194c2dd62385cb0cab91fbad9117

Observation c3d75179-36ee-422b-a25e-1390f51e1c74 · inbound

When LLMs get significantly worse: A statistical approach to detect model degradations cites this paper.

When LLMs get significantly worse: A statistical approach to detect model degradations MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:07:25.787095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:03:37.788377Z digest=sha256:be916a54d475bed22dfafa77de74c4ef643a23af9b3b90ba8c783debfff5a18f

Observation ab9cbd9a-41aa-410b-9fb2-8beda59b1d44 · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:00:40.744485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:15fb1e25a46e1993381aa4a8ff451c7bc2562656e90ed1bbe50284411f49db23

Observation b72c55e7-121f-469c-a4df-735975774b2a · inbound

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities cites this paper.

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:14.583482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:58:14.583482Z digest=sha256:cbf2a3079c98f3fd416f639b3c7eaeb1854640e67e0e79fe43c2a5b9ef3d2f98

Observation caeb7d52-bc16-4fac-a8f0-5664799f36bc · inbound

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation cites this paper.

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 349

Resolution
unresolved
no resolver link, observed 2026-08-02T22:31:50.154870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:31:50.154870Z digest=sha256:9596cc2674a71107b93b0f9398a7f2736206f0ecfddb59d89f2b0c6c2c75ba66

Observation f43f4cb7-6c5a-45bc-a725-120c585397cc · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.265009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.265009Z digest=sha256:016f35c9973f78a1b342a8d31c6be5504a3f9217e00f87804f4951d95ea8d217

Observation 2f83f8fb-db11-4b5a-b2bb-7a190450b54c · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:45:26.414366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:5dddd963bada3f2127f23c0e2667525786c33d326f61552b8695239c67e6a724

Observation 80efb0bd-9825-4256-92e4-a6d350c0c80c · inbound

Adaptively Robust LLM Monitoring via Activation Watermarking cites this paper.

Adaptively Robust LLM Monitoring via Activation Watermarking MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T17:39:43.379650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:39:43.379650Z digest=sha256:31cc423e0fa8698c4c058199efd1f01d6c6ab7720493b5ff572326c23e8dcedd

Observation 6a2f23db-da70-419b-984e-dde9af5f0f48 · inbound

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks cites this paper.

Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:28:14.085645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:24:17.788697Z digest=sha256:263ae390ec5ca8b15604d75294a7358d0ca41c5708239f278aa5dcdf36caa5c9

Observation 6c763adf-ed5f-4394-9af3-ae316fca3616 · inbound

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge cites this paper.

Structured Multi-Criteria Evaluation of Large Language Models with Fuzzy Analytic Hierarchy Process and DualJudge MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:13:01.112191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:12:07.852282Z digest=sha256:2c44aa7de56bd5f135da4aef6210ee058446718106b545360424f7b51f49d047

Observation d78aedea-1cda-495f-b03e-bc2ae3a7bcbd · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:08:01.249186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:be9180a347361156920649665145f30cfcca770a97d4803d17290fb1076375d7

Observation a1659414-7bad-4690-9125-b7331fa6d383 · inbound

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability cites this paper.

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:53:22.553430Z digest=sha256:886369c640376aa12bf3ca1913705b145a07a497a30838fac9fabb6b861b4c96

Observation 0ca98958-f293-46bc-aeb3-13faeff71956 · inbound

An Algebraic Introduction to Persistence cites this paper.

An Algebraic Introduction to Persistence MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T08:42:28.957233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:42:28.957233Z digest=sha256:7cf8af80b83e62ecf2ea33821964d42b08a9903e71d9b7945af02a28c4932c61

Observation 4e138d32-793c-4c50-bbc2-89e73ec0668f · inbound

MARS: Enabling Autoregressive Models Multi-Token Generation cites this paper.

MARS: Enabling Autoregressive Models Multi-Token Generation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:54:16.289777Z digest=sha256:cc480fc9e384a09de08a43f5afe0e4eb3431aefa8761f62859c49facad76a2cc

Observation 1d9c6c63-4821-4781-943b-05804e029e3b · inbound

EXAONE 4.5 Technical Report cites this paper.

EXAONE 4.5 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:47:34.692414Z digest=sha256:20309dfc8555b027b56548a41357050dade7531df7db6d2722fda7c8cd57c9bf

Observation 79092f1e-a2f9-4337-bc51-0762e73a6eab · inbound

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning cites this paper.

Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:42:57.596073Z digest=sha256:76ff7441c48c44cf8106ab913bc845920155d3ad681a29a4e3e3f7abe79897f6

Observation 4dfae951-1f02-4ee3-9604-a503b70e2e7a · inbound

SAGE Celer 2.6 Technical Card cites this paper.

SAGE Celer 2.6 Technical Card MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:59:36.852198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:54:24.133637Z digest=sha256:33cc8ca249daf7ae2558c49e007601728f3f2f1f4f081f883dca225c2d3c86b5

Observation bf083123-c623-418c-9b18-41476d5f7bac · inbound

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner cites this paper.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:00:10.325345Z digest=sha256:b932eba39f8f403a0b6441cc2bb05491441a465176f39c5dc10d40824714c25a

Observation c3b268b7-6730-49c6-b7a8-8003c6578aac · inbound

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner cites this paper.

Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:22.150586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:22.150586Z digest=sha256:3aad4b51ad7d652025e999f9522ce9453119e60e841cc194df929a405da9cfdc

Observation ce3e64bc-bbab-4053-b9e2-9ba504ee56bf · inbound

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models cites this paper.

Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:57:48.890310Z digest=sha256:c6a7456224b13636d43acab920569f11544d9ac9d5f7eb5c059795e96430092f

Observation f0e7dc22-40c6-45a3-be3e-5f64771b33ab · inbound

Super Apriel: One Checkpoint, Many Speeds cites this paper.

Super Apriel: One Checkpoint, Many Speeds MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T02:27:11.553553Z digest=sha256:d7219611326e85857a61ace27c14a4b764de29378bfedc5516d647871294bbfd

Observation c3b5426d-f612-473c-83dc-df844eb8b772 · inbound

Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees cites this paper.

Efficient Test-Time Inference via Deterministic Exploration of Truncated Decoding Trees MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T01:03:06.917872Z digest=sha256:d7934dcaf26b346a0fff78114b0f84ce62fe7fe913bfce81618952a4271f244a

Observation 400845a3-79bc-45ec-9893-386c1ea14dc4 · inbound

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows cites this paper.

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T23:52:47.833809Z digest=sha256:309f1e05ad1e896b81e3e07c5210e8084c319866f2692d4fd986505675a9961b

Observation 2cb20aaf-5490-4d74-aab3-b8f6c07abe71 · inbound

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling cites this paper.

COMPASS: COntinual Multilingual PEFT with Adaptive Semantic Sampling MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T01:14:16.831333Z digest=sha256:05b99f603083e26c64885b68444036bb4efc4226b50f9bd262758eea84d74b8a

Observation 45f9dfce-013a-4c58-8ccf-878cbf61be79 · inbound

Decoupled DiLoCo for Resilient Distributed Pre-training cites this paper.

Decoupled DiLoCo for Resilient Distributed Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:20:21.090246Z digest=sha256:33d785018c5d43c57c37c5766af46887316b06d4118249407985003babfcbe3e

Observation 2cfb4e89-8f68-4fcc-bf3b-8e3619686296 · inbound

Large Language Models Decide Early and Explain Later cites this paper.

Large Language Models Decide Early and Explain Later MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:21:09.335968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T12:06:33.200365Z digest=sha256:ba49079c650b365325b1efb814604372132ae323d47a4f5fadd72fe5e94e0b3e

Observation 6418bdca-2c56-4701-8d70-70b495f8ae2b · inbound

Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging cites this paper.

Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:51:24.455720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T13:38:37.045980Z digest=sha256:f3466ff53427e073d9794c151e8cf3d00e458d41e5a9d0badd11713c5766bf1b

Observation 0f200792-a149-4cb1-9cdc-88d7d63c1b3a · inbound

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation cites this paper.

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:46:28.452415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T09:08:04.485654Z digest=sha256:90602add45a13d4f42759f01a79cc46da6f2231f5c0bfb881858af8640b519a7

Observation e47e59a7-39f5-4d39-9311-06e1dc67e90f · inbound

On the Privacy of LLMs: An Ablation Study cites this paper.

On the Privacy of LLMs: An Ablation Study MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:25:05.586464Z digest=sha256:086fabcbd445536aa12439aea047817d0387c9d819664c0fbadad56d2ed38233

Observation c3543f14-d254-456a-bc90-e59de6e7622f · inbound

Analysis and Explainability of LLMs Via Evolutionary Methods cites this paper.

Analysis and Explainability of LLMs Via Evolutionary Methods MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T20:37:40.932811Z digest=sha256:56be4d459451f5bf400e23e4976a749651897d87701522d2710ebf1af4cea3d0

Observation 7c459372-68c1-4500-ad77-d9fc3986ac3a · inbound

MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge cites this paper.

MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:36:06.891413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T15:33:47.276997Z digest=sha256:128e17d84e1caabed56278d02b8651f617d3e4a353208217819aaed815462a56

Observation a1e9628e-5de0-4c35-82db-d6a5ed30e9a3 · inbound

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs cites this paper.

CXR-ContraBench: Benchmarking Negated-Option Attraction in Medical VLMs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:09.522954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:49:53.357083Z digest=sha256:5d984b852d75d7cbb0643438ceb2f49697a188a8952c26843f47a8cbad679f87

Observation 571dec5d-bd09-4e02-b987-0d54233a5938 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:01:17.462018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T12:57:24.822525Z digest=sha256:ffb00526d696d05e5ed823e9dfa3ca3d1910d584f3938cb95342a81f0665a6c4

Observation f5f17757-15a0-4b7f-b6e5-ced88756b3a4 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:51:29.021423Z digest=sha256:15ce9850feafece7ed014a491d804d416e408443669d892814be5becc4d06407

Observation 7005586a-ccad-47d4-b341-0ef14cc2ac73 · inbound

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level cites this paper.

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T21:17:59.726475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T21:14:22.718769Z digest=sha256:69431950b1668986c70c9368f9b182f0a7378b8e6ecc464cae3648f3bac3c7f2

Observation 2746932e-ad2b-475b-9cf0-99433ca2815b · inbound

An Interpretable and Scalable Framework for Evaluating Large Language Models cites this paper.

An Interpretable and Scalable Framework for Evaluating Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:08:25.577363Z digest=sha256:53a2e982c29ac9e06a2fa00378d54b5854c2cb32834210bcf59e5058be895ca4

Observation 309a17ee-5789-4d99-bb5d-11047c569070 · inbound

Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control cites this paper.

Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget Control MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:14.588885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:03:28.231935Z digest=sha256:fbddf18d623231a2322e78fa967908ff05ef0832e92b2830a46b8eef356a81fa

Observation 3ffc6f3c-aebb-48eb-8c30-b12165231b2a · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:28.647789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:34:10.370956Z digest=sha256:3755f1219f359c738b91d164b75b1c8fa9035a29a39dc3f7d91eb712f3412cef

Observation bb546503-eb3a-4299-9ff5-7413e0c653f4 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-20T23:23:51.131371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T23:22:51.808346Z digest=sha256:42744eeb3040dd0c828996c866f5453b08c7be26a9748cea78cf92c229c305e6

Observation 253b5cf3-edda-4978-82e5-d2ae8b28a274 · inbound

LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs cites this paper.

LLM-Guided Monte Carlo Tree Search over Knowledge Graphs: Composing Mechanistic Explanations for Drug-Disease Pairs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T03:06:18.277379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:04:57.158157Z digest=sha256:925ef6cce1c84ca5a073f119600112959185f4f8b4653a72cd0f1d69f5c1190c

Observation 1ad04c0f-798c-4c0c-8e84-b0354f2f0ee8 · inbound

Rotation-Preserving Supervised Fine-Tuning cites this paper.

Rotation-Preserving Supervised Fine-Tuning MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:27:24.449817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T06:26:20.393476Z digest=sha256:30d8bff1cfd5585307e8e81903d1ef65477a89ab247719bd5428692cc96a00ab

Observation 2aa65d82-511a-41e7-9f90-88520182bc1a · inbound

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability cites this paper.

Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:17:02.718420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T01:12:44.476894Z digest=sha256:9397edf1e5376eb6e5ac89119b6c06d18dac2ad173042ac4a3232def8a248ca0

Observation 73b0d0dd-1501-492f-ac0c-8a5e4ccaedbd · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T04:55:01.582395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:8497e04ddc82590755fa9f61ba616109f6189924f9ed5976e3008e07ea27d99a

Observation 96fa0261-503c-4882-9cc0-c9bbbce191d5 · inbound

TeachArena: Are Language Agents Ready for Realistic Teaching Work? cites this paper.

TeachArena: Are Language Agents Ready for Realistic Teaching Work? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T10:31:25.829749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:27:32.588593Z digest=sha256:a2cf471b1aa00abeccfe65f05ba29c54d73d0f71cbd2cdad9bd2b89f0a70e41a

Observation bc718502-e1b7-4c4a-8984-acb4bd8be7db · inbound

TeachArena: Are Language Agents Ready for Realistic Teaching Work? cites this paper.

TeachArena: Are Language Agents Ready for Realistic Teaching Work? MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T05:12:48.986173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:12:48.986173Z digest=sha256:97d380f6b12bddd48713e794c74220b85023403461f71270124ec38a89b6cb72

Observation 85c8b5fd-c397-4ecc-b0e4-63273f66d950 · inbound

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer cites this paper.

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:38:21.603192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T14:35:46.376752Z digest=sha256:8f0cd328f15c9925ba596498bd32710ce5b94f3e698194769b7a165514e0acff

Observation 7bafee52-7982-48f9-89b3-e066ae5ebfe5 · inbound

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer cites this paper.

MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T16:38:30.296706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:38:30.296706Z digest=sha256:5e28dd2f8c25a909e46f50e2f49381f1f9208e4fbcbfc39c061940ee9e53efa5

Observation 9787f1d0-ef75-477e-9f84-53dbdec6798f · inbound

Open-World Evaluations for Measuring Frontier AI Capabilities cites this paper.

Open-World Evaluations for Measuring Frontier AI Capabilities MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-21T06:39:43.714934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:38:51.427985Z digest=sha256:9d485339ae5462b7fd9aad01dac69c299cc57c37126bb361e84b78b35e0416d0

Observation f057cf59-39b0-4044-9dd9-d3d5d2de689e · inbound

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching cites this paper.

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:53:59.027513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T21:53:00.365846Z digest=sha256:9211429440672f7927a5c42545cad5b9c867ff485fd43e9592f11307a03d95fc

Observation ddaf38a6-2acb-40d3-9c98-4bb54e3b53c1 · inbound

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion cites this paper.

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.116222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T08:20:05.081980Z digest=sha256:24a0d209c4ccbae344fbcd948d0ae281195f9f779e1fa3251bc606d70a53598e

Observation 4627ff3a-a14e-4ad9-87fb-766af108731d · inbound

Mellum2 Technical Report cites this paper.

Mellum2 Technical Report MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-06-28T23:02:46.528915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:58:35.397914Z digest=sha256:a402a5e7f1b180d363b26d8e398935e32ff7b030a57d5331f98ab002fc0d2fe8

Observation 7ca2c20d-24b1-4b15-9f6b-18f515c3dd59 · inbound

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO cites this paper.

CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:26:01.079667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:25:04.067339Z digest=sha256:84e191eea67a83d25f66437fe759eb6c716ccf76731bd3f902970ee05e108776

Observation 511d0302-3f36-4d8a-b12e-2ecd0b4189c1 · inbound

SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving cites this paper.

SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T01:36:25.555466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:45:04.176857Z digest=sha256:10834e5ad4f542341adfa15907cced980f73fc187098d790abfd612449bd3aec

Observation 3f6cd5f6-0b5c-43e1-8c60-b263422dd82b · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:16:26.234549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T11:10:06.685591Z digest=sha256:9524b01bef42e9a2e63e36e6d479c708d5201550106f5577648b0b4ef03512c9

Observation 1d580256-1243-4dee-9619-9d1b8345ac96 · inbound

Trading Human Curation for Synthetic Augmentation in RLVR cites this paper.

Trading Human Curation for Synthetic Augmentation in RLVR MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T15:15:09.675778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:15:09.675778Z digest=sha256:f49ff4290b0302ff2a2dbf7323b3668d611cb05b1d0a9f94890220e060400636

Observation 2d6db900-ff05-4272-b511-3f274435a586 · inbound

Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models cites this paper.

Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:16:33.669027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:13:48.860754Z digest=sha256:73dd3f3f1e824e3bd82a73040784a19825d5b6423f8bf9a06e839f445601fcb1

Observation d4bbad33-11f8-4f66-96c4-61e9adfad71b · inbound

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) cites this paper.

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:26:59.401814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:15:20.867740Z digest=sha256:e8bc25d7ee5859cdadc6e4597426d8aa000132d141f7ea73f21372986a809064

Observation 26fa7b81-3a4b-462d-b437-bd9f0af0aa83 · inbound

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) cites this paper.

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version) MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T12:21:43.969544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:21:43.969544Z digest=sha256:ae5250e9585e02010bcde79683ea63d9086efa9aa2da62fc41217a4629bff01a

Observation 4669ed3c-54ee-41b2-ba87-bb23aafd33e2 · inbound

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces cites this paper.

Quantum-Inspired Trace-Augmented Evidence Selection for Reasoning over Structured Hypothesis Spaces MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T17:37:14.595165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:57:20.848868Z digest=sha256:8a934d9e7634c661e744d35e93855434ecb358dbe235d1335dacf4c3e5818dc2

Observation e69f046e-df83-451e-9bfa-37776fe2ac3d · inbound

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers cites this paper.

The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-06-29T14:13:30.249801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T14:07:26.366172Z digest=sha256:8592fd476893c2c20c85bbf53c5e00eb09b5bc9311da15034056043c64ecff10

Observation 3b741e2a-e8b5-491c-b6d6-8258f34bfc64 · inbound

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese cites this paper.

Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:07:17.918479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:40:16.052239Z digest=sha256:046a4a87402d7aff9c6f8f90654dd687a2bc35ebb397ad6bfffe974dd3c0b4af