Pith. sign in

Paper Citation Record · LEDGER

Are We Done with MMLU?

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2406.04127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.04127 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:26:28.825227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:05.653512Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8628d4db-1796-47ba-87cf-8b9561bda533 · inbound

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing cites this paper.

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing Are We Done with MMLU?

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:58:36.846203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T06:58:36.684583Z digest=sha256:fdd4f8b7aa9ff6be227a554076175245242bc879759354537284cb8552570c37

Observation 1bc2a4a8-1bae-43a1-bbdc-60a90eea74a2 · inbound

Qwen2.5-1M Technical Report cites this paper.

Qwen2.5-1M Technical Report Are We Done with MMLU?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:26:00.070159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T05:25:59.964619Z digest=sha256:ca1b0d901c0aed9724a7c512350af3b73a180a4ff351feeea182b3514847cb95

Observation 719eaea1-b645-4883-af87-a967570318c7 · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report Are We Done with MMLU?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.082104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:1251e9ee7c51b3f3c7a3720e38ce2e87af2877e120a34d5be4c0e7c27056f546

Observation f84b5750-48ae-4ae1-ae93-243fca0d0d55 · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report Are We Done with MMLU?

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:54:03.335098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:2abb2935bbe54d0c083a599fa3455f52ce903816d6721453a115a2cc7d3d0788

Observation 2f74c340-97fe-47bd-91f8-12ded746f83d · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report Are We Done with MMLU?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:28.499063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:4c905b78bdf2702e58f7aec03e8c309922b3afabee8a226ac9fc7cf959d13c26

Observation 80715de0-fd37-4ba8-9111-0e32dab9e258 · inbound

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource cites this paper.

Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource Are We Done with MMLU?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:05:47.753599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T00:05:08.916339Z digest=sha256:80cf65235e127554011c51cd0ba911f163c882ca825035c11aad26f1f7391aab

Observation c33513d7-bfc7-4987-b91f-2045fc6269e7 · inbound

Kimi K2: Open Agentic Intelligence cites this paper.

Kimi K2: Open Agentic Intelligence Are We Done with MMLU?

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:49:28.118074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:49:27.926646Z digest=sha256:5791fd3c448d2bf86db345bbdb95208df40959898255d33e3502ed38cf8fa30f

Observation 446b2309-0e25-4f51-801f-16872e3053de · inbound

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead cites this paper.

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead Are We Done with MMLU?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:12:55.724599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T02:12:48.586913Z digest=sha256:0b5b4d56f25896e4cc75d19011db200c91c3679f4bac940f436012baa4ab43f8

Observation 28221249-82a3-49e3-866f-6d6e97f7c153 · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications Are We Done with MMLU?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.825227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.825227Z digest=sha256:286d8a80038c7dc2743a62cb5412841cc9ff76b2e218d8843215def78c3e25dd

Observation eee64261-ae94-4e30-b07d-679bc36d1498 · inbound

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks cites this paper.

The Non-Determinism of Small LLMs: Evidence of Low Answer Consistency in Repetition Trials of Standard Multiple-Choice Benchmarks Are We Done with MMLU?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T05:30:11.831165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:30:11.831165Z digest=sha256:a1b148be87d6668dda335f2e499e09dcedf2ac679ba33426687360285dd96b09

Observation 197a9686-1107-4dd4-9a08-22e38a42be64 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Are We Done with MMLU?

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.411386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.411386Z digest=sha256:c9876af2f8944458d08ee82d1b5b841a7d03bc358ef708c526eebbe09f037343

Observation b667ddba-ab63-411d-9ccd-4f54b5bc5f1b · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Are We Done with MMLU?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:37.762357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:a3a417a218e5a110b77507a555e26c8f3c18681d0c971c6566c986fc6c0e80de

Observation ed9b015b-6b73-4953-b781-6081c3887c81 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Are We Done with MMLU?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.089956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:0cd37a8aa386d10c1b8df7e6f6316fd13ed52fc0cd8dcf0a0a1370b313d67164

Observation a7aa0688-3f8d-47c1-b3cd-dbc431932f70 · inbound

MiMo-V2-Flash Technical Report cites this paper.

MiMo-V2-Flash Technical Report Are We Done with MMLU?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:33:32.685167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T11:33:32.568261Z digest=sha256:9226d097469aef5be04437d0006cfcdb95ad708563e72ec33306e3ba046578e3

Observation 8973a4ee-ab89-474b-8c19-1c139bb6d251 · inbound

Ministral 3 cites this paper.

Ministral 3 Are We Done with MMLU?

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:12:24.776096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T19:12:24.627033Z digest=sha256:b56dca93637fc04bc55aab78c7c6d3d46389f7f1f2b2765c28584f34e8e1ed7e

Observation f4d28fa1-cf1e-4b17-a6f4-5dbde74f30ac · inbound

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining cites this paper.

SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining Are We Done with MMLU?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:00:40.711571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T05:58:03.113220Z digest=sha256:4dca4627ad39eab1595f33d35d9b76f29fa3b103bd4dfc725d6817d637d39f2f

Observation 2f30c67c-4b40-4b9f-83f2-0a35fc1ad07c · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report Are We Done with MMLU?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.498727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:f968828c226205f944a1ef8b2113ad72a54127885316e800ee0963d0d2e94663

Observation 15d84056-6f9e-425a-adb0-4e908ade3958 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Are We Done with MMLU?

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:36.632600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:e8151230892879c0f6ebcf2fb4c0d7f2e300e36b9b8a2ce3876e53dbe27fa252

Observation 440f7544-e51a-4987-8ce9-bbf591c80ef4 · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Are We Done with MMLU?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.208995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:34:10.370956Z digest=sha256:2706b4b9deada0b8f1eac32b34b064093b85e0e7f1994111d2536832fade8d1e

Observation 9d6350f6-144f-4a71-9b26-5f8c7200c00f · inbound

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training cites this paper.

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training Are We Done with MMLU?

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:23:51.225707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T23:22:51.808346Z digest=sha256:7bd7be4f3c52be237773f3b5140497c1ed4904737f645a486bf2a8e5fc5a711d

Observation 05ef44e4-7ff3-42dd-aa34-700f014e4aa1 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim Are We Done with MMLU?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:18:25.373549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:8c63f2c1b8af97750569fd308acca75aaa3b19301e81a8e59fc7e205842cfe6e

Observation ec0f4e96-df34-430b-87b3-cd0616800691 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report Are We Done with MMLU?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:23d53eacfdd8abb1e5cb551a0f1b42718b0b453b931b1b973eef031bd770c316

Observation 22961b7c-399f-4369-bfba-4e623e25d18b · inbound

Mellum2 Technical Report cites this paper.

Mellum2 Technical Report Are We Done with MMLU?

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:02:46.486953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:58:35.397914Z digest=sha256:23a2ca74921cf0e18e5ef55b8a193e0bed29d7b915e7fdcb9be2ac82af62a995

Observation 60ad93e2-6461-4c72-8427-cc59753e5a69 · inbound

Knowledge Index of Noah's Ark cites this paper.

Knowledge Index of Noah's Ark Are We Done with MMLU?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.276774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T06:33:33.539031Z digest=sha256:9327a17036ccf4ba05259ff913d00e125890fccd0763529d262594f49a6731a9

Observation 6844d1a6-624d-4896-96e6-779545bc6fb5 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Are We Done with MMLU?

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.957396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:34979b24daca1a260537613f37d4b104a7cc56b4a300e05b19ac2361c759eada

Observation 61ee3a97-ccc9-43f5-bef0-0a1980c88df4 · inbound

Improved Large Language Diffusion Models cites this paper.

Improved Large Language Diffusion Models Are We Done with MMLU?

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:05.655596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:34:45.620481Z digest=sha256:1219298c8f3ac8b4ea1a6cec619dba764a5956a234a3752d7f0d1f7e740d450c

Observation 435631d0-995c-4e93-9bad-d0337f6e532c · inbound

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety cites this paper.

Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety Are We Done with MMLU?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:52:55.778039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T00:46:03.210076Z digest=sha256:8bd9cf550da56803daec7889383d51ed3841dddd72af15594db77d04e1b1a481

Observation 2ad79e6f-baf2-4ae2-9574-8a56a4f471aa · inbound

Out-of-Distribution Generalization of Risk Aversion in Language Models cites this paper.

Out-of-Distribution Generalization of Risk Aversion in Language Models Are We Done with MMLU?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T07:15:51.883780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T07:15:51.883780Z digest=sha256:2159f07a5a0a2278d9dae3c829fb38c23e918f1c411bff6565d034cb8617c0c6

Observation ea9799b4-72b2-4307-a48a-0f45694db7d1 · inbound

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment cites this paper.

Faster AI, Uneven Frontier: Rapid Crossings, a Jagged Frontier, and the Repositioning of Human Judgment Are We Done with MMLU?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-15T07:32:36.632338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T07:32:36.632338Z digest=sha256:98ea8fe0b67db50b15bc2a26b0a3514eb5535a7ff9a5485bfed4ace8db3de531

Observation 9c6e7cc4-dc31-4e15-8447-b20525205fe9 · inbound

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements cites this paper.

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Are We Done with MMLU?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:48:02.409682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:48:02.409682Z digest=sha256:25b42652c4ebd7dea7ed181f078f7a5a8bd7c0e63559a070293e405bd522f74d