Pith. sign in

Paper Citation Record · LEDGER

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 10 inbound Pith citation observations for arXiv:2412.15194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15194 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:36:56.441502Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:17:29.508961Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f0890a7c-e109-4e22-94aa-ed6a2403c747 · outbound

This paper cites Phi-4 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.182704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.182704Z digest=sha256:1b95c80566a4ae07b9c9e130b4b5231251959a6570ba4439aa65bf252540614a

Observation 26b5a11c-43e8-4dfc-a1e8-91a3c016e4ad · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.188989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.188989Z digest=sha256:28f55ca62ce2831b8cd0e05566e6d3188a3859d49bd2bb66aed591451bfa3cfc

Observation 7ce5e8b0-b892-476b-9ae6-767594d0a466 · outbound

This paper cites GPT-4 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.194251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.194251Z digest=sha256:e1182609d56ba2e71c50f300e9df0293954b7b0cdfe580226f1eb8deed2027d7

Observation 58e88894-dbfa-49fd-845c-31fd7293e841 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.195843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.199571Z digest=sha256:d2e52e620af399d82e1decb95b4cafa0d91d9d76123f564931b62cdcf0286b80

Observation 91300db8-8176-441f-b4af-564e235d1e33 · outbound

This paper cites Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mmlu-pro+: Evaluating higher-order reasoning and shortcut learning in llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:57.181341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.204262Z digest=sha256:8c7945040d90f2ba75c073179bdbdd3bd5449b70c95e9d9a7cddcc4a48031922

Observation 4b6bb381-41e9-4fe8-aafc-ab2b2269ec75 · outbound

This paper cites Qwen Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Qwen Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.209191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.209191Z digest=sha256:bae7f7bb65f8a145108a8f8639fef4b7f1fb6931f6ab4a2911b4ec59936e40aa

Observation e6017d45-8fb9-46c1-b964-1ccb76709906 · outbound

This paper cites InternLM2 Technical Report.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark InternLM2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.214345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.214345Z digest=sha256:6ceb13f3f1808ff839622e768bb1b0cbe735cb459ff6c73a69d7f7f4370ef9c4

Observation 789a33e8-a9fe-49a7-ace9-b86a526694d9 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.219910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.219910Z digest=sha256:d0f8923fc2f288cfd4f1029b0d76c39cac2c74775184582512c8e6252d5481f8

Observation c3e537f3-adcd-4bd2-8d4b-e3e5527c7f11 · outbound

This paper cites Gonzalez, and Ion Stoica.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gonzalez, and Ion Stoica

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:57.157066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.224464Z digest=sha256:ed8f9dfdff87338786a03191f3d887c4af387305421c13df618c7b5793800a7d

Observation 27005d57-f694-4ae8-a7cc-a79ffe4bbd7c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.229210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.229210Z digest=sha256:0142160adcb0096314d5143e9fca3667d99a04c0579895d0a0c3b41515b1e875

Observation 4e4aa229-92c2-4e5d-ba14-dd8ef3b069e4 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.234166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.234166Z digest=sha256:51b74f519f3483e9519e023b8369928dc5e39003d10c138192ca44fcb992826b

Observation d9d890cd-546a-4cb3-a802-2e80e0234596 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.131661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.238498Z digest=sha256:46b63d2711e95eb3bf3f9c53e1b424f314d61cc88598879b7971d71311a819d0

Observation 4a3c1869-3fd5-4951-be65-8c3cf1198ccf · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.116782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.242949Z digest=sha256:c7e086f6c7e44100084f9d550df1bc3ae007546cceb4c5ae7b0a4003eb87adb0

Observation bca4d33e-5b27-4f94-9232-b6fe7992985a · outbound

This paper cites Are We Done with MMLU?.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Are We Done with MMLU?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.247364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.247364Z digest=sha256:013b175f9e15a4f96f626ecb0643004de41e2534f290f131e28f3cd3a3253627

Observation 405bd8cb-7834-4264-9d64-744e56d71f40 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.252020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.252020Z digest=sha256:1b2afff5a554350b5c881294b63f653fc928f6cf0fee8d62d43a7d83337a7289

Observation 18d512d1-bc91-4b72-9a74-b05d9b4d891c · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.256785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.256785Z digest=sha256:f875c92dd6de4f469450f4d62828375aabf32ede54d8296d443c0c3008e9596c

Observation a19bbbb6-d29e-45bf-bdab-8992c9615c43 · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Changing Answer Order Can Decrease MMLU Accuracy

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.261554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.261554Z digest=sha256:78387c6fd88048bca9710558673c584391d392ac2e1225a0002cbbcf2349fa17

Observation f8b19299-6703-4008-ae5e-a42b1d25934b · outbound

This paper cites Measuring massive multitask language understanding.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Measuring massive multitask language understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.266772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.266772Z digest=sha256:ce2cbc1bd35f8e5fae3eb3bc28187c769ea7b4d08351a5680aea1c5f57d396e6

Observation 98d789d2-915a-4377-9665-dbffb3a5dfde · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.090793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.271444Z digest=sha256:463ba856302d236474a1b4dce030320197d3a8e0f44562ff7f1939e23e988af7

Observation 30223fc1-2011-488f-81ce-38f977fad982 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.276076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.276076Z digest=sha256:1b636cc686c8a0d2221d7f115ef53486a7b64726b1f028cc01dcae1a59111eca

Observation 0961f84e-7a9a-420a-99a2-ebbaf7cf50a8 · outbound

This paper cites Mistral 7B.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.281024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.281024Z digest=sha256:a3b80d7104eea8751c9a3556db04a60a434cf3081b57710c5a1c5417b20c218a

Observation d4277810-f35f-445e-9f87-c5fd122b15b1 · outbound

This paper cites Mixtral of Experts.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.286099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.286099Z digest=sha256:6e2284a3290ba53575a0661dce44a4a87b2eefe8c83d81ec1205420623825d06

Observation 880b56f8-6794-4cbe-8d17-c9d57c88c580 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.291029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.291029Z digest=sha256:a874ab58d1d77763ed7db94411591e63ed94ba5db043ea8823f03687f6177432

Observation ee3ad207-45b6-4514-83e9-8c2afaffda59 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.295602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.295602Z digest=sha256:0c43bedbe4020f2e8b1545d5f6e67979fddacc7b5672c9956408c94b18836c38

Observation 2bd19fd2-1054-43db-84b9-b78a72dfade6 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.056345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.300303Z digest=sha256:097e4036e001ec831edf049bb0f6c3d06a19be3929a0207f65d957917e0f9284

Observation 4352b1cd-594d-474a-a045-fb2623415971 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.040129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.305291Z digest=sha256:5a1cffc95ed3c14ebf6d97ed060a3561cb789444e51527c0f32bf424f2bda732

Observation 6df5f9b2-41f5-463e-add9-14873c991060 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.310575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.310575Z digest=sha256:e56531424e1987122822da0297c44d07caea0baec85328bb83bb3be29494d58b

Observation fe50c6cc-c533-40e6-83c5-e5e82b2f2d73 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:36:57.014963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.315740Z digest=sha256:6b0931adfb5aaa303d178d8ebda23d6881cb48d5143ec5baf342f5e1af8901fd

Observation c243d202-091f-4772-ab7c-ead6da3e5bc2 · outbound

This paper cites Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.320629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.320629Z digest=sha256:4d53c89627db575835e708d5de555ec67de5df2f14d377b69541f754d0e46e94

Observation af01fa5c-200e-42fb-86e2-e97e5d88f467 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.325452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.325452Z digest=sha256:0f94d3aaa2a07924d13c1368741c00348b281f7f2c79434487455c70c68e4131

Observation fbc63022-021b-4b00-929c-376efddc9047 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.330309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.330309Z digest=sha256:b7b75c68c7e28a8f421c705dbde26295710f1d917e9fe9a92adb5bbad0783763

Observation 63284d49-6886-4e5c-b468-f5268d9a7b40 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.335120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.335120Z digest=sha256:3e9218808630377f65f45f31e4b8eb0c5838e1b363e856b44d9a3ccb4398473f

Observation c7234562-76aa-40a6-8fd8-ab989e73170a · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.339386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.339386Z digest=sha256:f8d9e5d7e136d6f0c1109d5eee879cbfd828b2f4552477ce9e0d9d8e8ebf5806

Observation a0cde24f-07c4-4f69-8bf9-a7c7692aebfc · outbound

This paper cites Detecting Pretraining Data from Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Detecting Pretraining Data from Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.344809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.344809Z digest=sha256:768999c0e0da6b8cdf3755ad1c127fe880f9ec2148e3102e3810b8185ba62497

Observation d63148f0-2746-45a7-a3dd-a1e910d835eb · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Gemma 2: Improving Open Language Models at a Practical Size

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.349638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.349638Z digest=sha256:4dc1242a682096fd7b59dd0692fa3c490f10662cfa18cea4898fb46d711484f7

Observation f70f6830-4f06-41e5-9bd6-e163fe394dd9 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.354281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.354281Z digest=sha256:7b601cb28c052c961fe1dd6c5fa5b38695f149128c2601360e6ceb82d0163978

Observation 5b175d5b-d8af-443c-82f1-1ad156104a58 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.358716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.358716Z digest=sha256:625efc1aa02509a17b6ed9deee8ed4bb723f434466c913d59e76cce4c06378a8

Observation 6153dfd6-d71e-4df5-9f13-6244ed606ea2 · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.363525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.363525Z digest=sha256:b721b1c37af627a41fab346bb28d8440727c4f7f305e26d3798472c016727b38

Observation 017b45bd-a3b3-405d-bacb-17e74bc2b72c · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.368110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.368110Z digest=sha256:2bac098e0215dd16cd8574d9ec04ba65243f23a9bc99c1a0f19b9a0e9ffea460

Observation 76973005-a47e-482f-bb3f-1ab01ddfcd28 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.373073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.373073Z digest=sha256:7f93a469f9a108499933f383e59a25d308fb75ea08d994c83ffff4538235fcdb

Observation e18a5227-3918-406d-a583-2a6aaa86fdd0 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Baichuan 2: Open Large-scale Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.377964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.377964Z digest=sha256:7b8b5f67b1f8f18eb6581fe78e4b36454baf7c898c39640ec706ea5c9a30ce25

Observation 08f10b90-0c52-45e5-b59e-f67d6f0a9ac9 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.383219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.383219Z digest=sha256:63e868ac8dc9b83e9345d4487cb73f979d4fa17691553a1a4543b0978ca7141d

Observation 3bf5126f-538a-4eff-a437-4f85b734d514 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Yi: Open Foundation Models by 01.AI

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.388284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.388284Z digest=sha256:604c76107cd6ab036acf2a5e4707489d1407fc0c5eb7f7e5b3c5253628faf3c5

Observation 213ca611-061c-4588-a1fb-35235ef37cb8 · outbound

This paper cites WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.393090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.393090Z digest=sha256:7a5840be2420e197a29c3a48c876835de9e83a0120a5c784b58a65f3454d67c0

Observation ccf77de9-01e4-4f77-bc08-92db4da8477e · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.397912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.397912Z digest=sha256:8e5fee2fcd52575e854f5ba76f7fc983ca1e7fa95e63ebef6fd148d77adbe202

Observation cd2e9f8d-ffb1-4668-ace2-09e617f8a2d7 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.402576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.402576Z digest=sha256:c0522c9cb9889be848bf1fceb83a262b3b345916ce6c13f0b6f05acce02b14f4

Observation fbf53d32-dbdc-4755-b585-f5e1011b97b8 · outbound

This paper cites xLAM: A Family of Large Action Models to Empower AI Agent Systems.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark xLAM: A Family of Large Action Models to Empower AI Agent Systems

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.407462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.407462Z digest=sha256:0bd1d7a60af5d1a71afbfa0d18f79a83bdb5e5054f953281d5cb448b6fae73e9

Observation 4f1471d4-5cf0-41a7-852a-1882ae3f10af · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.412708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.412708Z digest=sha256:19670655037b5b63d5d5b5e9cd1a1b5afdd20fd808c4fe4a2265584c6ad4471f

Observation efdc6df1-d7be-49ea-b549-5f6873b797ba · outbound

This paper cites an unresolved cited work.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.417153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.417153Z digest=sha256:6f85dd626057e1f7dfab883a07b0cf942e69634ed3fa68002ee4434134600df1

Observation 3044dc93-ad54-46d0-a910-2df71408c75d · outbound

This paper cites CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.422014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.422014Z digest=sha256:115d5fcc412ecefcc2b971894bb1f81fb86eeeb6bc69505c37ec9abe05c5dd1b

Observation 5c87fd0d-bc73-4c2e-ab97-6e0c3501b2d2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.426938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.426938Z digest=sha256:5d2d47b1f1723d2abc3d06ece9e9251dded81d1e1eaa4cb3443384019718b6d9

Observation f16f36bc-8bae-4637-a97a-d5ff722f0c77 · outbound

This paper cites Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.431608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.431608Z digest=sha256:5cb983400604d1d4c121082945e04a529d7069f7b9284996fe494ed4761da8df

Observation a973ea3c-e569-408c-9957-a2c7e4a653d0 · outbound

This paper cites online" 'onlinestring :=.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark online" 'onlinestring :=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.436392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.436392Z digest=sha256:3604d1e6d8e847db847d7118ef25961fc1654e139cda3350be2f6059ddc1f442

Observation 00c10ee9-dc97-4649-8e5b-12b23cd840c1 · outbound

This paper cites write newline.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark write newline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:36:56.933825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T11:36:56.441502Z digest=sha256:7f9ffbc6e592d0c0e43387121779d161add631d7a8d76f4f767d7250e6f70f32

Pith citing papers

Observation 90a15906-3214-4ead-8b2e-11f4b91575ee · inbound

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism cites this paper.

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:29.508961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:17:29.508961Z digest=sha256:ff7e17ab194585383a18fdeb7153392a75cade5d4074142ff749c13af966080a

Observation a54cd0d1-3536-4eff-83f8-93f98202586d · inbound

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities cites this paper.

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:26.145226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:26.145226Z digest=sha256:ca324a9ef25a7dbf2d7f1920d641f81ad9ff36db376650c6c04ac86bec93c284

Observation 5e135f8a-94d1-4679-acea-e22753019b80 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:47.983056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:47.983056Z digest=sha256:2a40a2db4a574b1f28d6002775ac0aa2e45dc6181174beb87afbb3a622fde5ab

Observation e689eac6-1808-4201-a231-73f69a312a71 · inbound

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models cites this paper.

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:36.032153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T12:21:39.237267Z digest=sha256:951761f4d7dd701099eed4dbdabc9be494c3b99d569594c8ccc1af9e9c463092

Observation 309c7c76-6f33-4ca6-bc7e-32a15024d37e · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.183812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:438198e54cd00b4069eb561612d6553378df65bfad637328bd3c33407954c17e

Observation 3ee3fae2-f0b0-40c6-b5a5-b7954c930d0f · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 180

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.471191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:f9d17e94f69303f6b59f9aea2324b0aa0d0be4b6fc3cbfe4af8737bebe8373ff

Observation c8b490c7-7a77-43b4-a1d2-0ef8d2636d42 · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T01:28:50.524986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:16f941a7688af9c1af2dd338ff794b0193edd9b5e7f85af1971b59341fa3c332

Observation cce7df17-337f-413b-8be4-643d567840e5 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:9cd6f2368a4177c2c79aebaed6e8afabf86010c762ae88cf0cde38ab1630306b

Observation e9e5c36a-ae8a-4a43-a059-dc7cc848603d · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.871998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.871998Z digest=sha256:eb11e49f37dc66cce4b20d54024a7cae24a32eeae76d297bb53428d879358349

Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.646103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.646103Z digest=sha256:5534b15b5f9b52d2c0548d4185bed946875c3c3969d73caa6b4503c5ebc69810