Pith. sign in

Paper Citation Record · LEDGER

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2412.15194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15194 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:27:26.145226Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a54cd0d1-3536-4eff-83f8-93f98202586d · inbound

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities cites this paper.

ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:26.145226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:26.145226Z digest=sha256:74ece65bb64dc3e6033b52a5fdb723372d568e8539c59c628e8923a964600ee6

Observation 5e135f8a-94d1-4679-acea-e22753019b80 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:47.983056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:47.983056Z digest=sha256:e97ce1d11f614a536f5edfdb93290cb6463accbf60846295854cd29bdd735b3a

Observation e689eac6-1808-4201-a231-73f69a312a71 · inbound

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models cites this paper.

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:25:36.032153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:21:39.237267Z digest=sha256:391c8cd1180bc87e2dc7d243e53a3bc6484a565ed71f8e83619cd3962ad269c6

Observation 309c7c76-6f33-4ca6-bc7e-32a15024d37e · inbound

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration cites this paper.

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:13.183812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:53:03.420926Z digest=sha256:614ae6bfe6e9f6da052ee6bfb2daec253f8d96a5e208087f8bbf97288189c1b4

Observation 3ee3fae2-f0b0-40c6-b5a5-b7954c930d0f · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 180

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.471191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:9d41b08e9985e91402a611439005943c95f4077511e241689dd99d18f3f099e6

Observation c8b490c7-7a77-43b4-a1d2-0ef8d2636d42 · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T01:28:50.524986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:c1e17df294f3a03574cd0a6c2141acefd15fab3a4278a7e274d31a6348471837

Observation cce7df17-337f-413b-8be4-643d567840e5 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:2bf578ca82d588308e227af8b8ec004634858bb0b0afb5a0b6377efdbc2cd963

Observation e9e5c36a-ae8a-4a43-a059-dc7cc848603d · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.871998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.871998Z digest=sha256:e3b1ef81d2bf79610279652002ceb559bbf93a5e02292393c008fc5a04f4d6ae

Observation 468f8bbb-9c8c-448c-accf-13bc3d193157 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:30.646103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:30.646103Z digest=sha256:eed314c0991697b982abb0470931957322e2ea13b1afbfd4c0ed83bf6d4ecb51