Pith. sign in

Paper Citation Record · LEDGER

MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2409.02257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.02257 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:16.178310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:27:56.414590Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1655239c-4031-40f4-a026-1ef5bfe7dc84 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.292773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:7ad7047fb5a98d352847b4a4b449bb31457e6a2f23f5cabce4eea632d67f774e

Observation f7a0cb8e-86be-4b03-ac03-4c9eb57a57ad · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.178310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.178310Z digest=sha256:cc61ddda81eb4d93e10fe462856aa994db834959c8444b1119bc20669eebcd4b

Observation 017184d3-39eb-4eca-aa18-9bf6a83c9bcc · inbound

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence cites this paper.

Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:15.598649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:15.598649Z digest=sha256:469b5df17ac88bfa91796b94e2ac587f53ab616303944de347bb0331c6db4492

Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · inbound

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks cites this paper.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.424893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.424893Z digest=sha256:40c546fb2bbbc24da08d5d262096626ba70e84068112f2abe497c82ff36c3591

Observation 563518c4-1ddd-463c-b440-57685e23b800 · inbound

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks cites this paper.

Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:13.013412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:13.013412Z digest=sha256:bc2e1d1649c7e75f85135eff87c2137c04fa41ea939df11615dc0df9a35c7279

Observation b946c5ab-c600-4479-b564-76e9885faa7d · inbound

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents cites this paper.

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.415938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:01:45.332920Z digest=sha256:ad42ee4b8a8e6b00731a1f5168f3eebcbf51727149d3d87d99faf14265f1df51

Observation de5ad9f1-22a0-4a63-b019-4a7a6bcd612b · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T03:01:56.282390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:01:56.282390Z digest=sha256:17f2f5a236df6b70a8dd9a90c0d6ac177a063754bfa2a876de3d3c5e50c93453