Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2405.00332.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:10.454387Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation a5a5ce5c-fedb-4db0-b1e5-5aab38d9b175 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 229
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b68c51b3-4b36-4ad5-a1e6-5c9a3b33c4b4 · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb1e3955-e78f-47d7-a156-1a631e15d31c · inbound
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ea7c5a03-6797-486f-abab-df38a4791870 · inbound
Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts? A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19155c10-d286-4379-bbb4-ca907877e2ee · inbound
INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d1392fc-f062-4c28-ad58-286ae3a831b5 · inbound
HARP: A challenging human-annotated math reasoning benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e601c16-9fa5-4f2c-9195-1860ba8669e1 · inbound
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd2e9f8d-ffb1-4668-ace2-09e617f8a2d7 · inbound
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14a60b25-e9db-49f9-bfcd-7a075ca3dba9 · inbound
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d7a177-961d-447e-b8a0-95c4f6d97ac3 · inbound
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b38f87b-4619-4d65-bd20-2c4c45f5b783 · inbound
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbf57a64-effd-4191-b7df-e58cf849db4e · inbound
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b58a3d6-b9b1-42c3-896e-a3f41e1f58e0 · inbound
A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e59564-af7f-40ee-bac5-de8cd31c81d3 · inbound
Investigating the Zone of Proximal Development of Language Models for In-Context Learning A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dda6688-f56f-4220-a0f9-58efe9f38463 · inbound
Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ca64ea-383a-474a-a83e-d399c0dea817 · inbound
Towards Contamination Resistant Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2fe13c-2714-4032-b3f4-6bdb30c1446e · inbound
MANBench: Is Your Multimodal Model Smarter than Human? A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1925b323-2750-4627-87d1-38c89723b7df · inbound
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a527933-1989-482c-8436-6261da33ad6e · inbound
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69e784c3-d604-4142-8de4-1b074e2caeaa · inbound
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 092b449a-b293-4b35-928f-1f367d0db27e · inbound
The Economics of AI Training Data: A Research Agenda A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d1a1d2-e08a-4cd7-a64c-669875314ea8 · inbound
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5778955a-9027-4da2-823b-f94c99130be7 · inbound
BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88bcc9df-2775-49e3-bc84-bf4f955def28 · inbound
Dataset Watermarking for Closed LLMs with Provable Detection A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d69e9fc3-e29d-429e-b78e-c04a4cf234eb · inbound
CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation eda4d12d-5f19-484c-9152-41a6d7234b22 · inbound
RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6a67174-8f9e-4167-8441-651f5ae44099 · inbound
RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6cef796c-8357-47a6-8a42-605166c1475d · inbound
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40289591-417c-4a50-aa07-395d7ed23850 · inbound
TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7da1d2-2edf-481c-8d62-2e0872c13e61 · inbound
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf2f86c7-c6f3-492d-bfa4-9deb7ac76e34 · inbound
Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3051fb7-1264-46c2-ae22-5c473129c3d2 · inbound
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.