Pith. sign in

Paper Citation Record · LEDGER

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 32 inbound Pith citation observations for arXiv:2405.00332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00332 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 32 of 32 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:10.454387Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a5a5ce5c-fedb-4db0-b1e5-5aab38d9b175 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 229

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:49.823811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:55dcc4c8406c5221ad5ab8c8e6c48eae4f47a09db44df4dcc43dde06800ee4ab

Observation b68c51b3-4b36-4ad5-a1e6-5c9a3b33c4b4 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:48:26.380848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:7803e10a880c21c02a8b0c97f123c0366cd62d77b71933d9ef17e7fce206eabf

Observation bb1e3955-e78f-47d7-a156-1a631e15d31c · inbound

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models cites this paper.

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:42:11.994710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-15T00:42:11.891829Z digest=sha256:ecabcfd290024a736e94f8ede5cbb8c7ddb4b7b806bfe892967a92739e27627c

Observation ea7c5a03-6797-486f-abab-df38a4791870 · inbound

Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts? cites this paper.

Do Large Language Models Perform Latent Multi-Hop Reasoning without Exploiting Shortcuts? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T12:53:38.959056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:53:38.959056Z digest=sha256:ab5b0f920b5a42ed9c4214d35efecd9862e4e843ca34e4cb0922e8628be2c208

Observation 19155c10-d286-4379-bbb4-ca907877e2ee · inbound

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge cites this paper.

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T05:56:43.280659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:56:43.280659Z digest=sha256:ff31423edbd01a9b17562399e66177d303208b638f9cdcf17eeba732d8944dde

Observation 5d1392fc-f062-4c28-ad58-286ae3a831b5 · inbound

HARP: A challenging human-annotated math reasoning benchmark cites this paper.

HARP: A challenging human-annotated math reasoning benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:34:19.737720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:34:19.737720Z digest=sha256:ee7f07d63894dbd2b620a3779ab5ea6da4928af95a4487ddf06b4530c69afc4c

Observation 3e601c16-9fa5-4f2c-9195-1860ba8669e1 · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.310224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.310224Z digest=sha256:5ea81b767d665736e1a4c702807d02f4428fdfda002c4ebc9190812d910cc0be

Observation cd2e9f8d-ffb1-4668-ace2-09e617f8a2d7 · inbound

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark cites this paper.

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:36:56.402576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:36:56.402576Z digest=sha256:22a00a81f69d5b117853d38950639f3a3f31a5c267e4fa22176467eec65212ec

Observation 14a60b25-e9db-49f9-bfcd-7a075ca3dba9 · inbound

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation cites this paper.

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:14.371302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:57:14.371302Z digest=sha256:d6b343ba5af852010b512056b36e8d1944fcedaefb699f62e6de18b6cc9f915f

Observation 29d7a177-961d-447e-b8a0-95c4f6d97ac3 · inbound

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models cites this paper.

UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T15:40:39.912031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:40:39.912031Z digest=sha256:d8e1d9e768d4fa2c3e415c5b8f631e3bbbc19394bff7e7ae2164d7d35d122110

Observation 2b38f87b-4619-4d65-bd20-2c4c45f5b783 · inbound

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models cites this paper.

UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T19:27:46.905311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:27:46.905311Z digest=sha256:25eb915c87147b5a1fd06fa39368eafadd87d1c2429203a346fc5eda05e2af0e

Observation cbf57a64-effd-4191-b7df-e58cf849db4e · inbound

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations cites this paper.

MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:26:07.800182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:26:07.800182Z digest=sha256:5353925b2631309707282290d1bbebe1fb6b773f330d2a5e12e27655da460878

Observation 0b58a3d6-b9b1-42c3-896e-a3f41e1f58e0 · inbound

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks cites this paper.

A Survey of Theory of Mind in Large Language Models: Evaluations, Representations, and Safety Risks A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:22:45.097354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:22:45.097354Z digest=sha256:31e4d4dd8e46e5974946760ff0688d56f3abc8ddd929587035fd7a68b1f9b2ec

Observation d3e59564-af7f-40ee-bac5-de8cd31c81d3 · inbound

Investigating the Zone of Proximal Development of Language Models for In-Context Learning cites this paper.

Investigating the Zone of Proximal Development of Language Models for In-Context Learning A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:11:54.680317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:11:54.680317Z digest=sha256:a9078408d544d711e9846b8afcf6cd79fdb66a0f11951d4fdee6618364f616fb

Observation 2dda6688-f56f-4220-a0f9-58efe9f38463 · inbound

Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators cites this paper.

Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:24.021886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:55:24.021886Z digest=sha256:8ec07af8a70b3abd60e2275fa50eb9d57f3ef334d8a28874f377f5c0c3edf651

Observation 95ca64ea-383a-474a-a83e-d399c0dea817 · inbound

Towards Contamination Resistant Benchmarks cites this paper.

Towards Contamination Resistant Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.454387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.454387Z digest=sha256:0b458a0869414def7c93148c2d44864acef3d7751c719a665ba3dcebd70e0e91

Observation 9a2fe13c-2714-4032-b3f4-6bdb30c1446e · inbound

MANBench: Is Your Multimodal Model Smarter than Human? cites this paper.

MANBench: Is Your Multimodal Model Smarter than Human? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:36.030597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:36.030597Z digest=sha256:b8e530bd0dca0fc7dba5b5259ba1539d045e3971d993e1aac15d040d51d9227c

Observation 1925b323-2750-4627-87d1-38c89723b7df · inbound

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards cites this paper.

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:08.467524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:16:08.467524Z digest=sha256:b058b9fae8a07ba745b063cadeea731b95dcac7a4b38023886db57b27a389880

Observation 2a527933-1989-482c-8436-6261da33ad6e · inbound

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind cites this paper.

The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:49:39.921570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:49:39.921570Z digest=sha256:b8d9c13c7e3368d451d6205aa59499e1e357528fb603805ecd1fe4c4e26f17eb

Observation 69e784c3-d604-4142-8de4-1b074e2caeaa · inbound

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? cites this paper.

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T13:48:53.795243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-12T13:48:53.691192Z digest=sha256:90343822064d1be1c62ba3693cbfdeea4dd5e1236c6fd5fc0f25600031132656

Observation 092b449a-b293-4b35-928f-1f367d0db27e · inbound

The Economics of AI Training Data: A Research Agenda cites this paper.

The Economics of AI Training Data: A Research Agenda A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:40:11.782844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:40:11.782844Z digest=sha256:2ae763ec0dfde75cd4c191ab3d872148756b29b0577404a4527a84d01ac96dd5

Observation 00d1a1d2-e08a-4cd7-a64c-669875314ea8 · inbound

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) cites this paper.

RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025) A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-14T22:03:02.990783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T22:01:21.565986Z digest=sha256:dfb9d3b92fe9b8d21b00154fe45180e6994d709654671c084da06c1e687a4525

Observation 5778955a-9027-4da2-823b-f94c99130be7 · inbound

BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models cites this paper.

BitCal-TTS: Bit-Calibrated Test-Time Scaling for Quantized Reasoning Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.385086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:01:32.816855Z digest=sha256:6b40b99946144bb051aa33aee0e86f87ac79a7fbdfddc6a27b8ee34c9ac5b05b

Observation 88bcc9df-2775-49e3-bc84-bf4f955def28 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:57.298388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:b8d70d9e1b8242a531473e0a46298d3163acc769a9c3cfb8b380a10f077f8df1

Observation d69e9fc3-e29d-429e-b78e-c04a4cf234eb · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.456819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:b841bb3d76cf604efaec7b0ccec7a66017ca82803e72a1cc5efb8faeecb79e84

Observation eda4d12d-5f19-484c-9152-41a6d7234b22 · inbound

RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents cites this paper.

RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:39:46.928751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T09:40:00.507836Z digest=sha256:2deb99a96d4a721c69441badcd2705fb9cea7b72266d4ad535973179c09e26c8

Observation c6a67174-8f9e-4167-8441-651f5ae44099 · inbound

RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents cites this paper.

RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:45:35.434123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-01T07:12:07.099036Z digest=sha256:58743765563d2d38d1d313cc8c69cdccf4550bcc3260540e86d466892cefb880

Observation 6cef796c-8357-47a6-8a42-605166c1475d · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:27:18.687360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:a5012b53e50962a5245cef8737937d36f0a43c2c0f0fd82b6d3da63792bdaf10

Observation 40289591-417c-4a50-aa07-395d7ed23850 · inbound

TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation cites this paper.

TLA+-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA Specification Generation A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-30T22:39:33.984319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:39:33.984319Z digest=sha256:0c654817e7cca6fb861df5952d4e6ae1d73fd22b519411cf6644e6ec41888b45

Observation 4a7da1d2-2edf-481c-8d62-2e0872c13e61 · inbound

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models cites this paper.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:51.475326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:51.475326Z digest=sha256:95f13547d9da72c431f66a728595f61b1e4095367dc0f3cf2d8b6f52a8d36c36

Observation cf2f86c7-c6f3-492d-bfa4-9deb7ac76e34 · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.425762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.425762Z digest=sha256:d5e01a4fed3da16d3801aeaffd8e2a7d63c36750851d317151136c00a3c06edb

Observation b3051fb7-1264-46c2-ae22-5c473129c3d2 · inbound

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination cites this paper.

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T06:02:27.990201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T06:02:27.990201Z digest=sha256:8961f53233316369a12307e2c59dc4a39d560e05855a763ea4fcf89fa53c3932