Pith. sign in

Paper Citation Record · LEDGER

Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2402.19450.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19450 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:05:07.629687Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.903971Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da975d38-1082-4a7f-8cb6-a4ee7715d135 · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.488804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:eb40ff4ae575f01941fd7dd2cef96c317160b88cfbdb1de07f1787f99c6d7fdb

Observation eab653d2-d2a8-4d6f-a349-0f105c367b7c · inbound

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models cites this paper.

GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:42:12.049878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T00:42:11.891829Z digest=sha256:f76416c906d4f5780537243e10ef92c9c08b3d138fa963d90fa567bf7373324d

Observation f329c6bc-ee4e-4f40-a39c-393130f64d25 · inbound

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark cites this paper.

ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:05:07.629687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:05:07.629687Z digest=sha256:44e7d75ceae6a53aaf139eb401c137667fae328376bbe3eec07c28e684dd125f

Observation 3826905d-f638-4ed1-aa2b-97dc4f193be1 · inbound

Probing for Arithmetic Errors in Language Models cites this paper.

Probing for Arithmetic Errors in Language Models Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T16:58:59.669192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:58:59.669192Z digest=sha256:ecb4476dc1f3a19974d7c36ac36ac10a38dc1452504798997ddfdca3e75ebf6e

Observation fac53161-6d61-45da-b0ef-7080273f2620 · inbound

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving cites this paper.

EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.703761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T15:08:03.793304Z digest=sha256:4a0a7c2cccb9f23ad48fb92035880e05e7914dfe63cbca0c00e6da74bccf084d

Observation ac9ef06e-0d3a-4efc-a2b4-f1446c679a46 · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:00.854306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:2dfa684bb63cfccbca4c4b8236d71bf4a17f21a295cc23dbfa105d4c0bb3f9b7

Observation 36339f75-5eda-42d9-a948-c9fb76348a22 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:08:21.713534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:06:08.545334Z digest=sha256:7f34bda1f08d21a9f2a16d1c33d860efb4a0fb6bcae6fb703d87052bac511e1d

Observation 6374d269-33f9-4537-8891-e3976752897d · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:21:28.862792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T11:18:51.642405Z digest=sha256:1e444b248a2b2fb6a1d60c0879d74c37be09303bdceb6ecb6e7afcb993995870

Observation 3e85165a-702b-4235-8e8b-4d0b8e12e8b4 · inbound

Robust Reasoning Benchmark cites this paper.

Robust Reasoning Benchmark Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T17:25:30.907575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:25:30.907575Z digest=sha256:88bde24c08f5f00a298acd2dde77fb0d9b8286499bea9991b6fc5031a6c693c4

Observation e40167f0-3073-4d43-b3d3-dcb104f12179 · inbound

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling cites this paper.

LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:40.203346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:33:30.850113Z digest=sha256:f66f327ae5cfee227c42d6d5be0fa752ef2018e5863ce0f0adc8198916bfad00

Observation 730dc534-79cb-4f5f-94b6-fd545a8a024e · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.905611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:39:44.696961Z digest=sha256:1192feb0aa83d41ec6495b7664dba218d78b8a6e837238fd3d16dd8e948c3ed2

Observation 9f4da5cd-79e4-435b-9d6a-4608fea40c7b · inbound

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery cites this paper.

Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 259

Resolution
unresolved
no resolver link, observed 2026-08-02T12:05:18.483003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:05:18.483003Z digest=sha256:e10ff7175151c5bddf1110de42bcd401732b1991b2b5ca650bd6cb5a44316346

Observation 6f402b3a-c51d-487d-ac1b-8d2bd36d255e · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.704626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:35:32.667865Z digest=sha256:86c9f78130afc353650434a607a276e9dd345da6acee5c2c7d60269e24e7406b

Observation f775b9b4-0b01-49ef-b1c7-7cf72a1db984 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T09:35:02.975384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:35:02.975384Z digest=sha256:11b57159a88b332eb70970265dba7b1dfd52f97f88a7bc2cacf555cf0f6be0c6