Pith. sign in

Paper Citation Record · LEDGER

MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.12209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.12209 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:12:38.725841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T20:13:24.799209Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a4fb8dfe-28f1-432b-9c51-6b8ead269c3b · inbound

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection cites this paper.

ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:13:24.804839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-23T20:10:59.264484Z digest=sha256:226aed2b5cba11371885dbfd3587deb60bc6040d99033a0b29477b34b4c869e3

Observation d02a36cf-61f9-4022-be44-d6e2618d4d31 · inbound

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems cites this paper.

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:57:13.405426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T22:55:34.238427Z digest=sha256:d473248a89193d7fe8a74a99a5d6d7de5f52c20ee5b1f9a37353cbd02ebfa7ff

Observation a4ac3de0-180e-4084-8b91-8ad34a3a22a8 · inbound

Computational Experiments in Number Theory cites this paper.

Computational Experiments in Number Theory MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:00.891133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:28:21.017594Z digest=sha256:8a07c14941f9447c76e864ed6b20238720f4e18ac666f8d76d12c8519431b776

Observation 2593c843-c04f-47c9-9431-139f71010cd0 · inbound

The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants cites this paper.

The Avengers: A Simple Recipe for Uniting Smaller Language Models to Challenge Proprietary Giants MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:38.725841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:38.725841Z digest=sha256:1f260ed612a703942b96e13d02bc8f563b4531a6367bd5ea745af97803659a79

Observation 74b9d6e3-901c-4cb0-827e-7c7c600334f7 · inbound

Evaluation of LLMs for mathematical problem solving cites this paper.

Evaluation of LLMs for mathematical problem solving MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:06.681146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:06.681146Z digest=sha256:5719d14f196c746a308a2956efee5ee1ca0363b5703181097469c5de7489c82e

Observation f9ff443f-a8d3-498e-b19a-149234f32a6a · inbound

SciDA: Scientific Dynamic Assessor of LLMs cites this paper.

SciDA: Scientific Dynamic Assessor of LLMs MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:18.377180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:18.377180Z digest=sha256:90cdd9dff6e90a2ea554f665733486a1251f5b9fbf09a86f92c3822a18b3df6c

Observation 193dd8cc-9d80-4f51-8211-c7db35f3a691 · inbound

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models cites this paper.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.838439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.838439Z digest=sha256:135464d63ac2dec001da3ab66c75f6864ab4612d3f80563424239c348616029c

Observation 2549b930-ffc8-4b3d-934c-a762e4745d65 · inbound

Potemkin Understanding in Large Language Models cites this paper.

Potemkin Understanding in Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:30:31.067924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:30:31.067924Z digest=sha256:2b1ebe4419e232245101b391f6618da9d5804eff65a127b3341ba82bcf5fa1dd

Observation b3be25f3-592f-4891-8055-e6dcb953c31e · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.235673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.235673Z digest=sha256:2330a2bf3aa7774cf11e9708bb41996034103a6207722464771abe4edfa97e82

Observation e8e54732-adc5-433c-a87a-d9e101bb78f9 · inbound

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics cites this paper.

StatEval: A Comprehensive Benchmark for Large Language Models in Statistics MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:23.844792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:36:23.844792Z digest=sha256:7c2dc9f1574b83e09e356f853dacd1aebdbbdf11547a66d4ec1bee922ce9f6f2

Observation 2701aed0-f1ed-42b5-a67d-fe4931bc0d23 · inbound

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy cites this paper.

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:12.995111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T08:53:13.427364Z digest=sha256:e4d4b42c82cb213f7d4c8bfb2469380949d19710ca8091e3fb308e8ba8a2443e

Observation e6a6050f-5bc0-4367-83fa-007eea97f816 · inbound

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy cites this paper.

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:53:53.592710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T00:50:11.410735Z digest=sha256:87bcdf1d61256974a00684a435f233a33bc8a2cdf578a5f86246302f43495d08

Observation d1a36c05-5cba-460a-9733-094f2a14c761 · inbound

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models cites this paper.

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:58:12.831388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T08:53:18.670312Z digest=sha256:bed82de1c58d621c1ffec59ef670e85ef0bee1a2dc54cd8bfb5952ff95dd7cb7

Observation 05f4609b-a3dc-4fc3-aae1-217c8bbbd5bc · inbound

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models cites this paper.

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:33:52.825418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T00:29:19.064290Z digest=sha256:c6a297f51c745268e3fc9b0ae6c6ea55889ac56c8bb356b9329189baff9d5995