Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Large Language Models for Math Reasoning Tasks

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2408.10839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.10839 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:05:03.187958Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:33:35.971242Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5d12f69-25cc-4900-a383-8c02d634fc85 · inbound

An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts cites this paper.

An Empirically-grounded tool for Automatic Prompt Linting and Repair: A Case Study on Bias, Vulnerability, and Optimization in Developer Prompts Benchmarking Large Language Models for Math Reasoning Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T17:14:06.591219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:14:06.591219Z digest=sha256:e24e3c95f862b4d88b6fec450f3edde5a628756f631dbf35c838850879eb9911

Observation 7474b217-b984-429b-a102-0ce9b547f260 · inbound

34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery cites this paper.

34 Examples of LLM Applications in Materials Science and Chemistry: Towards Automation, Assistants, Agents, and Accelerated Scientific Discovery Benchmarking Large Language Models for Math Reasoning Tasks

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:05:03.187958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:05:03.187958Z digest=sha256:a3c33a56c29f583c13bb286130a4df4a6469796c6ba1a1fea7ad3959f9c7efcb

Observation 29164d6c-7811-4749-b999-aaf74366e062 · inbound

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains cites this paper.

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains Benchmarking Large Language Models for Math Reasoning Tasks

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:33:36.068998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T11:33:32.804825Z digest=sha256:1ed9edbffd01015f6b369012a9de25615914fe561583a1200aa0b99ff7a79848

Observation aa8ae639-4b41-4c10-be6e-57b6d7b1fa75 · inbound

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials cites this paper.

SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials Benchmarking Large Language Models for Math Reasoning Tasks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:57.077734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:57.077734Z digest=sha256:d18f08fde9a677573b93e677a6ee76aec01ca383c65dff37c9a54bb33edf157c

Observation 913b81ff-ad3d-4393-9a90-73dccc64e506 · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Benchmarking Large Language Models for Math Reasoning Tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:03.146717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:03.146717Z digest=sha256:7aa2f59c7f2fa3d9efef500c34c2c23e79abf5d3f8785b48bdf338f5719de524

Observation 6036ef5c-4d9f-4230-892a-109d6ee7d315 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Large Language Models for Math Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.307485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.307485Z digest=sha256:f24be62861e20a7a45caae4bf4612a063d97acc7e2a7e14bb9ff81de51e46731