Pith. sign in

Paper Citation Record · LEDGER

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2504.17665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17665 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:28.303524Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62a960ec-5e32-450a-96d9-08c7fe067a1a · outbound

This paper cites online" 'onlinestring :=.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.159490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.159490Z digest=sha256:c9b7f66da1898d2f402ed98760c57cd225959d486dc8248bff50d2abf6725fec

Observation 506503f3-8331-40a1-9a64-ba58128b6830 · outbound

This paper cites write newline.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.256611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.256611Z digest=sha256:8a7c309ed5f289dad8559a79618153d7278c39745556dc9647d70a2ea59cd622

Observation e79519a5-0baa-460e-a67d-5feeff3aba02 · outbound

This paper cites GPT-4 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.277342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.277342Z digest=sha256:85896990a0593bdac521a7c070c67209487e2aa4a608f8a61ff2093d100b830f

Observation ff58524e-b33a-4208-8c31-6305b3699408 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.285575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.285575Z digest=sha256:931ed833c5e02531205963412573089e3d7154a21284fdc5fae4668565b6f3b1

Observation dada4289-5d98-4b00-82b1-e2e86c9b2d3f · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.305525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.305525Z digest=sha256:03c1220c62fb843c786660f179ed069a4d34b0bc3ccf0bbb3352028d7d48ecc6

Observation 2f794079-b4e6-4394-8ddf-6b4cdcceb854 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.911559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.408830Z digest=sha256:75d9f1f1362ca4a3524d9170cb0b2d22e2a61f7977fd73b01a7610567009926c

Observation 2d75e65c-7674-4401-aa0b-92df0c40e0f7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.435610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.435610Z digest=sha256:18d40d457f3962cefeee5bdc3af2110499707f6cfee1357efef68ae17b79d955

Observation 32c92f81-8e5f-4c35-b37f-9c1392d60637 · outbound

This paper cites MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.447288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.447288Z digest=sha256:4088ef00ba2f3bf74dfe1faf2aef6f65cdfd34ed77be20054b602fa34da7a064

Observation 4f422dc0-0654-44d6-bdbc-bda1a0d83758 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.534750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.534750Z digest=sha256:d519a1917cff847ab64d25f0bcac5f57809f82170638f5dce3c7dc6ffc8ad39b

Observation 8749a97e-aa6b-4800-bce8-8b7d742c5859 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.580437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.580437Z digest=sha256:de161fe1bb9cf3c1bff1e828f77876eeebc2af204de6ec76b20fa7a34f606a6c

Observation 7b4b1caf-91ca-4e0c-bb3d-20effa818813 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.842314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.678437Z digest=sha256:5f91ac05fbfa20926d1e7147b1bf769baa06e7de943f126c3a06bcef8b62692e

Observation 1d544743-a921-459d-993d-eb76b5a5ada0 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.705902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.705902Z digest=sha256:983b603cc190e4e054ddcd3b685aa52d63178fa6898f9f777323f5179b65f46d

Observation da02a7ba-5c01-4ac7-af52-ff898ee082b6 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.721732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.721732Z digest=sha256:f56fdb3d7aaf3870bc6d4608876349ed5f46b8ba9ff44d4692663bf249af7598

Observation 7b7701e4-daa6-4806-b3cb-94f779d066c7 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.731733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.731733Z digest=sha256:af50d2d26d4432fccadcf5f8e5cc7107f2b5f6115987c7dad68273e950542ffb

Observation 83f28c52-37e4-42fd-abdb-b02d042ec52e · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.817668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.817668Z digest=sha256:bd0a19c0aa144a64b951101e02fc7b420996898e290e96372ad9070605179b13

Observation a1e98042-fb50-45d1-aea0-72ac99a2a6e1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.856738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.856738Z digest=sha256:54e281b4b5a9ca07ff94571c504b9c128976479492797a4f66aa34b9c4715253

Observation 7fbaa48e-c3f8-409b-b597-87b03c0d2784 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MathPrompter: Mathematical Reasoning using Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.861159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.861159Z digest=sha256:963433a6072357bc9654012cfde685bdfa78d2966efadd302f49382672d5155b

Observation 8d49385a-1479-481a-a344-7af00853427a · outbound

This paper cites How Interpretable are Reasoning Explanations from Prompting Large Language Models?.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics How Interpretable are Reasoning Explanations from Prompting Large Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.876660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.876660Z digest=sha256:94ce7b43ab19f5fe6fcd2dc3d89138e2f03c02b1d3a8970090f3c527f9e8bdfa

Observation ac83fd1c-703b-4bb0-af44-4e9f7c456e1b · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.884720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.884720Z digest=sha256:b40d6086c810b529e40f22dd2fc043529a24c2af3812fd1d39a58f31a33c2d34

Observation 1b945a6d-a71f-40d9-a7c0-fc80f0ffb998 · outbound

This paper cites Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.976338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.976338Z digest=sha256:8a165c1b374cc7e8e9b9c8cb23237cb255710349e9d76355d46810be865204a8

Observation fdf7c8da-a15a-4dfa-8579-c6ceee6a6bf8 · outbound

This paper cites Let's Verify Step by Step.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.008119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.008119Z digest=sha256:df781322ca6e5cdb971bda9839baddb8267860690dab08676ea6eb740477c0c9

Observation a5df8063-c254-4073-ad11-dbe39036e38c · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.012469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.012469Z digest=sha256:553bdcd867ee73ef5374f68fe6600bd7e6ff6fda80e03194ead837afd5f69855

Observation db3513a4-b966-44bd-af55-159ec022d21b · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics StarCoder 2 and The Stack v2: The Next Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.017095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.017095Z digest=sha256:22cbff52162c9651984b6ee91bab2fd8ab7eec48e0ba6608c31cc2de3128a2dc

Observation 4cf74b68-30c9-40b6-8623-6fd74e7b2993 · outbound

This paper cites Faithful Chain-of-Thought Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Faithful Chain-of-Thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.021493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.021493Z digest=sha256:7c0ec82518d0b90927fb24ad04013f3cc3e72dc9a1012b239c916f536bd8933b

Observation 9ef4e6f1-ae80-4668-b1f4-36bf68506952 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.025866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.025866Z digest=sha256:7c9f0efd4e81b73a0782db457d380ecffb45ab10aef6a8f96233eac5684791c8

Observation 8453e4c9-49b1-47bb-8ecc-f7b35f7bbb24 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.069599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.069599Z digest=sha256:715dcb40ef0b60377dc8630fb90cfdb099d46aa2d7e73b551c39e902f2d0b227

Observation 1e4e2223-d431-4ae9-9af0-a9e31ef49758 · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.135870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.135870Z digest=sha256:cf30f107880af162669367ec8be410b08dce353d7bd37253dc5e32c432df42a6

Observation 10a79d37-2910-49f3-aff6-d11fd8c59d4c · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.140228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.140228Z digest=sha256:4259ea980638e3354e194b1ab953358bdd49eb93e591d4601a15a3081247b7b5

Observation 63a95e6a-6204-4928-98c8-5864781550c9 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Code Llama: Open Foundation Models for Code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.144373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.144373Z digest=sha256:5dd98e80640511054979512f07499479e8193ac830ea8bcbf55845bbdaabbd42

Observation 280b389d-db6e-44e7-bb5d-e48cbf25c1b3 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.148351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.148351Z digest=sha256:4531becbcddfec5a56af551816ed03dffd6f42f782897a9285091719b8a4166b

Observation 870b4498-931c-41cd-a6e7-abbc8839b71d · outbound

This paper cites https://lyz-code.github.io/autoimport/ autoimport.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics https://lyz-code.github.io/autoimport/ autoimport

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:38:28.666950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:28.151866Z digest=sha256:8e8fc28e0ff8dbc964e04f9ef3f46b455527576a2dfcff42c0289e14fff9b703

Observation 444747c2-aba4-48de-b594-2650b1891a9a · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.157210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.157210Z digest=sha256:944fe63892d9b6b79356fedb8d10faef29ccec993154d787180879c873173b93

Observation 0572d174-3895-4ae9-b9ed-00e10be25196 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.161671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.161671Z digest=sha256:0a5217abe3108fd61e82a4195aaa73864c1de30334ad3b1c8cbf5fddd0d8baa7

Observation 37de6bae-2ec7-4da3-805b-3b2633deb28f · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.165082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.165082Z digest=sha256:2735193a9b66185ee815dd896d836859e1e459abe25643acbe381a4f03da2960

Observation 386764f5-7076-40d2-9519-fa12012ec13d · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.255976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.255976Z digest=sha256:72c7bee57900d9b967b86b35209ca335fed67e295896ffdb48a5161576d4cf6b

Observation a7e1db0b-a0d7-4091-b74b-704b71752da3 · outbound

This paper cites CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.303524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.303524Z digest=sha256:d4dff48277e553b295259682a0ae97f573e8a63dc9b5c137e9e115a2708a1b74

Pith citing papers

No inbound Pith citation observations are available.