Pith. sign in

Paper Citation Record · LEDGER

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2504.17665.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17665 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:38:28.303524Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62a960ec-5e32-450a-96d9-08c7fe067a1a · outbound

This paper cites online" 'onlinestring :=.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.159490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.159490Z digest=sha256:59242bf53427b746d3bc3a7a96140a489a015286d928067fe6bd674b7d23d56c

Observation 506503f3-8331-40a1-9a64-ba58128b6830 · outbound

This paper cites write newline.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.256611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.256611Z digest=sha256:8156236eb355ef54c95b1f16b13a09e785a4fe8295081d1343e8232a6fb42159

Observation e79519a5-0baa-460e-a67d-5feeff3aba02 · outbound

This paper cites GPT-4 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.277342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.277342Z digest=sha256:ba38ff8e711c1db71e3a9ebbc57248661478dfb063b019f6ffcf6bca50b671d8

Observation ff58524e-b33a-4208-8c31-6305b3699408 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.285575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.285575Z digest=sha256:5a7c6c5cd0bdeb97ecdbf26a10ddacfb36df836a79c4f023c85443d4eeb74102

Observation dada4289-5d98-4b00-82b1-e2e86c9b2d3f · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.305525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.305525Z digest=sha256:c5f22d0fbe1e6e87a8c0b3b4b3b9292926a4a63bb0012441d9c6f12e5ac35759

Observation 2f794079-b4e6-4394-8ddf-6b4cdcceb854 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.911559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.408830Z digest=sha256:dfae046adb0c5ba108fd13ed6192018f94db12fb116599d78e55d4480c7099cf

Observation 2d75e65c-7674-4401-aa0b-92df0c40e0f7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.435610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.435610Z digest=sha256:71f5bf58d1047fdaa769cd83d67f14b1b68fa17f7521068d8f8db96006c17a9f

Observation 32c92f81-8e5f-4c35-b37f-9c1392d60637 · outbound

This paper cites MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.447288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.447288Z digest=sha256:397f1fdef9b57918a743fb6be001b34cf16fe3668a7d49914d2b527520ff9701

Observation 4f422dc0-0654-44d6-bdbc-bda1a0d83758 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.534750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.534750Z digest=sha256:7010f19c04bb6636aa33f231eb39c99b15ce3cbb931dd9deebdc9866e9aaf027

Observation 8749a97e-aa6b-4800-bce8-8b7d742c5859 · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.580437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.580437Z digest=sha256:8372d535656a8630195a42d210722c405a1db0964279d1b61d5caeadf772aac9

Observation 7b4b1caf-91ca-4e0c-bb3d-20effa818813 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:38:28.842314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:27.678437Z digest=sha256:0cf4500722b4688edbcd1527983215c5a1f980de2bf639c5ae8981528987540b

Observation 1d544743-a921-459d-993d-eb76b5a5ada0 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.705902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.705902Z digest=sha256:63b2ad99d06277a2f14068fe85587f9126ed98494f395c4fd94aeed56cddf8e3

Observation da02a7ba-5c01-4ac7-af52-ff898ee082b6 · outbound

This paper cites ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.721732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.721732Z digest=sha256:2c39215d594382c6c07963f308e99ee0fd0466395c19c401cd98d08c653e457a

Observation 7b7701e4-daa6-4806-b3cb-94f779d066c7 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.731733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.731733Z digest=sha256:e39abfb7007b9c8858aa885a0456b2d6db3345a2ca0efa71bc5215bf41be9d39

Observation 83f28c52-37e4-42fd-abdb-b02d042ec52e · outbound

This paper cites LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.817668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.817668Z digest=sha256:8b552f958b2eda14739e8042830e8608121b698f6b38aa2db0e2c62a9b18472f

Observation a1e98042-fb50-45d1-aea0-72ac99a2a6e1 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.856738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.856738Z digest=sha256:af29a3a19c74eb9eb0b29b4740060b9a65738dd917489d6afc7460357497095a

Observation 7fbaa48e-c3f8-409b-b597-87b03c0d2784 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MathPrompter: Mathematical Reasoning using Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.861159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.861159Z digest=sha256:9f66d8a3446033e04043cb9860790e02cd07ad0d3050b7d97e5eb7d00ce3d01b

Observation 8d49385a-1479-481a-a344-7af00853427a · outbound

This paper cites How Interpretable are Reasoning Explanations from Prompting Large Language Models?.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics How Interpretable are Reasoning Explanations from Prompting Large Language Models?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.876660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.876660Z digest=sha256:692bfb21fed412c99b4696663523811732dbcd1ce46a2ee92c9936411c646fe8

Observation ac83fd1c-703b-4bb0-af44-4e9f7c456e1b · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.884720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.884720Z digest=sha256:c5c1adf13e2db2fd4698d13a55cd6d31d8064d2f86175f80a8b342937ea2785e

Observation 1b945a6d-a71f-40d9-a7c0-fc80f0ffb998 · outbound

This paper cites Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:27.976338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:27.976338Z digest=sha256:7291fcb89a81e1b79dda3e9b39980099cc59848d8e35e9ef50cdfc1dd56051bd

Observation fdf7c8da-a15a-4dfa-8579-c6ceee6a6bf8 · outbound

This paper cites Let's Verify Step by Step.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Let's Verify Step by Step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.008119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.008119Z digest=sha256:b7e44dd01c44343234ec7f84aa0e8da124a9575f0b01bd75b755f1fed327a54e

Observation a5df8063-c254-4073-ad11-dbe39036e38c · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.012469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.012469Z digest=sha256:90d4eb9f4edd76cca6d6282d15f1282bcd507799995765a778ce9877ea629a35

Observation db3513a4-b966-44bd-af55-159ec022d21b · outbound

This paper cites StarCoder 2 and The Stack v2: The Next Generation.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics StarCoder 2 and The Stack v2: The Next Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.017095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.017095Z digest=sha256:c825735a2e1d85d700daf8f1a55dc0358bd066def27dc8f8d2caf82eee342add

Observation 4cf74b68-30c9-40b6-8623-6fd74e7b2993 · outbound

This paper cites Faithful Chain-of-Thought Reasoning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Faithful Chain-of-Thought Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.021493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.021493Z digest=sha256:30c9e46f8ba76d0d38f79a9b1c08b080153ce19b262eac76daf5bf6e20892212

Observation 9ef4e6f1-ae80-4668-b1f4-36bf68506952 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.025866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.025866Z digest=sha256:ac7bb4275336b44c0f6e7135ddd4e76cce78f488e057520b8dcf5ef6b3247f84

Observation 8453e4c9-49b1-47bb-8ecc-f7b35f7bbb24 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.069599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.069599Z digest=sha256:ec5708998466b490ce35742513fe9558ce6a467418ade318a1fa34d2099968a5

Observation 1e4e2223-d431-4ae9-9af0-a9e31ef49758 · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.135870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.135870Z digest=sha256:65eb6cefd59ddace14ec8a83711b8b710426553c560b98faf15830b1c633bb7a

Observation 10a79d37-2910-49f3-aff6-d11fd8c59d4c · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.140228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.140228Z digest=sha256:c667655c745d13753d7b23f3a1e1cb99673830c6da679801f7b4c59b3ae95e1d

Observation 63a95e6a-6204-4928-98c8-5864781550c9 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Code Llama: Open Foundation Models for Code

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.144373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.144373Z digest=sha256:037e28f79b054d25743f466fda04ae77d7cb9f7868ed428f4843706d7d79a806

Observation 280b389d-db6e-44e7-bb5d-e48cbf25c1b3 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.148351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.148351Z digest=sha256:8b0a27ec28025aa2f27856c93c62619334b0ac6e60cb5d1fe3faf08bfd6af785

Observation 870b4498-931c-41cd-a6e7-abbc8839b71d · outbound

This paper cites https://lyz-code.github.io/autoimport/ autoimport.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics https://lyz-code.github.io/autoimport/ autoimport

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:38:28.666950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T10:38:28.151866Z digest=sha256:2c908c24df59cbec09753769ae2db3dcf8759e96f236cc5766491d4c3a974200

Observation 444747c2-aba4-48de-b594-2650b1891a9a · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.157210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.157210Z digest=sha256:c0743f9d0451ca9aab369a124f537cf621563f4b2bfa1465b5740b15fad964a7

Observation 0572d174-3895-4ae9-b9ed-00e10be25196 · outbound

This paper cites an unresolved cited work.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.161671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.161671Z digest=sha256:760123923b85bb89b21be0e31eee3be8ce8d06b2097dead8d3bcb091e00ee850

Observation 37de6bae-2ec7-4da3-805b-3b2633deb28f · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.165082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.165082Z digest=sha256:fcd804f3929d21f8a68520ce62b03e3de72ecf93be5377df337a87f7413d8277

Observation 386764f5-7076-40d2-9519-fa12012ec13d · outbound

This paper cites Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.255976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.255976Z digest=sha256:e018b614d117ca575e9898f9fa6f35d41c768f8e53113ae3e94f839d4e4fc2a5

Observation a7e1db0b-a0d7-4091-b74b-704b71752da3 · outbound

This paper cites CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code.

Evaluating Intermediate Reasoning of Code-Assisted Large Language Models for Mathematics CodeBERTScore: Evaluating Code Generation with Pretrained Models of Code

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T10:38:28.303524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:38:28.303524Z digest=sha256:b9fe477de0ae156f83c45b3a35512989e7dc99fee5575caf9697e7699b1f523e

Pith citing papers

No inbound Pith citation observations are available.