Pith. sign in

Paper Citation Record · LEDGER

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning

As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2506.17533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17533 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:36.411682Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41ab05ea-3da2-47f7-a8c3-c53c66092af6 · outbound

This paper cites Language Models are Few-Shot Learners.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Language Models are Few-Shot Learners

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.483065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.483065Z digest=sha256:e77f7b72112867c8ceb8ac3275677b5a0c824448188b20c2ee8996ba0a431963

Observation d64b9b35-f1af-4bd3-9ce1-7878d1d93f5b · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.507985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.507985Z digest=sha256:0d462d4f2057839975955dfc2c5ae780a1c2f49411d976bf82dcf0d34eccedf7

Observation 81238a6e-2b0e-48b9-9012-9cd10c027c1c · outbound

This paper cites Step-level Value Preference Optimization for Mathematical Reasoning.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.544750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.544750Z digest=sha256:032d9f38086bb8d41afe5a7d05dfd50aa20b31d42dff562feb13d08d63e58c8c

Observation 4a52c34c-84db-4bf4-9755-9fa0e33a6d3d · outbound

This paper cites Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.567499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.567499Z digest=sha256:c039911c52b9eb1980a0fa3ade66ff5550953d6a8e6c46c56e8d778e389492f6

Observation 68d06ddb-0468-4029-9fe4-fc2d250ae7ef · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.601535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.601535Z digest=sha256:2caf046313229dfb82ba9e19c30d897dd41120fef9f92af4c1cfb57904da72a3

Observation 10739e12-f9ee-48a2-a056-224bb4de76e2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.639595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.639595Z digest=sha256:5b409cbf6940bbfada3c1d6e6171de4965998e6602fbe60c28b8b5f5e9baaa07

Observation 720a4782-eba4-43a7-ac84-3231ea75b7a2 · outbound

This paper cites DeepSeek-V3 Technical Report.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.678037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.678037Z digest=sha256:fd207e095ac7caa178a15a714692824944fa444a9855f016bdec15d7bb242a2c

Observation 1d7a5eaa-f1a8-42fe-bf64-020f12a41351 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.725780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.725780Z digest=sha256:c6e671379aadbb82d429236cbef77cbd26a3b7b1ebfbc751b99e6c1885b7a821

Observation 3a7b1593-0cd5-4afa-a5c5-c110cac517ef · outbound

This paper cites The Llama 3 Herd of Models.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.751872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.751872Z digest=sha256:1b977b4b0ad6eb56282dd5791dd556c12d5077b52d3db034494463dc04c8d93f

Observation 2656d7cb-e0c5-402f-bb80-5de2452c58f7 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.759763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.759763Z digest=sha256:af9d0bc9d17f5dbd68f25ebe2debf5fa124a0c4418dcc009838a10a28aa77330

Observation 03448270-1e72-4723-891b-a331e5063da9 · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.768822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.768822Z digest=sha256:b6d2db30c75f65488de90de8da3986386fdcd9182e56eb03d18209066d037f54

Observation 62a4bf07-7775-4597-8813-dd15972aca7f · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.780907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.780907Z digest=sha256:ef82314876d94f57cb2312300a9297b68eeb4f59f13b841545849be766583ccb

Observation a7d3a5ce-1b44-4037-8856-4a59efb214bb · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.784417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.784417Z digest=sha256:d4a9c61dfeabb41e321bbd7993acf4719df564cefe2f5967fc1a68d0517e9295

Observation 6ec01f6e-1f34-4817-96fa-5ed0a5f76acb · outbound

This paper cites Mistral 7B.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.789430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.789430Z digest=sha256:0506ae077eba15e3ccde63a630518bd2efb5bb1cfea6f4e4bd5aad8fb92c781e

Observation abef9125-2b0d-482e-9369-b428b999344d · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.793524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.793524Z digest=sha256:8225f2a9c86e2ceaec5aeb61bbe391ad009b11ef72a07ba160d03432b8a4512d

Observation 6e971a98-5922-4918-8cb6-525083fda904 · outbound

This paper cites Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:35:38.564837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:35.798164Z digest=sha256:b2d57965e271b136eac8f9b911f5931fda63d2a22df12147c90e30493cccb5b9

Observation eb671bb0-6aa8-41c0-a761-e18225df1bdf · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.801629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.801629Z digest=sha256:a4e5c38b2d46d4475dad583b8fb8115cd493e80da97a774334b4cd6df78f899c

Observation 36b2294f-7e56-478e-93bf-cf838495dc01 · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:35:39.615578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:35.807515Z digest=sha256:a3c45296b84cc594e688f91089122f926da3d3878a10b811217e51ac5f84dcaa

Observation de1130a1-afeb-4eb2-8811-9d0c36ee29fa · outbound

This paper cites GPT-4 Technical Report.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.816879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.816879Z digest=sha256:d143772bf968bc2a48811c038e8dc3706349b40ad9f49e9c150a0a1311bbbbfd

Observation 62cf131a-8fe0-483a-bded-1171765540df · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:35:39.487177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:35.834748Z digest=sha256:fc6a0426d780268a1e5560778ff96a115e18d1bbe0f58e60159cc23335ca083f

Observation 37d3195f-9ffc-4d6d-b76d-a2436a12b8b3 · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:35:39.414823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:35.888675Z digest=sha256:af8da298acae0a3c9f5edb74feb8351ac4c46e40b293aab4799e1dd97282efc3

Observation d7d2d4c6-4d02-4a06-8453-ac859996c402 · outbound

This paper cites Self-Reflection in LLM Agents: Effects on Problem-Solving Performance.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Self-Reflection in LLM Agents: Effects on Problem-Solving Performance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.944885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.944885Z digest=sha256:2dcb7d3c39f302c041484f3910b414a81f984fcef0b92e78c3d6117e949e0be7

Observation e1279313-417a-47ed-a1a8-01a313703505 · outbound

This paper cites RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-Fold

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.967705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.967705Z digest=sha256:890e2f0e2636cbf09f00813ab255e182320540a865944c06be21e99c56e21b40

Observation 988349bb-edc9-4dde-9e90-5b46c67ce724 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.995617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.995617Z digest=sha256:1f1af3dadba026178357707abda81b3a3631acc1cd456e125ac5df7b878a741d

Observation b1b765de-9231-43d1-80ee-ba10b373c0d2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.030468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.030468Z digest=sha256:30d7a26c4e0a0cfb97a2a540efb698fb25c7278c0ae6c9ad12fb29c63698334c

Observation b4958542-339b-4fa8-9173-945add6a594a · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Solving math word problems with process- and outcome-based feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.065794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.065794Z digest=sha256:93d556468f087646c58efd93867c3faf7ca0e063d60af61890d03364ed0e6041

Observation e82a1853-8285-4461-8166-fcdc7998f07b · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.079485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.079485Z digest=sha256:9230d119630024a01c309527bc61d4bea64c33e17543976a86d3bd67aecefee8

Observation 901ede36-cd7b-492e-a496-01f068123fc7 · outbound

This paper cites MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.084979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.084979Z digest=sha256:13f0d682e6974b11b7c85fb801c9d12b70f35c5fd61f958af0a1c336a1938016

Observation 461889e5-8629-4eeb-b096-1f79bdd524cb · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.090853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.090853Z digest=sha256:46971aa6db33a8c5be7cbb34e2f9663309097bc9f3520a1e23b8cc9a99504cdf

Observation 3c63e26d-c456-4990-a7b7-049a1d1ea470 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.097924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.097924Z digest=sha256:3ca034b6725b385aba97fcdf3b8386717342a3ade1e9fd1d405c408765d02939

Observation 6c555cf2-80ac-426f-a5b6-1c650347b6d3 · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:35:39.375548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:36.117846Z digest=sha256:5e9aa4d4bd550bc8f283da8a384654f6a1abc9b8650bcfbcf47216c34c2bb3a0

Observation 8ec35155-ec7d-4c13-ae8a-e6c9f64f860f · outbound

This paper cites Qwen2 Technical Report.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Qwen2 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.123399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.123399Z digest=sha256:6f17ecdd82538b5541dde4e5537d8d4ec45067586c2c28e2fa226ad121bbc9a9

Observation e2fe4bf6-d873-41f4-b517-db7e12f85681 · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T23:35:36.563521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:36.146704Z digest=sha256:80b32fdbdb3a0f69525db408a13e5b38447154ab14468ff8dbb1bd93338f1265

Observation 0edfeb99-8766-4a19-b24a-7c9e4350575e · outbound

This paper cites an unresolved cited work.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:35:39.286735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T23:35:36.167852Z digest=sha256:69d629136fc88d6f4accaa6c04fc47f05b257b99ca0a865c946f0ecf5c60e286

Observation 0ed12f59-2c2a-4be0-9baf-e34ad72ae4f7 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.214815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.214815Z digest=sha256:006b95542f3f38afc4079b614502f3457db23ac9851a68814e916344bfcf7b59

Observation 25598257-3867-4653-b43a-0808e163a3a8 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.257590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.257590Z digest=sha256:6d7b824ccd2becbb6b32c7df79eb60f7e1930d292521794ecb86447184575cfb

Observation bf3d2dc8-ec3f-4f87-a879-840526c9c941 · outbound

This paper cites Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.314813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.314813Z digest=sha256:aff48cd1f31ed2529a17595e211e471aff06bf04c1fdcf810b84003fb7924e87

Observation c310e405-2a45-4796-a435-0df9ae1b9660 · outbound

This paper cites URL: " 'urlintro :=.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning URL: " 'urlintro :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.364845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.364845Z digest=sha256:cc91eab33ee1ba0832e435dba5eaea4ef656fd9fa052216e3e59ac20d650fdd6

Observation 5430bc80-ac98-4ed9-a62e-5d5fc9d2bded · outbound

This paper cites write newline.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.411682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:36.411682Z digest=sha256:5362718aac25bf196ae70faa18fb89f624448fcc8c6e392eb34e3b62da645176

Pith citing papers

No inbound Pith citation observations are available.