Pith. sign in

Paper Citation Record · LEDGER

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 5 inbound Pith citation observations for arXiv:2505.14391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14391 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:38:50.332584Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:32:55.155806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:47:28.400784Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 56ea353b-2cf6-437e-a849-016d3e2fb7b9 · outbound

This paper cites online" 'onlinestring :=.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.491408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.491408Z digest=sha256:8560d6148f3e89165f6352c58c7ab9495347c2c5f432988920bfb3d283026822

Observation 5ba3d376-3125-4cb4-bef0-b951c523e66b · outbound

This paper cites write newline.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.646682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.646682Z digest=sha256:c60bf0bb8fe79f6a253122e01a67dfb55f370574694511302c4015934a2d3b5f

Observation 99ca64c4-e157-473c-9c38-c563923aad0b · outbound

This paper cites GPT-4 Technical Report.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.810278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.810278Z digest=sha256:5fb3d446d7acdcc8d66cabb9e538fd8b556e45dab8df19eb29fcd023691e9517

Observation e6f6a7e9-f55a-4881-9803-c1929b3d812e · outbound

This paper cites AlphaMath Almost Zero: Process Supervision without Process.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning AlphaMath Almost Zero: Process Supervision without Process

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:45.934365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:45.934365Z digest=sha256:d5822af5a74a495b97cb57556f1824b2e08f1eb562d879f1f549cf2c80354f5f

Observation 2913f427-467c-4a51-9ead-0fd50fbf8d14 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.053508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.053508Z digest=sha256:3e4c05956f3d54666293f1df9e6f905def20766202698f8f6edab73f44f0d609

Observation 60664621-518a-454d-9e7e-1498e3ce7270 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.198208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.198208Z digest=sha256:7b78783fcc89c5016186bf7f1eb0a01ea8d44d4ecdaa26fff4c88cba2163e4db

Observation a605f1c9-900b-4e70-878d-96b1991697b1 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.337373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.337373Z digest=sha256:8bba9c17b2dbce6063b929d33934142c9297a2241372eae8280aa865cbd7830e

Observation 06c8340c-dd92-4d0b-98e8-fbecc9948eeb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.503915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.503915Z digest=sha256:87bc60fddd3b6e0a2466dfbb2165af06f49598da7d94d4429a85843d7f00cc20

Observation c1bbf3ff-8196-4307-b5f3-6b2c0075e88a · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.642136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.642136Z digest=sha256:b44c1a626560a5737792f2b8e489fae87ef6ca383bcfb009a2d1cf398971504f

Observation 16e8b4a0-21bc-41e3-a6cc-843eb4a2c98c · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.750931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.750931Z digest=sha256:5221022d24ae1a9b390ce9792065576652646011aa1dff204dfc9cfa4abb579f

Observation 1dbca456-136a-40dc-9da3-d3c8609eb920 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:46.889283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:46.889283Z digest=sha256:fe6a314d752d1a44f877815fe5be25afdce9c04ca7e0f26e54ce88c85c863ac1

Observation 6a2055b3-5ad8-4212-9f18-dd114fb85d6a · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.040946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.040946Z digest=sha256:dd4901e2f98403e75d85539404557bc1da374aecdd08c98d477859310d50135d

Observation 4f78a47e-f84a-49b0-a976-e0dae7057545 · outbound

This paper cites Let's Verify Step by Step.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.152880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.152880Z digest=sha256:0084c7491e0fd079add0bfb0c544cc4acf282eeec77b0ab90e6f41d76f5573aa

Observation eb3f3a78-6ed7-4ae3-8649-6891effedeb9 · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Augmenting Math Word Problems via Iterative Question Composing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.312631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.312631Z digest=sha256:eb5e187c6e9eed29e820e68973fe7d43b6588f87706ffec7cd0f4c9d27d28cf9

Observation 8a944585-18f4-4afa-b251-22318ae3ad1c · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.463728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.463728Z digest=sha256:f5df5ab38d279b7f14ff4893fd0f442f3ab581645d6b25ffe08811709648ea36

Observation 41a83954-97b4-4a1c-b509-692ce2a2caf9 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.611667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.611667Z digest=sha256:4b09c8a25f31cf7cc7af0bf3b01f3be6f17bfc5e745693fe4bbfb794a75e52bc

Observation f06f3f27-1078-4c93-ab80-8998c3a269fd · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.846806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:38:47.742752Z digest=sha256:855c028406e2f985320e2ccbb02aaa1e66187cc9466d44528ff418874bab85ee

Observation c24a8b09-f2c4-4ca0-bcaf-e6991b214e00 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.858019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.858019Z digest=sha256:db2af70c2e4ea0d316da3e95de51bbd52f3d6fdbdb20848d3cb5d8bf90ad5698

Observation 5b37e09c-c8cb-43dd-b03e-b0f7df8d9fea · outbound

This paper cites s1: Simple test-time scaling.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning s1: Simple test-time scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:47.943305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:47.943305Z digest=sha256:a346a10006581974c01dc71359908984f4de1074ef8e2c74641426b69d3a8080

Observation 3acdf7a3-f627-4f9a-9956-ffd8bdd99dbb · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.144397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.144397Z digest=sha256:f6402a837e5937013fd82531c5f028ebc338b24c5f164b0ec0c64204bd322dd5

Observation c5437f33-0c61-4f58-ad92-432aa0b5a7c4 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.712051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:38:48.277073Z digest=sha256:592908e03beb596b00deeb60e52593db6de2c770c15ae366046d767d96f71a4e

Observation 326fe5bc-3c46-494f-8c9c-400f7d0fc181 · outbound

This paper cites Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.422552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.422552Z digest=sha256:425f163201a74e362511206737ce94bddd617ef961f12d27d1c4c755800cb10f

Observation e003c48c-8e24-486b-a8ce-5a086ee337bd · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.565575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.565575Z digest=sha256:2544132eb8aee52cdd83546c8baf9f90f96620f875f27f2f3bfd7dc4399a1bc6

Observation 3a7b88cd-be1a-4b3a-bf1a-3897f50274cc · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.738925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.738925Z digest=sha256:de4a9cd8133a3e7b86a0dd9b3a0f0537a24883070af23e8c72eb368ed8aee9ca

Observation fc93cedd-75b4-4182-8da9-3dd5f9219ee9 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:48.946473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:48.946473Z digest=sha256:a4365cee64b6579d4a78bd5bd933110755ac568ea7a7e2db53e8f2f244467ce2

Observation ca3e96b9-940b-4df8-8af2-26d93ae340af · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.448514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.158853Z digest=sha256:9208c8afd0e2d1f1249bd958b385ab6fea50fb98c1779b0304b761ffbc50fbb0

Observation a30162a0-c2bd-4ab1-9c0c-bbce5076f09d · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:51.195028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.306606Z digest=sha256:5bd0c920f0df03a5c43c949128cd91b6204ff45aef856670a533c46383300776

Observation 5eab2423-d5b3-4a0d-be5b-c291cdbcc078 · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.425108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.425108Z digest=sha256:c50af82e7eb05e7d278b8137373fe57a46adfc1ff5d0c30e44ebc926ebdca592

Observation abf49cd1-0d55-431d-8f3f-9634012cd3ee · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.624831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.624831Z digest=sha256:88de235d46713c86327a5528ff70f2d0a9f04bbd8bbe3ffd2effc613ff70433a

Observation fbafd562-03d0-411e-8ae2-bf6c467f52eb · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Solving math word problems with process- and outcome-based feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.778223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.778223Z digest=sha256:5b12eabdeb005c0aa55f70c45414ea721150cc49caf2a169ea136f3c2b43020d

Observation 64034c53-e4af-46bc-87c3-a756aa20b3df · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:38:50.970290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:38:49.897802Z digest=sha256:74f593ebf98a8371305c5e78912a1cd0ba7a6acaa05e24abe1a9fa3ac49fc539

Observation c9c70a28-981a-443c-b618-94a132a8e05b · outbound

This paper cites Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:49.957959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:49.957959Z digest=sha256:44e133ecd8f52395bbf0a782c0e13b405b35457b38af32adfbd12ecfe532aec8

Observation 4d06186d-a39e-4144-ba60-7034a482949e · outbound

This paper cites an unresolved cited work.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.036904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.036904Z digest=sha256:d657786f636cf45aa0d90cd47064f63a342c5e0a638226672daaca975d93a6d7

Observation 5ab26c4c-3faf-40a7-a2b6-f62068ae0116 · outbound

This paper cites ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.099971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.099971Z digest=sha256:72b918157601099856d0f89af90608c5a0ced8a7565b967a001028cdc8759f62

Observation 53065805-6222-4cb9-827c-c369526c799f · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.179847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.179847Z digest=sha256:b37cce40fb6503a9cc5bccc4a3a980f8484c56a2974b7e614e03d68a8237abdf

Observation 630a15a5-4fd0-4948-9489-2af1c8d60d94 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning Automatic Chain of Thought Prompting in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.251756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.251756Z digest=sha256:c957713ef14789a2ee4d339e7f969a2633e0f95a47d370ae5a7594f01a731ecc

Observation 32fdfd0e-e858-4043-823d-a09a5e0bef91 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:50.332584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:38:50.332584Z digest=sha256:5afd676f5e1b805edac1dae3125a770b15956a790290f9217083345dd4658337

Pith citing papers

Observation 4bdfda50-1c54-4dda-9721-25b4f9f42947 · inbound

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning cites this paper.

Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:55.155806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:32:55.155806Z digest=sha256:90569c67e2d318df0f35bb374a98735331a162b55c9a1c7c913dacf82cf6a6e7

Observation 7e729a3c-92e4-4d7c-a159-81f0384848bf · inbound

LLM Reasoning with Process Rewards for Outcome-Guided Steps cites this paper.

LLM Reasoning with Process Rewards for Outcome-Guided Steps Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:47:28.403595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T06:43:19.869111Z digest=sha256:f5171ab69e0dc5f6167775527e70bcb106bff466c508c1fb35ef26957d3703d5

Observation fe7a82cc-eca3-4c9f-bacd-f3766a158e7b · inbound

Improving Medical VQA through Trajectory-Aware Process Supervision cites this paper.

Improving Medical VQA through Trajectory-Aware Process Supervision Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:06:06.051115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:18:17.846140Z digest=sha256:31fd9550786e126ea50e3024f82f7daa06abb2f2bca1ffc8502308f84a462069

Observation 16ffdba1-42f0-4c17-8c90-bd1ee5a4ba35 · inbound

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning cites this paper.

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:26:27.809955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T06:13:09.898530Z digest=sha256:162f507ec684c0d702b519afa28737f94fcc63fc48c7f3340b4a376034eca2db

Observation e49ff7ad-a288-4b99-92e4-8b31bf123d1c · inbound

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning cites this paper.

PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:22:50.477238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T19:20:32.435135Z digest=sha256:6da87c73edf6eaf8b1a9be5f15f03e1e49e154c3556dda36d7728293837181dc