Pith. sign in

Paper Citation Record · LEDGER

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2505.19706.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19706 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.519840Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:49.973078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:53:06.955702Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5c6f086-183d-422b-a08b-8f31bb64ebc8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:18.968255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:18.968255Z digest=sha256:f3a6e491d1d1c19af1b487bf29f6bb4ef22676006491e6d13fd97e5cacd7d8ba

Observation 6d751b58-6292-4e64-b21a-6bb4a7c6ffe9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.156320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.156320Z digest=sha256:1360c92ed56bc24abe92d927665564feea5e182ed799c259c987afac7dbb068c

Observation 5011179a-17b5-44e5-85fa-4f9aeb89f488 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.237664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.237664Z digest=sha256:5c413fe0c7313a123b8cb5de05dd91800df424c9707c52c18a78ecd758ca45db

Observation 2adfa8dd-79b8-4883-979a-77031df66ba4 · outbound

This paper cites Let's Verify Step by Step.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Let's Verify Step by Step

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.319766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.319766Z digest=sha256:b5164da4a6eb58ea074acdfd75409a46b42839b99c6c802d6927d60286cfcb7a

Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · outbound

This paper cites Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.428465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.428465Z digest=sha256:c00a605e1c5e6fb36a5b758c162e869df80b352838f91f92edcf78daa2ead0b5

Observation efd378be-ec6a-4293-9a0f-93cb4fb7e8f6 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.554437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.554437Z digest=sha256:503d6cfb4429ef54b39fa623b3e615c4d62cd213e9e38f8d492d301683710526

Observation 8513cbaa-decc-4914-953b-2c653543a98a · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:22.011172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.674351Z digest=sha256:80ea8ef880e87758ab5bc5ba29fbcbd800bea6724a6ebfdb4b4a8224d558dd8d

Observation b66f5c06-a081-477e-87d8-c6ef3911f85e · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:14:21.835635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T14:14:19.783614Z digest=sha256:50f32a0e51dc808aef0cd7ef32d74f2718587028a27e1af31fd25dc99c7ee689

Observation a86bc2bd-c4a8-4c07-8e90-39319c29e13c · outbound

This paper cites R-PRM: Reasoning-Driven Process Reward Modeling.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision R-PRM: Reasoning-Driven Process Reward Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.876368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.876368Z digest=sha256:9e6fca4a2d08bcfbbc37606c5509ca819243fa18edcbfa4ce54ea841b81cdb0b

Observation 58760b42-ef1c-4624-b54b-aa01c9a01cad · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:19.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:19.989779Z digest=sha256:ebdd2107a390cd1bddaed48b6a34727fd5e4b34011d58d3305249b7e794b639b

Observation 56a5b2ec-712a-4545-8dd9-8ba54ba819fc · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.136880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.136880Z digest=sha256:5ece8aab534a172648c658b47fd26104671e282fb54cd4b53766d43a1874fc0f

Observation 88c8303b-0f73-4d43-88c7-98bc4cfb0d3d · outbound

This paper cites AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.254758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.254758Z digest=sha256:5721e7243b394f986aaa8b4a128c9f87d8ec6caba0240bd4f544832ba0c17f5c

Observation b52878ae-1e85-4ec6-9057-86ea82c1a8c9 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Solving math word problems with process- and outcome-based feedback

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.364504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.364504Z digest=sha256:9f01fb9b1b6a2f1ba110bcbe19643ccb4a11dcd3dd22367f3b89cd28ef3ef22c

Observation 879e5c96-25ab-4a97-b84f-7b061a93b65f · outbound

This paper cites OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.522619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.522619Z digest=sha256:b26b71a9a1d573d6124e683a4c88dfbef03bd77158ffc28ec3b75638299c44f7

Observation 89155cee-1c19-481b-823d-169683158bd1 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.610933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.610933Z digest=sha256:a16d9452694b0cc3973d64f49a1cae3ce0ecf7f8b89f0a9a0bdec7d6f3494bf3

Observation 55b9c388-e4ae-40c2-8c1a-b682b88306f3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.694287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.694287Z digest=sha256:8f508ceab4950ff54a01393a31218df820979a7ac8ffd6341f791c46c7b97746

Observation c640f3d6-9084-40a9-a532-b9e2b78e8fd3 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.773480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.773480Z digest=sha256:58f8fba58a9e936247cdfa87ac765ae72c11ffc3098de900094abea3744ff567

Observation cd0f8c5f-24aa-41c6-a3e4-cdccf5033517 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.861017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.861017Z digest=sha256:e42053660736b59d1f5da4bc22ec8a9aa2876f7081da97a5f45bfff8bd4d7bee

Observation ba321121-99fe-4196-8d6d-d1a4d1c5fce0 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:20.931464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:20.931464Z digest=sha256:bcf4a698897d6896dea16c402926c2df4d9174bbe86929a07f9506c6828f39f6

Observation 828d72b4-303a-42e1-bcc8-3f04d1b990c8 · outbound

This paper cites an unresolved cited work.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.023530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.023530Z digest=sha256:ced03c3d3c66371d31e5c87d6cc2685d6f8db34a07052594ad2a0db527444bac

Observation 9120d802-a4f3-4240-ba10-d949d515721a · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.092331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.092331Z digest=sha256:83be933e4d894b22f4b2b0fc88e89cc64c014665e9ffba27cc001ecae43661d2

Observation 85eae2ab-d0f4-483b-9e3f-d46c1d24eeb0 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.168455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.168455Z digest=sha256:0d05a654af07c4ee9950a0ebf206083f25b1b924dacd23bbbd30a28bdff8ab49

Observation 0606978e-3802-4f8d-be98-70a2a53d2731 · outbound

This paper cites GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.245461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.245461Z digest=sha256:6dd065430aea3595075849f2a5f55087712d13a09f66294a8027bf5d828ad4ef

Observation 54282e4b-965d-4c47-8922-efafabd17d13 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.314240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.314240Z digest=sha256:4b3e6a64fbd1155754ea218d924decd3f764a3f2a955d46b2e0f1abba446d81e

Observation c8b28158-9ba8-4576-bf51-3fbe86659bd2 · outbound

This paper cites online" 'onlinestring :=.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision online" 'onlinestring :=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.388599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.388599Z digest=sha256:cad45b2edf73c6be0d10bc46f466747639fefe5879bfb30ce67ba2c181f83944

Observation c2acf720-a52c-4464-9441-157725fd18b0 · outbound

This paper cites write newline.

Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision write newline

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.519840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.519840Z digest=sha256:cb4579674510bbbca76231cc5ce48aa39a8588cc70a5065b22589f29bd810506

Pith citing papers

Observation a1b42452-0fdc-42ee-af57-006b3dd497e1 · inbound

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models cites this paper.

AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:49.973078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:49.973078Z digest=sha256:1c5ba28bcf7d7a803ad7272bc003954a1b42276daf3d5c1e851f40f56bf59953

Observation bc392757-c65a-47c2-ae0a-5adced636e00 · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.957200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:24de49329cbcaebe98f9b9760e1437a4c40fc4e2c458d65fcd84328ef8ec9daa