Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.519840Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2505.19706.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.519840Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:49.973078Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T14:53:06.955702Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5c6f086-183d-422b-a08b-8f31bb64ebc8 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d751b58-6292-4e64-b21a-6bb4a7c6ffe9 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5011179a-17b5-44e5-85fa-4f9aeb89f488 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2adfa8dd-79b8-4883-979a-77031df66ba4 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Let's Verify Step by Step
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9834023-d4e3-4163-b2fb-b9b35d000c63 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efd378be-ec6a-4293-9a0f-93cb4fb7e8f6 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8513cbaa-decc-4914-953b-2c653543a98a · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b66f5c06-a081-477e-87d8-c6ef3911f85e · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a86bc2bd-c4a8-4c07-8e90-39319c29e13c · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision R-PRM: Reasoning-Driven Process Reward Modeling
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58760b42-ef1c-4624-b54b-aa01c9a01cad · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56a5b2ec-712a-4545-8dd9-8ba54ba819fc · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c8303b-0f73-4d43-88c7-98bc4cfb0d3d · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision AURORA:Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse Verification
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b52878ae-1e85-4ec6-9057-86ea82c1a8c9 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Solving math word problems with process- and outcome-based feedback
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 879e5c96-25ab-4a97-b84f-7b061a93b65f · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89155cee-1c19-481b-823d-169683158bd1 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55b9c388-e4ae-40c2-8c1a-b682b88306f3 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c640f3d6-9084-40a9-a532-b9e2b78e8fd3 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0f8c5f-24aa-41c6-a3e4-cdccf5033517 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba321121-99fe-4196-8d6d-d1a4d1c5fce0 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 828d72b4-303a-42e1-bcc8-3f04d1b990c8 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9120d802-a4f3-4240-ba10-d949d515721a · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85eae2ab-d0f4-483b-9e3f-d46c1d24eeb0 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision The Lessons of Developing Process Reward Models in Mathematical Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0606978e-3802-4f8d-be98-70a2a53d2731 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54282e4b-965d-4c47-8922-efafabd17d13 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision ProcessBench: Identifying Process Errors in Mathematical Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b28158-9ba8-4576-bf51-3fbe86659bd2 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision online" 'onlinestring :=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2acf720-a52c-4464-9441-157725fd18b0 · outbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision write newline
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1b42452-0fdc-42ee-af57-006b3dd497e1 · inbound
AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc392757-c65a-47c2-ae0a-5adced636e00 · inbound
Process Rewards with Learned Reliability Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.