Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2504.00891.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:44:05.432691Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 76ff2eb6-6d36-49cd-8018-38ebf2b9cdc7 · inbound
CEC-Zero: Chinese Error Correction Solution Based on LLM GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0606978e-3802-4f8d-be98-70a2a53d2731 · inbound
Error Typing for Smarter Rewards: Improving Process Reward Models with Error-Aware Hierarchical Supervision GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a05919-d3b7-4492-afd8-954af94a293b · inbound
RewardAnything: Generalizable Principle-Following Reward Models GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7891e27b-3d91-4bbf-8f8b-c79929b233b8 · inbound
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0c21f7-4a31-4d49-aa00-3063fb45931f · inbound
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 498826c1-e615-4c11-9bb6-b9d0c2b1e64c · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d9c29762-dc39-4dfa-94d2-f5d15191f27b · inbound
VRPRM: Process Reward Modeling via Visual Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 09feaf30-5189-45e0-ba59-9a6d8c2827e0 · inbound
VRPRM: Process Reward Modeling via Visual Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21956489-8630-4de6-be2c-d86c29306425 · inbound
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b17e8c60-dd13-46b4-9bf3-6a46a0cd282b · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ff55b3-f345-4883-b400-6c450a78b296 · inbound
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4cfd4f57-14b9-4ecf-b437-59dc73f21466 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 248
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cfcb354-9fa0-46ac-8c05-00e28e555a11 · inbound
Rethinking Reward Models for Multi-Domain Test-Time Scaling GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fa7736a-1c50-4957-948d-cc2c222b8e8f · inbound
OpenClaw-RL: Train Any Agent Simply by Talking GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5071457f-3064-4b6d-9403-5c11d55f4eb2 · inbound
Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 150769c8-fde7-4410-96b5-e1bbb638392f · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 55fb33bd-8353-4aa3-9d71-e11beff6a47a · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 902af584-c4ad-43c0-bd49-f7955164ed3a · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 180
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 21a0c66c-7784-4175-989a-a8e94a6ee721 · inbound
Unsupervised Process Reward Models GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88537622-2749-48fa-9e25-640d33ac42cf · inbound
Not only where, But when: Temporal Scheduling for RLVR GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fa6bd11-97f4-4a59-9fd0-a64de946cd65 · inbound
Trust Region On-Policy Distillation GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 196
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 554fc23f-eb35-46a2-b280-ccf2419f76d4 · inbound
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fc0c13be-1a6b-4d1a-afd0-91d8030837c1 · inbound
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 275c856d-35b4-46b1-83da-a774dafbc60f · inbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890c36e8-81db-45f0-b5d5-67bbcd0237eb · inbound
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4066ecfe-4bb9-4e54-a632-564d9fc718a5 · inbound
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.