Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2510.08049.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:45.669787Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 2d850bde-0282-4b18-a290-f3dc3a01aeff · inbound
ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a81ab97-54aa-4679-9cc2-b1894b2689f6 · inbound
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f309b32f-ba66-4810-b73b-d98d9592556a · inbound
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fc80b76-4701-4c90-9dd6-7e8aca71f89c · inbound
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 196
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 932b2a9a-9e3c-4721-8568-38eda0d226e2 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fdcce4e-9326-46b4-bab9-8b0875ca0526 · inbound
Process Supervision of Confidence Margin for Calibrated LLM Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 434f1bb5-6ed0-4fca-b823-fca0e54f6391 · inbound
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ed8a920-2437-4abc-a660-a8e17e2a1942 · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d58d4ea6-070a-4d41-826e-141868c78f62 · inbound
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6d71e468-2da3-4ea2-9322-7474b7952a3b · inbound
Improving Vision-language Models with Perception-centric Process Reward Models A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 59e34ba7-0921-45af-bf08-f04947d5aacd · inbound
SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b36b1c9-172a-4812-8301-95552cf2bfd8 · inbound
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ccb30e36-f67b-4ef9-9dcf-5c0cc37f6ded · inbound
Learning from Saturated Data: Signals Beyond Correctness for LLM Training A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea62d952-3925-4f51-9dc8-d2fe78c4813d · inbound
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 49477570-1edd-4303-b655-6eca31098101 · inbound
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4945d923-54f8-4128-837e-3e623c309379 · inbound
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69719b0b-ac22-4bba-9b34-be23546ea564 · inbound
Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4aefbf9-c8c4-4576-afdb-3f49d95be509 · inbound
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ce994c-2139-4f6a-a705-8a1eebf1e7ce · inbound
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8900606-bcef-415b-ab27-00dbdcdac4e5 · inbound
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c4fc5c-8ba0-4182-8d2b-6574abb25afd · inbound
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.