Pith. sign in

Paper Citation Record · LEDGER

A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2510.08049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.08049 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:45.669787Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2d850bde-0282-4b18-a290-f3dc3a01aeff · inbound

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling cites this paper.

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:30:59.555716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T06:30:39.858246Z digest=sha256:9768569294fdf7ceea71716dd32a688d33b257133ef23fed2368d3223518b1f8

Observation 6a81ab97-54aa-4679-9cc2-b1894b2689f6 · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:33.949237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:33.949237Z digest=sha256:3413173b52ab9bf1bc492e671b22b78163f2646043976ebaa4cb65369b805ced

Observation f309b32f-ba66-4810-b73b-d98d9592556a · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:30:53.448267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:242d02e9ddc750616ef56346d18e060e4dd60217088d17cea2d0c4f502d8b0aa

Observation 5fc80b76-4701-4c90-9dd6-7e8aca71f89c · inbound

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering cites this paper.

Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 196

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:20:59.523271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T17:40:14.733882Z digest=sha256:97f9399d48bd7da6e32380df516e5ca91334dd253af4d8baf143030ff8897bc3

Observation 932b2a9a-9e3c-4721-8568-38eda0d226e2 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 126

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.516058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:bfdd89febdc2ed18b51a6d7655f02e7de54f36eabb24e81701224a5665b28796

Observation 5fdcce4e-9326-46b4-bab9-8b0875ca0526 · inbound

Process Supervision of Confidence Margin for Calibrated LLM Reasoning cites this paper.

Process Supervision of Confidence Margin for Calibrated LLM Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:41:12.237411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T08:19:09.437464Z digest=sha256:b4fffa4d7381b88373b6aa757a346b9bb11276063b6e52a0bfc4fee3beb96a16

Observation 434f1bb5-6ed0-4fca-b823-fca0e54f6391 · inbound

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models cites this paper.

LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T21:16:25.962737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T06:06:53.959061Z digest=sha256:f26d8aa3cb9e026bff96c53e111110f092a6553880bdc1cd3de1d942cc1a7c7f

Observation 1ed8a920-2437-4abc-a660-a8e17e2a1942 · inbound

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis cites this paper.

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-09T00:19:25.993336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:47:34.897401Z digest=sha256:c8422fed0cdcd5e2dd17ccfc4f23040b3bac40bdd34db23e8911b3f0e309a617

Observation d58d4ea6-070a-4d41-826e-141868c78f62 · inbound

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis cites this paper.

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:15:42.558536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T09:13:20.071265Z digest=sha256:4a7531a3ce935a8b3bee857f35396ce26ff274a207862cdfa9edc0cc1952bd5f

Observation 6d71e468-2da3-4ea2-9322-7474b7952a3b · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:16.840951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:9cf5a6b16cc859c2d834f79ea90170353a6cef418e059ac8aab3158d678a8be1

Observation 59e34ba7-0921-45af-bf08-f04947d5aacd · inbound

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering cites this paper.

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:55:43.621498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T17:57:18.130401Z digest=sha256:003da4c158ebb9436f63ced19bb7b6d7620d5326b4416c0949afc52a8a85d246

Observation 4b36b1c9-172a-4812-8301-95552cf2bfd8 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:43:38.769945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:7c9420ca893b90a0fe6893815ca8eafe8475294a3027a63825c94854a973108a

Observation ccb30e36-f67b-4ef9-9dcf-5c0cc37f6ded · inbound

Learning from Saturated Data: Signals Beyond Correctness for LLM Training cites this paper.

Learning from Saturated Data: Signals Beyond Correctness for LLM Training A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:12:24.561892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T17:10:05.592596Z digest=sha256:9eb44494b445073b1a2fa83187d06f6c86bf98a372b70080b28af127b8203c49

Observation ea62d952-3925-4f51-9dc8-d2fe78c4813d · inbound

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs cites this paper.

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:18:58.377454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T00:48:56.892634Z digest=sha256:649ad321a9ef59f90529f506a8b8debad7b3643b4c34dfcdd8ea3850a628a8dd

Observation 49477570-1edd-4303-b655-6eca31098101 · inbound

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments cites this paper.

MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:39:31.181688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:43:03.092837Z digest=sha256:15b01ff5a8a0cdb3cc561943bd8a5a38893783c26db0506dc661eb5c318a289b

Observation 4945d923-54f8-4128-837e-3e623c309379 · inbound

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling cites this paper.

KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T05:02:55.608400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:02:55.608400Z digest=sha256:992fb532c4cfaea95de180035d6b431f3e175656e98a19787e5f8b65b913b5ed

Observation 69719b0b-ac22-4bba-9b34-be23546ea564 · inbound

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems cites this paper.

Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T15:43:24.809948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:43:24.809948Z digest=sha256:b074d3bbc07dc4122763ec98a98d3bfe6348a56d44cb5b833573723e94d5bc9c

Observation e4aefbf9-c8c4-4576-afdb-3f49d95be509 · inbound

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning cites this paper.

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T13:31:02.965705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T13:31:02.965705Z digest=sha256:dfb99b1f35eb102b239ae308bb045a5de3ee1e9bec9692f4ef53c95aa36a37aa

Observation c8ce994c-2139-4f6a-a705-8a1eebf1e7ce · inbound

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability cites this paper.

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:55:02.896387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:55:02.896387Z digest=sha256:03ba9bc45f1c0fee201e2a274cc1f8a72d484f73d557d927c81ed1a9ff21da71

Observation f8900606-bcef-415b-ab27-00dbdcdac4e5 · inbound

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability cites this paper.

NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T04:24:56.708873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:24:56.708873Z digest=sha256:f59bde63f43cf7bcaa84529dedc2247782fbf117ef2d567ca6f0c99f88aad2e0

Observation 51c4fc5c-8ba0-4182-8d2b-6574abb25afd · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning A Survey of Process Reward Models: From Outcome Signals to Process Supervisions for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:45.669787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:45.669787Z digest=sha256:8f056e4009dd595104f3a3736a98d6b5fd7b94b8f17f3cd5e27c092c5ce54021