Pith. sign in

Paper Citation Record · LEDGER

R-PRM: Reasoning-Driven Process Reward Modeling

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2503.21295.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21295 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:43:42.331901Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4bcda4fe-1756-49ff-af56-4157e78c42d3 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models R-PRM: Reasoning-Driven Process Reward Modeling

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.184410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:1bea99f118e5e132ef769345bc9263378895f9a265960c4988f5604ae5a3039f

Observation a393c6a9-5bd3-4a51-8cf1-56aecfb7a3fa · inbound

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling cites this paper.

ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling R-PRM: Reasoning-Driven Process Reward Modeling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:30:59.606855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T06:30:39.858246Z digest=sha256:e87e9237097985dfb3adafc11d443e289b180ab8b05a39923c82a5a8a0d1d108

Observation 2a30bd6c-6d1a-4085-ae26-802f37b13c27 · inbound

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs cites this paper.

HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs R-PRM: Reasoning-Driven Process Reward Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:43:42.331901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:43:42.331901Z digest=sha256:7d32427831ab182f028f3c805e261a437eed51f4767663ad3b420b53fb87ba6c

Observation 4fd6eea1-b460-4f99-8258-c69097cc812c · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation R-PRM: Reasoning-Driven Process Reward Modeling

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.303909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.303909Z digest=sha256:d2803e982789fa56ba499eb6be85bbb61cd4b7ef93cb244fbf96394cfb6d1439

Observation b4ee02a4-07ff-42d9-81b8-cb80e3007941 · inbound

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing cites this paper.

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing R-PRM: Reasoning-Driven Process Reward Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:08:26.234113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T08:24:43.494005Z digest=sha256:788c199954f52f673dce1f43ff9bf9385cb123f7bac7712775203101f01af190

Observation 8576d876-f414-40ab-ae32-bc9bbf4c0dd7 · inbound

Pause or Fabricate? Training Language Models for Grounded Reasoning cites this paper.

Pause or Fabricate? Training Language Models for Grounded Reasoning R-PRM: Reasoning-Driven Process Reward Modeling

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:03:36.662326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T03:01:58.366028Z digest=sha256:046157f13a7bde0bb96755d7bfed34beb7c27be4e4775ddc6c07342c0b5ff4d6

Observation 63e4b679-207e-4c0d-a4bd-32fea6b1e59b · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding R-PRM: Reasoning-Driven Process Reward Modeling

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.772927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:a82c43ec53728d1e6ea4ae266e2e010ebcf8e66dd35e5684df6b8e88d336dc42

Observation 24989b75-f624-45b4-a5ac-7d75c5c20002 · inbound

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction cites this paper.

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction R-PRM: Reasoning-Driven Process Reward Modeling

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:03.356407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T09:44:38.871250Z digest=sha256:3038d5564f12e6110dfe7bf955ad04a9dc7f65ca369c8375a5445adf59238e73

Observation 8a8debba-b69b-46d0-8a60-8c98e2030c48 · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs R-PRM: Reasoning-Driven Process Reward Modeling

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.863366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:822aa4022a260a67602f62344146271cde8192b473cb4e02f5fe926db6b5db80