Pith. sign in

Paper Citation Record · LEDGER

FreePRM: Training Process Reward Models Without Ground Truth Process Labels

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 4 inbound Pith citation observations for arXiv:2506.03570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03570 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:42.841133Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:57:32.948217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.479928Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b692a8df-3cdc-4c89-be02-b60657262463 · outbound

This paper cites Alphamath almost zero: Process super- vision without process.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Alphamath almost zero: Process super- vision without process

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.644217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:40.848523Z digest=sha256:a1aa4d6b71d60d60dd682efac1ed643d6ad3280da1be13b07aee000e9ca812ca

Observation 1f6013f9-299f-4d41-8515-9aba3e27ce84 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.892903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.892903Z digest=sha256:bd95580e3ac74188809d8d93ad895986c34258be361b7d94f7d246c9f9d91ac6

Observation 409cec1e-f603-4b34-9b8c-778ef84edb9a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.984203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.984203Z digest=sha256:03bdf4885402b0575a511b8d94f263e07bfd1a337b711ba5310102cf130aa14b

Observation ce41dbb2-d935-40d8-917c-c8f0874ef52c · outbound

This paper cites The Llama 3 Herd of Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.062201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.062201Z digest=sha256:de5fdb4cc091c26a2f2912d29a762fe8189cfa5fe4bb6b7eda94504fceb0f6ac

Observation 60386761-00f8-4d14-9105-f477aa201bba · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.126923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.126923Z digest=sha256:064d5443c34312ac4cb8acbc3171ddb1316889f27f420679877a835c5bd22aa4

Observation ae69aec8-0373-439b-ab6c-ce25584e52db · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Measuring mathematical problem solving with the MATH dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.562843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:41.206238Z digest=sha256:d446c4a9dd3873e1308345e52e278cb4822aa9e975a838b05c18f31ec9b390d0

Observation 07e9f2ba-67cb-4956-8cec-c8d6ef4c59e1 · outbound

This paper cites Qwen2.5-Coder Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Coder Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.250937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.250937Z digest=sha256:91ad96f8ebf4ff3d46f792022d647d9269ef25fd54ac37c2f5cbcb905a2129f8

Observation 5d6ddf36-bc9b-479f-b39c-69e0daab2ff2 · outbound

This paper cites Mistral 7B.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.343356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.343356Z digest=sha256:761b9f8f9c176a9f898e624370f9e88043e23908eb6f79f82f92710abed8d13c

Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · outbound

This paper cites Process Reward Model with Q-Value Rankings.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.397499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.397499Z digest=sha256:7b44faefcd5e7393247775594cf4706fe59a98911eecd852766aef002de9531e

Observation 41423e53-92bb-439f-b9d5-05f4325a7b4b · outbound

This paper cites Let’s verify step by step.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Let’s verify step by step

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.434667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:41.502360Z digest=sha256:77c2b133f86e010a54c21c7a2462afa2f5c35e175a490e42df33b3fb61d131fe

Observation 768ef0a6-eb07-4df4-ba05-4b74f7a63012 · outbound

This paper cites Autopsv: Automated process-supervised verifier.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Autopsv: Automated process-supervised verifier

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.321785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:41.560493Z digest=sha256:2f4cfc97f305e44ea63a130c24fca1e768619628bf14a47daed5046f83fbd59e

Observation 71f147f5-9355-44db-9ffd-cf487da4571e · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.654194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.654194Z digest=sha256:b031deb5d3342b02cd3bea01a8e7fd1736d2fe3ae01fd3ce8f7f114729df3a01

Observation 2107b35b-0ad3-45ea-b7eb-01f31268c4ff · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.718662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.718662Z digest=sha256:3f783264d747ad5c2d157154edc8aa6a5c964469d5432e45bb82b7fcfdfbf3ef

Observation c138dd29-192a-467a-aada-eb6e5c7cabaa · outbound

This paper cites Skywork-o1 open series.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Skywork-o1 open series

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.213012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:41.762843Z digest=sha256:55851b01f0fc155df261a2bc84610bbacb754031f83c0d0d360b778d26ddecf6

Observation b94bc281-9988-4294-afbb-d713b7ea1435 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.911560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.911560Z digest=sha256:a91bf2309f999bca8d09adc9cdd430bddf04df93b771b44acdf0e2cf2de4db76

Observation e88e6dc7-a111-4a8e-8439-e4c2192d5cd3 · outbound

This paper cites Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.008548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.008548Z digest=sha256:63340e5e25cdeaf967ef09ed598ce4fa62b3c4cdd5d57ca14a222c2108733728

Observation 67cac0ea-15ea-4643-b7c7-f04d64faecf2 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.963514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.057605Z digest=sha256:6d463aa7a2ae8f9353bf9a87f23475ca88d413a8636254f3bb60ebfc363799fa

Observation 9a29aa4f-7abc-4e3f-89ef-5b847debc3a2 · outbound

This paper cites Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.838113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.140691Z digest=sha256:3d3be234c578e5bd09baf6c2564d69e68dad680745be025d9b5f3db906db805b

Observation 67831e97-c489-455b-9ff2-72921486e6ab · outbound

This paper cites Training large language models for reasoning through reverse curriculum reinforcement learning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training large language models for reasoning through reverse curriculum reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.758576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.205649Z digest=sha256:903e296bf29a426d2977aa23b7868d11624d9adb4e3190e32cd013a0aeeb3875

Observation e37d8d9a-6eea-411f-a8c9-0c9fc9a8503f · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Evaluating mathematical reasoning beyond accuracy

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.633494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.272829Z digest=sha256:e115894f6b90703536a285e9f6b37cbc2ee8ad20b8f7fe3d622077ee5e7cf1c2

Observation f102bf8e-4392-41a8-91f9-5159c003ec7f · outbound

This paper cites An implementation of generative prm.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels An implementation of generative prm

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.509443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.339628Z digest=sha256:2317f023f55fa9f3b74de2c1edc59d601fe1003719f27a753df420ebb7a44f5b

Observation cbc7beb3-2e64-4474-9d5b-61baa3ed4d88 · outbound

This paper cites Qwen2.5 Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.395649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.395649Z digest=sha256:eba08e82adb31670dac13030a0cca8bfab414acd3984bbbed4bc081240ce92c8

Observation a586e8ac-6143-45c3-a3dc-3f523d1d4f41 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.449307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.449307Z digest=sha256:288909e34ffbb49188e716f673f15083cba1a99f44b3fd2eb1cdb814c6e1bcab

Observation 64cba8e6-cf3c-4a15-a310-e833bd8a7636 · outbound

This paper cites Ovm, outcome-supervised value models for planning in mathematical reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Ovm, outcome-supervised value models for planning in mathematical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.348491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.539085Z digest=sha256:a7315b9ee852e59ae8101baf1c4c41b05eb3a84dc2665bf08f4694cd4fb02c98

Observation aef94d07-53cf-45ff-8dfe-8b2e80688aae · outbound

This paper cites Free Process Rewards without Process Labels.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Free Process Rewards without Process Labels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.575866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.575866Z digest=sha256:cc70ffecf555c60dadd2b37eeed9c3a60492d3e41d72655eebf0a9cc7c43defe

Observation 944cfdcc-93f7-47b8-82ea-1109dd1f95d8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.642734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.642734Z digest=sha256:53adc2637246a20a9a60c38d2ee733b634477a66bdb185417c2ce4b4ffd1b93b

Observation 814b695d-53dc-4b7b-b30c-487d35b11c74 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.722080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.722080Z digest=sha256:4e985748f5e6333ceb8272eec3dd3f8a5a8e5e0e68b924d796a9ad5c8fa04730

Observation 63478cf3-b4f8-4893-85c5-e0d097331ed0 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:06:42.769583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.769583Z digest=sha256:90371f276edebd5095b215d44a6c4eb23cf2acd15fe83bb99d015a8893237977

Observation 62e255ff-a0d4-48b2-8583-4ac3d54511ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.161376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:42.841133Z digest=sha256:5ddae965ba12bcecd32dac13c4d8579b818ef87d595f1d304ce3d1ca7c7c456d

Observation 8e862179-7b9f-476a-ab48-33c92b429597 · outbound

This paper cites an unresolved cited work.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.086411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:06:41.854876Z digest=sha256:ee08345d2282700fa82f04dbebffeb324c97a9a13196569e90d2efdf592f42e5

Pith citing papers

Observation 355ce1e2-f6ca-4d58-983a-4b788d6b6edb · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:32.948217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:32.948217Z digest=sha256:1c36fda67c1f12ac90956531aeb62639877b00eeb6907f933c7b4f9cb20604dd

Observation d1a55f1a-dc7c-49ae-9cea-5bfb614c2268 · inbound

Unsupervised Process Reward Models cites this paper.

Unsupervised Process Reward Models FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:18.892131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:19:04.275069Z digest=sha256:a328d22009fda93f98821e895b95e5c4a87bf4929bcc1477fd966479b4540a71

Observation 2dc7b6a4-248e-44e1-b692-e4fa63a8c41b · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.933322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:0714b10a093108df68d54c6d63a30b7bd5f325ab49b42931375f6c8f762cf4fc

Observation 32756f27-4328-4fa0-ac5b-0dd79cc3f4d8 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.481368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:246e4bd8661ffbbb2939b2958d3af661d800a84f1dc35f2c560f0db4b3a18680