Pith. sign in

Paper Citation Record · LEDGER

FreePRM: Training Process Reward Models Without Ground Truth Process Labels

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 4 inbound Pith citation observations for arXiv:2506.03570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03570 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:42.841133Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:57:32.948217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.479928Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b692a8df-3cdc-4c89-be02-b60657262463 · outbound

This paper cites Alphamath almost zero: Process super- vision without process.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Alphamath almost zero: Process super- vision without process

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.644217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:40.848523Z digest=sha256:100ca4f0638131a4981d5feb04d88e499e3200655afe48b3f9838db7bb1aecf5

Observation 1f6013f9-299f-4d41-8515-9aba3e27ce84 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training Verifiers to Solve Math Word Problems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.892903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.892903Z digest=sha256:bd95580e3ac74188809d8d93ad895986c34258be361b7d94f7d246c9f9d91ac6

Observation 409cec1e-f603-4b34-9b8c-778ef84edb9a · outbound

This paper cites Process Reinforcement through Implicit Rewards.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reinforcement through Implicit Rewards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:40.984203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:40.984203Z digest=sha256:03bdf4885402b0575a511b8d94f263e07bfd1a337b711ba5310102cf130aa14b

Observation ce41dbb2-d935-40d8-917c-c8f0874ef52c · outbound

This paper cites The Llama 3 Herd of Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.062201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.062201Z digest=sha256:de5fdb4cc091c26a2f2912d29a762fe8189cfa5fe4bb6b7eda94504fceb0f6ac

Observation 60386761-00f8-4d14-9105-f477aa201bba · outbound

This paper cites LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.126923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.126923Z digest=sha256:064d5443c34312ac4cb8acbc3171ddb1316889f27f420679877a835c5bd22aa4

Observation ae69aec8-0373-439b-ab6c-ce25584e52db · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Measuring mathematical problem solving with the MATH dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.562843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:41.206238Z digest=sha256:974a3c05c14c75e7d3c9b66cadadcd506000708d23f72e1c9fc8008ba6ab1999

Observation 07e9f2ba-67cb-4956-8cec-c8d6ef4c59e1 · outbound

This paper cites Qwen2.5-Coder Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Coder Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.250937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.250937Z digest=sha256:91ad96f8ebf4ff3d46f792022d647d9269ef25fd54ac37c2f5cbcb905a2129f8

Observation 5d6ddf36-bc9b-479f-b39c-69e0daab2ff2 · outbound

This paper cites Mistral 7B.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Mistral 7B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.343356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.343356Z digest=sha256:761b9f8f9c176a9f898e624370f9e88043e23908eb6f79f82f92710abed8d13c

Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · outbound

This paper cites Process Reward Model with Q-Value Rankings.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.397499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.397499Z digest=sha256:7b44faefcd5e7393247775594cf4706fe59a98911eecd852766aef002de9531e

Observation 41423e53-92bb-439f-b9d5-05f4325a7b4b · outbound

This paper cites Let’s verify step by step.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Let’s verify step by step

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.434667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:41.502360Z digest=sha256:7ad781f80d27ac7b98e76d2caf82e86eb0ad238b3f4f0e0eec4c49e95857b3df

Observation 768ef0a6-eb07-4df4-ba05-4b74f7a63012 · outbound

This paper cites Autopsv: Automated process-supervised verifier.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Autopsv: Automated process-supervised verifier

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.321785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:41.560493Z digest=sha256:6e6229dddd63e2ef203625705388cfb2cbdfd28a8c723ca56b8de16ad18933ea

Observation 71f147f5-9355-44db-9ffd-cf487da4571e · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.654194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.654194Z digest=sha256:b031deb5d3342b02cd3bea01a8e7fd1736d2fe3ae01fd3ce8f7f114729df3a01

Observation 2107b35b-0ad3-45ea-b7eb-01f31268c4ff · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.718662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.718662Z digest=sha256:3f783264d747ad5c2d157154edc8aa6a5c964469d5432e45bb82b7fcfdfbf3ef

Observation c138dd29-192a-467a-aada-eb6e5c7cabaa · outbound

This paper cites Skywork-o1 open series.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Skywork-o1 open series

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.213012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:41.762843Z digest=sha256:c23adeca94a7f765b0fed96552eac7807d6b1055fdcd1fdfd7c43d80a84553b6

Observation b94bc281-9988-4294-afbb-d713b7ea1435 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.911560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.911560Z digest=sha256:a91bf2309f999bca8d09adc9cdd430bddf04df93b771b44acdf0e2cf2de4db76

Observation e88e6dc7-a111-4a8e-8439-e4c2192d5cd3 · outbound

This paper cites Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.008548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.008548Z digest=sha256:63340e5e25cdeaf967ef09ed598ce4fa62b3c4cdd5d57ca14a222c2108733728

Observation 67cac0ea-15ea-4643-b7c7-f04d64faecf2 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.963514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.057605Z digest=sha256:5a442dd7d2bc01c76c4a4efc04ec9ce9951d1ad5542e3769c4e204e80301dfc1

Observation 9a29aa4f-7abc-4e3f-89ef-5b847debc3a2 · outbound

This paper cites Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Multi-step problem solving through a verifier: An empirical analysis on model-induced process supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.838113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.140691Z digest=sha256:cf22cfe2532aaf37fc4a4844f39d36687b33e1ee5a08610f27feb19d8b9dd929

Observation 67831e97-c489-455b-9ff2-72921486e6ab · outbound

This paper cites Training large language models for reasoning through reverse curriculum reinforcement learning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Training large language models for reasoning through reverse curriculum reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.758576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.205649Z digest=sha256:9474a3080a8d161225ca3a2ac9ff31bc9bbcff0a0bbca96e0e5b3dadb87c713d

Observation e37d8d9a-6eea-411f-a8c9-0c9fc9a8503f · outbound

This paper cites Evaluating mathematical reasoning beyond accuracy.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Evaluating mathematical reasoning beyond accuracy

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.633494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.272829Z digest=sha256:965f8ad1d21591d14905bf4a1366004503445906bb34d48438c8bc2820785f4d

Observation f102bf8e-4392-41a8-91f9-5159c003ec7f · outbound

This paper cites An implementation of generative prm.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels An implementation of generative prm

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.509443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.339628Z digest=sha256:9a5ee1467a3f711974f5b6aa8015f7b62fc419b8306a35ad331919a54cda731a

Observation cbc7beb3-2e64-4474-9d5b-61baa3ed4d88 · outbound

This paper cites Qwen2.5 Technical Report.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.395649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.395649Z digest=sha256:eba08e82adb31670dac13030a0cca8bfab414acd3984bbbed4bc081240ce92c8

Observation a586e8ac-6143-45c3-a3dc-3f523d1d4f41 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.449307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.449307Z digest=sha256:288909e34ffbb49188e716f673f15083cba1a99f44b3fd2eb1cdb814c6e1bcab

Observation 64cba8e6-cf3c-4a15-a310-e833bd8a7636 · outbound

This paper cites Ovm, outcome-supervised value models for planning in mathematical reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Ovm, outcome-supervised value models for planning in mathematical reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.348491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.539085Z digest=sha256:48f71db83c8077efd187544e7f41256a5870fa4363f5cbbf8d85ca493c95cb80

Observation aef94d07-53cf-45ff-8dfe-8b2e80688aae · outbound

This paper cites Free Process Rewards without Process Labels.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Free Process Rewards without Process Labels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.575866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.575866Z digest=sha256:cc70ffecf555c60dadd2b37eeed9c3a60492d3e41d72655eebf0a9cc7c43defe

Observation 944cfdcc-93f7-47b8-82ea-1109dd1f95d8 · outbound

This paper cites Generative Verifiers: Reward Modeling as Next-Token Prediction.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Generative Verifiers: Reward Modeling as Next-Token Prediction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.642734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.642734Z digest=sha256:53adc2637246a20a9a60c38d2ee733b634477a66bdb185417c2ce4b4ffd1b93b

Observation 814b695d-53dc-4b7b-b30c-487d35b11c74 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:42.722080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.722080Z digest=sha256:4e985748f5e6333ceb8272eec3dd3f8a5a8e5e0e68b924d796a9ad5c8fa04730

Observation 63478cf3-b4f8-4893-85c5-e0d097331ed0 · outbound

This paper cites ProcessBench: Identifying Process Errors in Mathematical Reasoning.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels ProcessBench: Identifying Process Errors in Mathematical Reasoning

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:06:42.769583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:42.769583Z digest=sha256:90371f276edebd5095b215d44a6c4eb23cf2acd15fe83bb99d015a8893237977

Observation 62e255ff-a0d4-48b2-8583-4ac3d54511ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.161376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:42.841133Z digest=sha256:887076330c57a11c7e6e335d331453a25ceb5a07c8b31298495e6f00f2f24cba

Observation 8e862179-7b9f-476a-ab48-33c92b429597 · outbound

This paper cites an unresolved cited work.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.086411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:06:41.854876Z digest=sha256:2677cf962426672a9415b539d69a9a2aef76d725854a147240d732481daad2f4

Pith citing papers

Observation 355ce1e2-f6ca-4d58-983a-4b788d6b6edb · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:32.948217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:32.948217Z digest=sha256:1c36fda67c1f12ac90956531aeb62639877b00eeb6907f933c7b4f9cb20604dd

Observation d1a55f1a-dc7c-49ae-9cea-5bfb614c2268 · inbound

Unsupervised Process Reward Models cites this paper.

Unsupervised Process Reward Models FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:21:18.892131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:19:04.275069Z digest=sha256:c931d3abef451028fed3dc8ef1b07605ddf942a2497c71f31714bbd43f6cc37d

Observation 2dc7b6a4-248e-44e1-b692-e4fa63a8c41b · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.933322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:4ea44b6a614acd1cc3bdc8cdea68e761c72ff20df605cb261c18ab023beed3d7

Observation 32756f27-4328-4fa0-ac5b-0dd79cc3f4d8 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation FreePRM: Training Process Reward Models Without Ground Truth Process Labels

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.481368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:112cd84e4d56851aa875a30d29b546b22001f27638bf5a8b8bed5c7d2da31e08