Pith. sign in

Paper Citation Record · LEDGER

Process Reward Model with Q-Value Rankings

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2410.11287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11287 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:31:35.195492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:51.866161Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 48f26aaa-f025-40cb-8a7f-01a2f6e712ee · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Process Reward Model with Q-Value Rankings

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:45:17.597534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:ec7f0d4ae14a8ee32b3b3af4d5b385c21c62e4888dc3286441b7243ceac26c1e

Observation 71dbc219-e964-4d7b-b9f2-442416720d72 · inbound

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning cites this paper.

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:44.016691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:49:44.016691Z digest=sha256:de179d4ce89e55f8a6043aca5864c64ab1a721f84cb0bfff19215607ef6c2859

Observation 8c3df4bf-21af-4c82-87ea-c8af838614f8 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reward Model with Q-Value Rankings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.988761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.988761Z digest=sha256:1320019215ddc49699137ccae682ba141c3d5a7c2c71b9831f17702335f54349

Observation 7443a436-cbb5-447d-b931-7f039b861a59 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Process Reward Model with Q-Value Rankings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.397141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:4149ee0ec87b5b1d04f27c8a8356be85366b6137dba16fb87291ccbff97bda19

Observation 9041ee50-4f15-4d4f-84a2-55162dffceb1 · inbound

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning cites this paper.

Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning Process Reward Model with Q-Value Rankings

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:31:35.195492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:31:35.195492Z digest=sha256:54606b644b5b708ac886ea8c8b86935f812df0822be5cc8a2667808e19d49fc4

Observation 9e0e3dc3-97e1-4e1e-b2b4-f66c1636120f · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Process Reward Model with Q-Value Rankings

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.413629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.413629Z digest=sha256:d0efb8c42ab96b56e4c9e34fb0eb76702d68c3f55003dd0972d13766a491e66e

Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · inbound

FreePRM: Training Process Reward Models Without Ground Truth Process Labels cites this paper.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.397499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.397499Z digest=sha256:cfda5ad65e96eddb43bb467c111a09f585316457ea47b9760f216c446da54323

Observation 8e28e85a-a52d-48c4-81a8-bcdde58098ea · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Process Reward Model with Q-Value Rankings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.788301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.788301Z digest=sha256:524e044255481f854c8268342b0aff2e212a160ea508a4f448e95e2ac98e414c

Observation 1dff3c38-1b04-47eb-a071-2298d56aea64 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.336123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:c941e0eded7ce460f3cf53bfe88775ec5e0b1b9004e2c4613dcd4b1c4be703fc

Observation 6c701981-9b3a-4eff-8e56-2ae0d2d47507 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.140332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.140332Z digest=sha256:2c2118f791d0aa8bdb17592f6f02c46675e902f4992f2d21b76d6204a817b698

Observation 50027bf8-af34-475d-ac06-6d1b5a69ec1e · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.714002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:4525d00c5ba495a4b5b52e2dacb294140de794329d7bfe423859966dc5fa429d

Observation afac84e0-9a3c-4970-9a04-08d9b45a8c40 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.297668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:7bc6be73eb4ed24a7cbc175fab7f8ab5efe94324cf21295b90e360b011f5cbe6

Observation 23d205d3-bf76-4552-b38b-6417f0701040 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Process Reward Model with Q-Value Rankings

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.791764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:f9d11ce1ce29a50df9bb18750a42f09b87229b2bff066ed40e17654610029602

Observation 53b48197-71da-42f8-9961-2a058b7e843c · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Process Reward Model with Q-Value Rankings

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.867763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:89d383545217c87627898bf8f06e34401f34fd858fccafbcbec9d87b35ffa736

Observation 6498a7aa-eda0-4b4b-8336-9839b323321c · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Process Reward Model with Q-Value Rankings

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:45.786097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:45.786097Z digest=sha256:77171d7357c3167652ef9b5063bf8383294c82adcac7a8345c25cc72c56b811b